Method, device, apparatus and storage medium for interaction
Patent Information
- Application Number
- CN202610728435.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-08-18
AI Technical Summary
[0008]以此方式,可以在与虚拟对象的对话交互过程中,在恰当的时刻生成并播放与当前对话相关的媒体内容,可以有效的提高信息的传递效率。另外,生成的媒体内容与第一对话内容相关联,可以提高媒体内容与对话内容的匹配程度,从而提高了媒体内容的质量。
Smart Images

Figure CN122593659A_ABST
Abstract
Description
Technical Field
[0001] The examples in this article generally relate to the field of computer science, and in particular to methods, apparatuses, devices, and computer-readable storage media for interaction. Background Technology
[0002] With the development of computer technology, various forms of electronic devices have greatly enriched people's daily lives. For example, people can use electronic devices to interact, such as generating or playing media content. How to improve interaction efficiency is a key concern. Summary of the Invention
[0003] In a first aspect, an interactive method is provided. The method includes: presenting a first interface for dialogue interaction with a virtual object; and, in response to first dialogue content of the dialogue interaction satisfying a triggering condition, playing first media content on the first interface, the first media content being generated based on the dialogue interaction, the first media content including screen content associated with the virtual object, the screen content being related to the first dialogue content.
[0004] In a second aspect, an apparatus for interaction is provided. The apparatus includes: a first presentation module configured to present a first interface for dialogue interaction with a virtual object; and a playback module configured to play first media content on the first interface in response to first dialogue content of the dialogue interaction satisfying a trigger condition. The first media content is generated based on the dialogue interaction and includes screen content associated with the virtual object, the screen content being related to the first dialogue content.
[0005] In a third aspect, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method of the first aspect.
[0008] In this way, media content relevant to the current conversation can be generated and played at appropriate times during dialogue interaction with virtual objects, effectively improving information transmission efficiency. Furthermore, associating the generated media content with the initial dialogue content enhances the matching degree between the media content and the dialogue content, thereby improving the quality of the media content.
[0009] It should be understood that the content described in this section is not intended to limit the key or important features of the examples in this article, nor is it intended to restrict the scope of the solution. Other features will become readily apparent from the following description. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the various examples herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A schematic diagram of the example environment is shown; Figures 2A to 2E Example interfaces for some scenarios are shown; Figure 3 The flowcharts show example processes of interactions in some scenarios; Figure 4 Schematic block diagrams of example devices for interaction in several scenarios are shown; and Figure 5 A block diagram of an electronic device capable of implementing multiple illustrative scenarios is shown. Detailed Implementation
[0011] The examples in the text will now be described in more detail with reference to the accompanying drawings. While some examples are shown in the drawings, it should be understood that solutions can be implemented in various forms and should not be construed as limited to the examples presented herein. Rather, these examples are provided to provide a more thorough and complete understanding of the solutions. It should be understood that the drawings and examples in this document are for illustrative purposes only and are not intended to limit the scope of protection of the solutions.
[0012] It should be noted that the headings of any section / subsection provided herein are not restrictive. Various examples are described throughout this document, and examples of any type may be included under any section / subsection. Furthermore, examples described in any section / subsection may be combined in any way with any other examples described in the same section / subsection and / or different sections / subsections.
[0013] In the description of the examples in this document, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an example" or "the example" should be understood as "at least one example". The term "some examples" should be understood as "at least some examples". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0014] The examples in this article may involve user data, data acquisition, and / or use. All of these aspects comply with relevant laws, regulations, and rules. In the examples presented here, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, when implementing each example, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained through appropriate means, in accordance with relevant laws and regulations. The specific methods of notification and / or authorization can vary depending on the actual situation and application scenario; the scope of the solution is not limited in this regard.
[0015] In this manual and the sample solutions, any processing of personal information will be conducted only under legal grounds (such as obtaining the consent of the data subject or being necessary for the performance of a contract) and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.
[0016] An interactive scheme is proposed. The scheme includes: presenting a first interface for dialogue interaction with a virtual object; and, in response to the first dialogue content of the dialogue interaction satisfying a trigger condition, playing first media content on the first interface. The first media content is generated based on the dialogue interaction and includes screen content associated with the virtual object, and the screen content is related to the first dialogue content.
[0017] In this way, media content relevant to the current conversation can be generated and played at appropriate times during dialogue interaction with virtual objects, effectively improving information transmission efficiency. Furthermore, associating the generated media content with the initial dialogue content enhances the matching degree between the media content and the dialogue content, thereby improving the quality of the media content.
[0018] The following describes various examples of this scheme in further detail with reference to the accompanying drawings.
[0019] Example Environment Figure 1 A schematic diagram of example environment 100 is shown. (e.g.) Figure 1 As shown, example environment 100 may include electronic device 110.
[0020] In this example environment 100, electronic device 110 may run an application 120 that supports interaction. Application 120 may be any suitable type of application for interaction, including but not limited to: social applications, media applications, or other suitable applications. User 140 may interact with application 120 via electronic device 110 and / or its attached devices.
[0021] exist Figure 1 In environment 100, if application 120 is active, electronic device 110 can use application 120 to present interface 150 for supporting interaction.
[0022] In some cases, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some cases, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).
[0023] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for interactive applications 120 in electronic devices 110.
[0024] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection can include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections. In some cases, server 130 and electronic device 110 can exchange signaling information through their communication connection.
[0025] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the scheme.
[0026] The following description of the example will continue with reference to the accompanying drawings.
[0027] Example Interaction Figures 2A to 2E Example interfaces 200A to 200E are shown, illustrating interactions under various scenarios. Interfaces 200A to 200E can, for example, be... Figure 1 The electronic device 110 shown is provided.
[0028] like Figure 2A As shown, the electronic device 110 can present an interface 200A (e.g., referred to as a first interface). In some cases, the first interface can be used for dialogue interaction with virtual objects.
[0029] In some cases, virtual objects can also be any suitable processing entity, such as an agent or bot. As an example, virtual objects can be implemented based on any suitable machine learning model, such as a generative model like a language model.
[0030] In some cases, the primary interface can be a conversational interface between the user and the virtual object. For example, the user can input messages using input controls in the conversational interface (the content indicated in the message is the dialogue content), and the virtual object can generate a reply message based on the user's message, or the virtual object can proactively output messages to the user. In some cases, messages associated with the virtual object and / or messages associated with the user can be any appropriate type of message, such as including, but not limited to, at least one of the following: text messages, audio messages, video messages, image messages, document messages, etc.
[0031] by Figure 2A As an example, electronic device 110 can present dialogue content 201-1 associated with virtual objects in interface 200A.
[0032] In some cases, virtual objects can be associated with any appropriate configuration information, which may include, but is not limited to, the identification information, visual appearance information, character information, voice information, etc. of the virtual object, etc., which will not be elaborated here.
[0033] In some scenarios, virtual objects can interact with users through conversations based on their own configuration information. As an example, a user can input messages related to the plot or story using the input components of interface 200A. The virtual object can then output a response message related to the plot or story based on the message and its own configuration information. Of course, users can also input other types of messages, such as knowledge-based inquiries or casual conversations, which will not be elaborated upon here.
[0034] by Figure 2B As an example, electronic device 110 can receive user-input dialogue content 201-2 via input component 202-1 of interface 200A.
[0035] It should be noted that, due to the limited amount of content that interface 200A can display at a time, messages 201-1 and 201-2 currently displayed by interface 200A are only a portion of the conversation messages between the user and the virtual object. Other conversation messages that are currently invisible can be updated and displayed in interface 200A based on preset operations associated with the conversation interface. Preset operations can be any appropriate operation, such as screen swiping operations (up and down swiping) received by interface 200A, or message search operations received by interface 200A, etc., which will not be elaborated upon here.
[0036] To improve the efficiency of information transmission, the electronic device 110 can present second media content on the first interface. The second media content can be associated with virtual objects.
[0037] In some cases, the second media content can be presented in any appropriate location. For example, the first interface may also include an area for presenting visual elements related to the virtual object, which can be configured to present the second media content associated with the virtual object.
[0038] In some cases, the second media content can be any appropriate content, such as images, videos, audio, etc.
[0039] In some cases, the second media content can be presented on the first interface at any appropriate time. For example, when outputting dialogue content to a virtual object, that is, when the message output by the virtual object is presented in the form of a bubble on the first interface, the second media content is presented simultaneously.
[0040] In some cases, the second media content can correspond to any appropriate type, such as the second type. As an example, the second media content can be pre-defined media content, characterized by low resource consumption and real-time responsiveness to user actions.
[0041] As another example, the second media content may include background images and dynamic elements. The background image may be an image from an image set that matches the virtual object. The image set may include multiple background images generated offline beforehand. Dynamic elements may be any suitable elements, such as facial expressions, motion elements, special effects elements, etc. Facial expressions may be based on facial expressions from an expression set that match the virtual object (e.g., a character's emotions, facial expressions). Special effects elements may include, but are not limited to, at least one of the following: weather effects (e.g., falling snowflakes), atmospheric effects (e.g., fluttering petals, changing light), etc.
[0042] While the virtual object outputs dialogue content, the background image of the second media content in the first interface can match the dialogue content output by the virtual object. For example, if the virtual object is describing that it is waiting for coffee in a coffee shop, the background image can be associated with the coffee kiosk. Dynamic elements can be determined by the dialogue semantic tags corresponding to the dialogue content output by the virtual object. For example, if the dialogue content output by the virtual object indicates that the currently purchased coffee has a free gift, the dynamic element can be a happy animation of the virtual object.
[0043] like Figure 2A As shown, the electronic device 110 can display media content 220 on the interface 200A. The media content 220 can include the visual image of a virtual object (e.g., a character portrait or half-body portrait), a background image (default white background, but can be any other suitable background), and weather elements (snowflake elements).
[0044] In some cases, the second media content can be lightweight, so the corresponding second type of media content can be maintained for most of the conversation time, ensuring smooth interaction and low resource consumption.
[0045] like Figure 2A As further shown, interface 200A may include input component 202-1. Users can use input component 202-1 to input text or voice messages to interact with virtual objects.
[0046] Furthermore, the electronic device 110 can determine whether each dialogue content of the dialogue interaction meets the triggering conditions based on the dialogue interaction. The triggering conditions can be any appropriate conditions, such as indicating whether the moment associated with the current dialogue content is a highlight moment. In some cases, highlight moments can indicate key nodes in the dialogue, such as plot twists, emotional climaxes, character contrasts, and dramatic changes in the scene.
[0047] In some situations, electronic device 110 can use a predetermined model to analyze each dialogue content in the dialogue interaction to determine whether each dialogue content meets the triggering conditions. To improve accuracy, for each dialogue content, electronic device 110 can use a predetermined model to analyze the dialogue content and its corresponding context information in the dialogue interaction to determine whether the dialogue content meets the triggering conditions.
[0048] Furthermore, the electronic device 110 can play first media content on the first interface in response to the first dialogue content of the dialogue interaction satisfying the triggering condition. The first dialogue content can be any appropriate content, which can be the dialogue content output by the virtual object, or the dialogue content corresponding to the current user interacting with the virtual object.
[0049] In some cases, the primary media content can be a highlight video dynamically generated based on the current dialogue interaction, and its content can be closely related to the primary dialogue content; that is, the primary media content can be generated based on the dialogue interaction.
[0050] In some cases, the primary media content may include visual content associated with the virtual object, which is related to the primary dialogue content. The visual content may include, but is not limited to, at least one of the following: close-ups of the virtual character's face, body movements, facial expressions, etc.
[0051] In some cases, the first media content can correspond to the first type, and the second media content can correspond to the second type. The first type is different from the second type.
[0052] Specifically, if the second type indicates that the second media content is preset media content, the first type can indicate that the first media content is media content generated through dialogue interaction.
[0053] For example, if the second type indicates that the second media content includes background images and dynamic elements, then the first type can indicate that the first media content is video content.
[0054] For example, the first type can indicate that the first media content is video content generated through dialogue interaction, while the second type can indicate that the second media content can be preset media content, and the preset media content can include background images and dynamic elements.
[0055] As an example, electronic device 110 can stop displaying the second media content and play the first media content. Figure 2B and Figure 2C As shown, the electronic device 110 can stop displaying media content 220 and play media content 230.
[0056] In some situations, electronic device 110 can obtain first information based on the content of the first dialogue. The first information may be related to at least one object corresponding to the dialogue interaction. For example, the first information may indicate, but is not limited to, at least one of the following: the object characteristics of the virtual object (such as the degree of contrast between its persona and its actual characteristics), and the degree of interaction between the virtual object and the current user (such as intimacy, favorability, etc.). Furthermore, the electronic device 110 can respond to the first information fulfilling the first requirement and play the first media content on the first interface.
[0057] As an example, electronic device 110 can play first media content on a first interface in response to the virtual object's corresponding object characteristics meeting a first requirement. Object characteristics may include, but are not limited to, the virtual object's and / or the current user's persona, personality, etc. For example, if the first dialogue content is the dialogue output by the virtual object, and the current first dialogue content indicates that the virtual object's image does not match the preset persona information (e.g., a normally aloof virtual character suddenly becomes shy, or a normally indifferent virtual object suddenly shows concern for the current user, etc.), then it can be determined that the object characteristics meet the first requirement.
[0058] As another example, electronic device 110 may play first media content on a first interface in response to the degree of interaction between the virtual object and the current user meeting a first requirement.
[0059] For example, if the first dialogue indicates that the virtual object is confessing its feelings to the current user, or that the virtual object and the current user are arguing or making up, or that the ambiguous relationship between the virtual object and the current user is escalating, then it can be determined that the level of interaction meets the first requirement.
[0060] In other scenarios, electronic device 110 can acquire second information based on the first dialogue content. The second information may be related to dialogue attributes of the dialogue interaction. Dialogue attributes may include, but are not limited to, at least one of the following: dialogue location, dialogue plot, etc. The dialogue location may indicate the virtual dialogue location where the virtual object is located. For example, the dialogue location may include, but is not limited to, at least one of the following: bedroom, coffee shop, rainy night street, etc. The dialogue plot may include, but is not limited to, at least one of the following: argument, confession, farewell, etc.
[0061] Furthermore, the electronic device 110 can respond to the second information satisfying the second requirement by playing the first media content on the first interface.
[0062] As an example, electronic device 110 may play first media content on a first interface in response to the dialogue location meeting the second requirement. For instance, electronic device 110 may determine that the dialogue location meets the second requirement in response to a change in the dialogue location (e.g., from indoors to outdoors, or from a public space to a private space).
[0063] As another example, electronic device 110 may play first media content on a first interface in response to the dialogue plot meeting the second requirement. For example, electronic device 110 may determine that the dialogue plot meets the second requirement in response to an unexpected turn in the dialogue plot (such as the sudden appearance of a third party, the occurrence of a sudden event, the exposure of a secret, etc.).
[0064] In some cases, electronic device 110 can obtain the rating information corresponding to the first dialogue content based on the first information and the second information.
[0065] For example, electronic device 110 can determine the rating information corresponding to the first dialogue content based on the following formula: highlight_score=w1×plot_twist_score+w2×emotional_peak_score+w3×character_contrast_score + w4×scene_drama_score; Where w1, w2, w3, and w4 are preset weights; highlight_score is the scoring information; plot_twist_score corresponds to the sub-scoring information of the plot twist dimension; emotional_peak_score corresponds to the sub-scoring information of the emotional climax dimension; character_contrast_score corresponds to the sub-scoring information of the character contrast dimension; and scene_drama_score corresponds to the sub-scoring information of the scene drama dimension. The first information is associated with the emotional climax dimension and the character contrast dimension, and the second information is associated with the plot twist dimension and the scene drama dimension.
[0066] Furthermore, the electronic device 110 can play first media content on the first interface in response to the rating information being greater than a threshold. The threshold can be any appropriate value; for example, if the value range of the sub-rating information corresponding to each dimension is [0,1], then the threshold can be 0.7.
[0067] In order to improve the presentation quality of media content, in some cases, electronic device 110 can present animated content that can indicate that the second media content is stopped from being presented and the first media content is started playing.
[0068] Specifically, before the first media content is played, the electronic device 110 can execute a dissolve transition effect to stop playing the second media content and start playing the first media content. The dissolve transition effect can indicate a smooth, gradual transition between two visual modes, which can avoid abrupt screen cuts and thus maintain the user's immersion.
[0069] For example, a dissolve transition effect can indicate that the last frame of the second media content begins to fade out (e.g., the opacity changes from 0% to 100% in 1 second), while the first frame of the first media content gradually fades in (the opacity changes from 100% to 0%).
[0070] by Figure 2D As an example, once the dissolve transition effect is complete, the electronic device 110 can play the first media content in full screen.
[0071] Taking the generation of the first media content by the server as an example, the following explains the generation process of the first media content.
[0072] To resolve the conflict between the long video generation time and the real-time nature of dialogue, electronic device 110 can send a request to the server when the rating information of a certain dialogue content is greater than a first threshold and less than a second threshold. This request can instruct the pre-allocation of computing resources (such as GPU resources) and the loading of model weights, but does not execute the complete video generation task. The first threshold is less than the second threshold; for example, the first threshold could be 0.5 and the second threshold could be 0.7. Furthermore, when electronic device 110 responds to the fact that the rating information of the first dialogue content is greater than or equal to the second threshold, the server can trigger the generation of the first media content using the video generation model.
[0073] It should be noted that the entire process does not block the front-end dialogue flow, and users can continue to have subsequent conversations with the virtual object during the generation of the first media content.
[0074] In some situations, the server can obtain third-party information based on the content of the first dialogue. This third-party information describes the scene information corresponding to the dialogue interaction. Scene information may include, but is not limited to, the virtual dialogue location where the virtual object is located, information about plot reversals, information about object contrasts, information about changes in the dialogue plot, etc.
[0075] Furthermore, the server can obtain prompt information based on third-party information.
[0076] As an example, the server can generate prompts based on a predefined model and third-party information. The predefined model can be any suitable machine learning model, such as a generative model, or of course, others.
[0077] As another example, the server can pre-configure multiple intelligent systems (Agents), each corresponding to a highlight type (e.g., a character contrast agent, an emotional climax agent, a plot reversal agent, a dramatic scene agent, etc.). Each intelligent system can also correspond to a specific plot type under a particular highlight type; for example, an Agent might correspond to a plot where the character's personality changes from lively to quiet under a character contrast type. Each intelligent system can be associated with its corresponding preset prompts, i.e., system prompts.
[0078] Furthermore, the server can respond to the third information by matching it with the first intelligent system among multiple intelligent systems, and obtain the preset prompt information corresponding to the first intelligent system. For example, if the third information indicates that the virtual object suddenly changes from playful to worried, then the third information can be matched with the character contrast agent. Furthermore, the electronic device 110 can obtain the preset prompt information corresponding to the character contrast agent.
[0079] Furthermore, the server can obtain the prompt information based on preset prompt information. Specifically, the server can directly determine the preset prompt information as the prompt information. The server can also obtain the prompt information by editing the preset prompt information.
[0080] For example, the server can combine preset prompts with dynamically acquired information (such as dialogue text, virtual object identification information, and scene identification information) to generate the final prompt.
[0081] Furthermore, the server can provide this prompt to the first intelligent system to generate the first media content. Specifically, the server can provide this prompt to the first intelligent system so that the first intelligent system can invoke its associated media generation model to generate the first media content. The media generation model can be any suitable machine learning model, such as a language model.
[0082] In some cases, the prompt information can be structured descriptive data used to guide the media generation model in generating specific media content, which may include one or more of the following: first descriptive information, second descriptive information, and third descriptive information.
[0083] In some cases, the first descriptive information may correspond to at least one storyboard, and the first media content includes at least one storyboard. The first descriptive information may also be referred to as storyboard description information. The first descriptive information may refer to a sequence of shot instructions divided along a timeline, with each shot containing a timestamp, image description, camera movement, character actions, dialogue / tone, background music description, etc.
[0084] In some cases, the second descriptive information can be information corresponding to the style of the first media content; the second descriptive information can also be referred to as video style description information. For example, the second descriptive information may include: first-person perspective, modern urban cinematic feel, emotional reversal, and high-gloss slow-motion visual blockbuster.
[0085] In some cases, third descriptive information can be associated with the visual effects of the primary media content. Third descriptive information can also be referred to as visual effect description information. For example, third descriptive information may include: cinematic soft focus, dreamy blur, low-contrast softening, etc.
[0086] To improve the quality of the generated primary media content and avoid significant differences in visuals before and after switching between them, the server can also obtain reference content. This reference content can be related to the primary dialogue content and is used to ensure that the generated primary media content closely matches the current dialogue in terms of character appearance consistency and scene accuracy.
[0087] For example, the reference content may include at least one of the following: a first image and a second image. In some cases, the first image may indicate the visual appearance of the virtual object. For example, the first image may be a three-quarter view of the virtual object or a close-up of its face. The second image may indicate the virtual scene in which the virtual object is located, such as the background image of the scene where the current dialogue takes place (e.g., a cafe, a park, a rainy night street, etc.).
[0088] Furthermore, the server can generate first media content based on the reference content and prompts. Specifically, the server can provide the reference content and prompts to the media generation model to generate the first media content.
[0089] Once the first media content is generated, the server can cache it locally or send it to electronic device 110. Furthermore, electronic device 110 can, in response to the current time meeting the playback requirement, play the first media content on the first interface. The playback requirement can indicate that the current time is the optimal time to play the first media content. The optimal playback time can include, but is not limited to: the interval between when the user has just entered a new dialogue and is waiting for a virtual object's reply, the interval between scenes or actions described in the virtual object's reply, or a brief idle period when the user actively pauses the interaction, etc.
[0090] As an example, electronic device 110 can receive second dialogue content, which is a response to the first dialogue content. To avoid interrupting the user's reading or input behavior and to ensure that the appearance of the highlight video matches the user's attention rhythm, the second dialogue content can be a response entered by the user to the first dialogue content. Furthermore, electronic device 110 can play the first media content on the first interface.
[0091] by Figure 2C and Figure 2D As an example, if the dialogue content 201-1 output by the virtual object meets the triggering condition, then after the user inputs dialogue content 201-2, the following will be displayed: Figure 2C and Figure 2D The highlight video 230 is shown. Specifically, if the dialogue content 201-1 instructs the virtual object to accidentally spill a drink, then the scene presented in the highlight video 230 may include close-ups of the spilled drink, close-ups of the virtual object's micro-expressions switching from a playful expression to a panicked and distressed expression, the virtual object's frantic action of grabbing a tissue and handing it over, and the virtual object asking in a low voice, "Are you okay?"
[0092] Furthermore, in response to completing the playback of the first media content, the electronic device 110 can present third media content on the first interface. The third media content can be associated with a virtual object. Specifically, in response to completing the playback of the first media content, the electronic device 110 can output other dialogue content, either through the virtual object or by the current user interacting with the virtual object, and at this time, the third media content can be associated with the other dialogue content.
[0093] In some cases, the third type of third media content can be the same as the second type. As one example, the third media content can be preset media content. As another example, the third media content can include background images and dynamic elements. As yet another example, the third media content can include preset media content, and the preset media content can include background images and dynamic elements.
[0094] like Figure 2E As shown, in response to the completion of playback of the first media content, the electronic device 110 can perform a reverse dissolve transition. For example, the reverse dissolve transition can instruct the last frame of the first media content to gradually fade out (e.g., within 0.8 seconds), while simultaneously restoring the presentation of the third media content, i.e., presenting media content 240. At this time, media content 240 can be associated with dialogue message 201-3 output by a virtual object, such as including dynamic elements (hearts), background images (preset white images or others) associated with dialogue message 201-3. Dialogue message 201-3 can be a response to dialogue message 201-2 input by the user.
[0095] It should be noted that the third media content corresponds to the same type as the second media content, except that the specific content (such as character expressions and scene backgrounds) may have been updated during the playback of the first media content. For example, during the playback of the first media content, the user may not have sent a new message, but the virtual object may have generated new dialogue content, and the electronic device 110 can match the corresponding character expressions and backgrounds based on the latest dialogue content.
[0096] To facilitate quick access to historical media content, in some cases, the electronic device 110 may, in response to receiving a viewing request, display at least one piece of historical media content on a first interface.
[0097] In some cases, at least one piece of historical media content may include first media content, and at least one piece of historical media content is generated in response to historical dialogue content meeting triggering conditions. Historical media content also corresponds to the first type. Specifically, the generation process of historical media content is the same as that of first media content, and will not be elaborated upon here.
[0098] As an example, the electronic device 110 may present a viewing entry point on a first interface, which is used to view historical media content. Furthermore, the electronic device 110 may receive a viewing request in response to a user clicking on the viewing entry point.
[0099] In some cases, at least one piece of historical media content can be presented in multiple ways.
[0100] In some examples, electronic device 110 can present at least one piece of media content based on the generation time of at least one piece of historical media content. For example, electronic device 110 can sort and display at least one piece of historical media content in reverse chronological order.
[0101] In other examples, the trigger conditions may be associated with different trigger types (e.g., the first trigger type is "character contrast highlight", and the second trigger type is "emotional climax highlight"). At least one piece of historical media content may include historical media content corresponding to the first trigger type and historical media content corresponding to the second trigger type.
[0102] At this time, the electronic device 110 can display a first content item on the first interface, which may include historical media content corresponding to the first trigger type. The electronic device 110 can also display a second content item on the first interface, which may include historical media content corresponding to the second trigger type.
[0103] In some cases, the first content item and the second content item can be presented in different areas of the first interface.
[0104] Furthermore, the electronic device 110 can replay the corresponding historical media content in response to receiving a selection of a certain historical media content.
[0105] In this way, media content relevant to the current conversation can be generated and played at appropriate times during dialogue interaction with virtual objects, effectively improving information transmission efficiency. Furthermore, associating the generated media content with the initial dialogue content enhances the matching degree between the media content and the dialogue content, thereby improving the quality of the media content.
[0106] Example process Figure 3 A flowchart of an example process 300 for interaction under certain conditions is shown. Process 300 can be implemented at electronic device 110. See below for reference. Figure 1 To describe process 300.
[0107] like Figure 3 As shown, in box 310, the electronic device 110 presents a first interface, which is used for dialogue interaction with virtual objects.
[0108] In frame 320, electronic device 110 responds to the first dialogue content of the dialogue interaction meeting the triggering condition and plays the first media content on the first interface. The first media content is generated based on the dialogue interaction and includes screen content associated with virtual objects. The screen content is related to the first dialogue content.
[0109] In this way, media content relevant to the current conversation can be generated and played at appropriate times during dialogue interaction with virtual objects, effectively improving information transmission efficiency. Furthermore, associating the generated media content with the initial dialogue content enhances the matching degree between the media content and the dialogue content, thereby improving the quality of the media content.
[0110] In some cases, playing first media content on a first interface includes: presenting second media content on the first interface, the second media content being associated with a virtual object; and stopping the presentation of the second media content and playing the first media content, the first type of the first media content being different from the second type of the second media content.
[0111] In this way, second media content associated with the virtual object can be presented first on the first interface, and then the user can switch to the first media content. The two media contents are of different types, but both are related to the virtual object, which avoids visual conflict and maintains the user's immersion.
[0112] In some cases, the first type indicates that the first media content is media content generated through dialogue interaction, and the second type indicates that the second media content is preset media content.
[0113] In this way, by differentiating the generation methods, highlight videos can be highly personalized, while regular content can maintain low latency and low resource consumption, thus balancing content freshness and system performance.
[0114] In some cases, Type 1 indicates that the first media content is video content, and Type 2 indicates that the second media content includes background images and dynamic elements.
[0115] In some cases, process 300 may also include: in response to the completion of playing the first media content, presenting third media content on the first interface, wherein the third type of the third media content is the same as the second type.
[0116] In this way, after the first media content has finished playing, the third media content of the same type as the second media can be presented, ensuring a seamless return to the normal interaction mode after the highlight moment, maintaining narrative continuity, and improving interaction efficiency.
[0117] In some cases, playing the first media content on the first interface includes: receiving second dialogue content, which is a response to the first dialogue content; and playing the first media content on the first interface.
[0118] In this way, highlight videos can be inserted during natural gaps in the conversation (such as the gap while waiting for an AI response), avoiding interruptions to the user's reading or input behavior and improving the naturalness of the playback timing.
[0119] In some cases, playing first media content on the first interface in response to the first dialogue content of the dialogue interaction meeting the triggering condition includes: obtaining first information based on the first dialogue content, wherein the first information is related to at least one object corresponding to the dialogue interaction; and playing the first media content on the first interface in response to the first information meeting the first requirement.
[0120] In this way, the playback condition can be based on whether the first information related to the interactive object meets the first requirement, thus realizing fine-grained triggering decisions and avoiding the waste of computing power caused by invalid playback.
[0121] In some cases, playing first media content on the first interface in response to the first information satisfying the first requirement includes: playing first media content on the first interface in response to the object characteristics corresponding to the virtual object satisfying the first requirement; or playing first media content on the first interface in response to the degree of interaction between the virtual object and the current user satisfying the first requirement.
[0122] In some cases, responding to the first dialogue content of the dialogue interaction meeting the triggering condition and playing the first media content on the first interface includes: obtaining second information based on the first dialogue content, the second information being related to the dialogue attributes of the dialogue interaction; and responding to the second information meeting the second requirement and playing the first media content on the first interface.
[0123] In some cases, playing first media content on the first interface in response to the second information fulfilling the second requirement includes: playing first media content on the first interface in response to the dialogue location fulfilling the second requirement, wherein the dialogue location indicates the virtual dialogue location where the virtual object is located; or playing first media content on the first interface in response to the dialogue plot fulfilling the second requirement.
[0124] In some cases, process 300 further includes: in response to receiving a viewing request, presenting at least one piece of historical media content in a first interface, the at least one piece of historical media content including first media content, the at least one piece of historical media content being generated in response to historical dialogue content meeting triggering conditions.
[0125] This method provides a quick way to access historical media content, effectively improving the efficiency of information transmission.
[0126] In some cases, presenting at least one piece of historical media content in the first interface includes: presenting at least one piece of historical media content based on the generation time of the at least one piece of historical media content.
[0127] In this way, historical media content can be sorted and presented based on the generation time. The time index structure optimizes the retrieval efficiency of historical data, shortens the waiting time for users to find specific highlight videos, and improves the efficiency of information transmission.
[0128] In some cases, the triggering condition is associated with a first trigger type and a second trigger type, and presenting at least one piece of historical media content in the first interface includes: presenting a first content item in the first interface, the first content item including historical media content corresponding to the first trigger type; and presenting a second content item in the first interface, the second content item including historical media content corresponding to the second trigger type, wherein at least one piece of historical media content includes historical media content corresponding to the first trigger type and historical media content corresponding to the second trigger type.
[0129] In this way, historical media content can be divided into first and second content items based on the trigger type and presented separately. The time complexity of content retrieval is reduced by the classification index, and efficient filtering and loading by type is achieved, which improves the efficiency of targeted information delivery.
[0130] In some cases, the first media content is generated based on the following process: obtaining third information based on the first dialogue content, the third information being used to describe the scene information corresponding to the dialogue interaction; obtaining prompt information based on the third information; and generating the first media content based on the prompt information.
[0131] This approach enables the creation of an automated processing flow from dialogue text to video generation instructions, reducing the intervention costs associated with manual annotation of prompts and improving the efficiency of media content generation. Furthermore, obtaining prompts based on contextual information corresponding to the dialogue interaction can effectively enhance the quality of generated media content.
[0132] In some cases, obtaining prompt information based on third information includes: in response to the third information matching with a first intelligent system in a plurality of intelligent systems, obtaining preset prompt information corresponding to the first intelligent system; and obtaining prompt information based on the preset prompt information.
[0133] In this way, the conditional generation capabilities of the pre-trained intelligent system can be utilized to shorten the time for the video generation model to parse and adapt the prompt information, thereby improving the generation efficiency.
[0134] In some cases, generating first media content based on prompt information includes: providing prompt information to a first intelligent system to generate first media content.
[0135] In this way, prompts can be provided to the matching intelligent system to generate primary media content, which can effectively ensure the quality of the generated media content.
[0136] In some cases, generating first media content based on prompts includes: obtaining reference content related to the first dialogue content; and generating first media content based on the reference content and prompts; wherein the reference content includes at least one of the following: a first image indicating the visual image corresponding to the virtual object; and a second image indicating the virtual scene in which the virtual object is located.
[0137] In this way, the primary media content can be generated by combining reference content and prompts. This avoids a significant difference between the generated primary media content and the media content associated with the virtual object presented during the dialogue interaction, thus ensuring the quality of the generated primary media content.
[0138] Example devices and equipment A corresponding apparatus for implementing the above methods or processes is also provided. Figure 4A schematic structural block diagram of an example device 400 for interaction under certain scenarios is shown. Device 400 may be implemented as or included in electronic device 110. The various modules / components in device 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0139] like Figure 4 As shown, the device 400 includes: a first presentation module 410 configured to present a first interface for dialogue interaction with a virtual object; and a playback module 420 configured to play first media content on the first interface in response to the first dialogue content of the dialogue interaction meeting a trigger condition. The first media content is generated based on the dialogue interaction and includes screen content associated with the virtual object, which is related to the first dialogue content.
[0140] In some cases, the playback module 420 is also configured to: present second media content associated with a virtual object on a first interface; and stop presenting the second media content and play first media content of a different type than the second type of the second media content.
[0141] In some cases, the first type indicates that the first media content is media content generated through dialogue interaction, and the second type indicates that the second media content is preset media content.
[0142] In some cases, Type 1 indicates that the first media content is video content, and Type 2 indicates that the second media content includes background images and dynamic elements.
[0143] In some cases, device 400 also includes a second presentation module configured to: in response to completion of playback of the first media content, present a third media content on a first interface, wherein the third type of the third media content is the same as the second type.
[0144] In some cases, the playback module 420 is also configured to: receive second dialogue content, which is a response to the first dialogue content; and play the first media content on the first interface.
[0145] In some cases, the playback module 420 is also configured to: obtain first information based on the first dialogue content, wherein the first information is related to at least one object corresponding to the dialogue interaction; and play the first media content on the first interface in response to the first information satisfying the first requirement.
[0146] In some cases, the playback module 420 is also configured to: play first media content on the first interface in response to the object characteristics corresponding to the virtual object meeting the first requirement; or play first media content on the first interface in response to the degree of interaction between the virtual object and the current user meeting the first requirement.
[0147] In some cases, the playback module 420 is also configured to: obtain second information based on the first dialogue content, the second information being related to the dialogue attributes of the dialogue interaction; and play the first media content on the first interface in response to the second information meeting the second requirement.
[0148] In some cases, the playback module 420 is also configured to: play first media content on the first interface in response to the dialogue location meeting the second requirement, wherein the dialogue location indicates the virtual dialogue location where the virtual object is located; or play first media content on the first interface in response to the dialogue plot meeting the second requirement.
[0149] In some cases, device 400 also includes a history presentation module configured to: in response to receiving a viewing request, present at least one piece of historical media content in a first interface, the at least one piece of historical media content including first media content, the at least one piece of historical media content being generated in response to historical dialogue content meeting trigger conditions.
[0150] In some cases, the history presentation module is also configured to present at least one piece of historical media content based on the generation time of at least one piece of historical media content.
[0151] In some cases, the triggering condition is associated with a first trigger type and a second trigger type, and the history presentation module is also configured to: present a first content item in a first interface, the first content item including historical media content corresponding to the first trigger type; and present a second content item in a first interface, the second content item including historical media content corresponding to the second trigger type, at least one historical media content including historical media content corresponding to the first trigger type and historical media content corresponding to the second trigger type.
[0152] In some cases, the first media content is generated based on the following process: obtaining third information based on the first dialogue content, the third information being used to describe the scene information corresponding to the dialogue interaction; obtaining prompt information based on the third information; and generating the first media content based on the prompt information.
[0153] In some cases, obtaining prompt information based on third-party information includes: obtaining reference content that is related to the first dialogue content; and generating first media content based on the reference content and the prompt information.
[0154] In some cases, generating first media content based on prompt information includes: providing prompt information to a first intelligent system to generate first media content.
[0155] In some cases, generating first media content based on prompts includes: obtaining reference content related to the first dialogue content; and generating first media content based on the reference content and prompts; wherein the reference content includes at least one of the following: a first image indicating the visual image corresponding to the virtual object; and a second image indicating the virtual scene in which the virtual object is located.
[0156] The modules included in device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some cases, one or more modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 400 can be implemented at least partially by one or more hardware logic components. By way of example, and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.
[0157] Figure 5 A block diagram of an electronic device 500 in which one or more examples may be implemented is shown. It should be understood that... Figure 5 The electronic device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the examples described herein. Figure 5 The electronic device 500 shown can be used to implement the electronic device 110 discussed above.
[0158] like Figure 5 As shown, electronic device 500 is in the form of general-purpose electronic device 110. Components of electronic device 500 may include, but are not limited to, one or more processing units or processors 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processor 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.
[0159] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.
[0160] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various examples.
[0161] The communication unit 540 enables communication with other electronic devices 110 via a communication medium. Additionally, the functionality of the components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, networked personal computers, or another network node.
[0162] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. External devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device (e.g., network card, modem, etc.) that enables electronic device 500 to communicate with one or more other electronic devices 110. Such communication can be performed via input / output (I / O) interface (not shown).
[0163] A computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. A computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0164] The flowcharts and / or block diagrams of the methods, apparatus, devices, and computer program products referred to herein describe various aspects. It should be understood that each block of the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0165] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0166] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0167] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0168] Various examples have been described above. The foregoing descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. An interaction method, comprising: A first interface is presented, which is used for dialogue and interaction with virtual objects; as well as In response to the first dialogue content of the dialogue interaction satisfying the triggering condition, first media content is played on the first interface. The first media content is generated based on the dialogue interaction and includes screen content associated with the virtual object. The screen content is related to the first dialogue content.
2. The method according to claim 1, wherein playing the first media content on the first interface includes: On the first interface, second media content is presented, and the second media content is associated with the virtual object; as well as Stop presenting the second media content and play the first media content, wherein the first type of the first media content is different from the second type of the second media content.
3. The method according to claim 2, wherein the first type indicates that the first media content is media content generated through the dialogue interaction, and the second type indicates that the second media content is preset media content.
4. The method of claim 2, wherein the first type indicates that the first media content is video content, and the second type indicates that the second media content includes background images and dynamic elements.
5. The method according to claim 2, further comprising: In response to the completion of playing the first media content, a third media content is presented on the first interface, wherein the third type of the third media content is the same as the second type.
6. The method according to claim 1, wherein, Playing the first media content on the first interface includes: Receive the second dialogue content, which is a response to the first dialogue content; On the first interface, the first media content is played.
7. The method according to claim 1, wherein, The first dialogue content responding to the dialogue interaction satisfying the triggering condition, and playing the first media content on the first interface, includes: Based on the first dialogue content, first information is obtained, wherein the first information is related to at least one object corresponding to the dialogue interaction; In response to the first information satisfying the first requirement, the first media content is played on the first interface.
8. The method according to claim 7, wherein, The step of responding to the first information satisfying the first requirement and playing the first media content on the first interface includes: In response to the virtual object's corresponding object characteristics satisfying the first requirement, the first media content is played on the first interface; or In response to the virtual object's interaction level with the current user satisfying the first requirement, the first media content is played on the first interface.
9. The method according to claim 1, wherein, The first dialogue content responding to the dialogue interaction satisfying the triggering condition, and playing the first media content on the first interface, includes: Based on the first dialogue content, second information is obtained, and the second information is related to the dialogue attributes of the dialogue interaction; In response to the second information satisfying the second requirement, the first media content is played on the first interface.
10. The method of claim 9, wherein, The response that the second information satisfies the second requirement, playing the first media content on the first interface includes: In response to the dialog location satisfying the second requirement, the first media content is played on the first interface, wherein the dialog location indicates the virtual dialog location where the virtual object is located; or In response to the dialogue scenario satisfying the second requirement, the first media content is played on the first interface.
11. The method according to claim 1, further comprising: In response to receiving a viewing request, at least one piece of historical media content is presented on the first interface. The at least one piece of historical media content includes the first media content. The at least one piece of historical media content is generated in response to historical dialogue content meeting the triggering condition.
12. The method according to claim 11, wherein, The presentation of at least one piece of historical media content in the first interface includes: Based on the generation time of the at least one historical media content, the at least one historical media content is presented.
13. The method of claim 11, wherein the triggering condition is associated with a first trigger type and a second trigger type, and wherein presenting at least one piece of historical media content in the first interface comprises: In the first interface, a first content item is presented, which includes historical media content corresponding to the first trigger type; as well as In the first interface, a second content item is presented. The second content item includes historical media content corresponding to the second trigger type. The at least one historical media content includes historical media content corresponding to the first trigger type and historical media content corresponding to the second trigger type.
14. The method of claim 1, wherein the first media content is generated based on the following process: Based on the first dialogue content, third information is obtained, which is used to describe the scene information corresponding to the dialogue interaction; Based on the aforementioned third information, a prompt message is obtained; as well as Based on the prompt information, the first media content is generated.
15. The method according to claim 14, wherein, The process of obtaining the prompt information based on the third information includes: In response to the third information being matched with the first intelligent system in a plurality of intelligent systems, a preset prompt information corresponding to the first intelligent system is obtained; The prompt information is obtained based on the preset prompt information.
16. The method according to claim 15, wherein, The step of generating the first media content based on the prompt information includes: The prompt information is provided to the first intelligent system to generate the first media content.
17. The method of claim 14, wherein, The step of generating the first media content based on the prompt information includes: Obtain reference content, which is related to the content of the first dialogue; and The first media content is generated based on the reference content and the prompt information; The reference content includes at least one of the following: The first image indicates the visual representation of the virtual object; The second image indicates the virtual scene in which the virtual object is located.
18. A device for interaction, comprising: The first presentation module is configured to present a first interface, which is used for dialogue and interaction with virtual objects. as well as The playback module is configured to play first media content on the first interface in response to the first dialogue content of the dialogue interaction meeting a trigger condition. The first media content is generated based on the dialogue interaction and includes screen content associated with the virtual object, which is related to the first dialogue content.
19. An electronic device comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 17 when executed by the at least one processor.
20. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 17.