Method and apparatus for interaction in session, and device and storage medium
By generating dynamic media content within the conversation and utilizing image fusion technology from multiple participants, the problem of monotonous conversation interaction methods is solved, improving message interaction efficiency and user experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- LEMON INC(GB)
- Filing Date
- 2024-11-15
- Publication Date
- 2026-05-21
Smart Images

Figure CN2024132504_21052026_PF_FP_ABST
Abstract
Description
Methods, apparatus, devices and storage media for interaction during a session Technical Field
[0001] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to methods, apparatus, devices, and computer-readable storage media for interaction in a session. Background Technology
[0002] With the development of computer technology, more and more users are using the internet for conversations. For example, users can use instant messaging applications or instant messaging services provided by other applications to interact with other users. During these interactions, users can support multiple modalities of messaging. For example, users can send text messages, voice messages, or image messages during a conversation. Summary of the Invention
[0003] In a first aspect of this disclosure, a method for interaction in a session is provided. The method includes: in response to an interaction request from a first participant in the session, sending an interaction message in the session, the interaction message being associated with a first image indicated by the interaction request; and in response to an interaction action from a second participant in the session in response to the interaction message, presenting dynamic media content generated based on a first image and a second image in a session interface of the session, wherein the second image is determined based on the interaction action.
[0004] In a second aspect of this disclosure, an apparatus for interaction in a session is provided. The apparatus includes: a sending module configured to send an interaction message in the session in response to an interaction request from a first participant in the session, the interaction message being associated with a first image indicated by the interaction request; and a presentation module configured to present dynamic media content generated based on a first image and a second image in a session interface of the session in response to an interaction action by a second participant in the session in response to the interaction message, wherein the second image is determined based on the interaction action.
[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
[0007] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0009] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;
[0010] Figures 2A to 2C illustrate example interfaces according to some embodiments of the present disclosure;
[0011] Figures 3A to 3C illustrate example interfaces according to other embodiments of the present disclosure;
[0012] Figure 4 illustrates a flowchart of an example process of interaction in a session according to some embodiments of the present disclosure;
[0013] Figure 5 shows a schematic structural block diagram of an example device for interaction in a session according to some embodiments of the present disclosure; and
[0014] Figure 6 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure. Detailed Implementation
[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0016] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0017] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0018] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0019] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0020] As mentioned above, text and / or image interaction are important interaction methods in conversations. For example, in a conversational scenario, participants can send text messages or image messages. According to traditional methods, participants can only perform limited types of interactions on the messages sent in the conversation, such as replying or forwarding. This somewhat affects the efficiency of message interaction in a conversation.
[0021] Embodiments of this disclosure propose a scheme for in-session interaction. The scheme includes: in response to an interaction request from a first participant in the session, sending an interaction message in the session, the interaction message being associated with a first image indicated by the interaction request; and in response to an interaction operation by a second participant in the session in response to the interaction message, presenting dynamic media content generated based on a first image and a second image in the session's session interface, wherein the second image is determined based on the interaction operation.
[0022] In this way, by utilizing images provided by multiple participants in a session to generate and provide dynamic media content, embodiments of this disclosure can enrich the interaction methods in a session scenario, improve the efficiency of message interaction in a session, and thus enhance the user experience.
[0023] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.
[0024] Example Environment
[0025] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in Figure 1, the example environment 100 may include an electronic device 110.
[0026] In this example environment 100, electronic device 110 may run an application 120 that supports in-session interaction. Application 120 may be any suitable type of application for in-session interaction, examples of which may include, but are not limited to, instant messaging applications or other suitable applications that provide instant messaging services. User 140 may interact with application 120 via electronic device 110 and / or its attached devices.
[0027] In environment 100 of Figure 1, if application 120 is active, electronic device 110 can use application 120 to present interface 150 for supporting interaction in the session.
[0028] In some embodiments, electronic device 110 communicates with server 130 to provide services to application 120. Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 can also support any type of user-facing interface (such as "wearable" circuitry).
[0029] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 in electronic devices 110 that support in-session interaction.
[0030] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, server 130 and electronic device 110 can achieve signaling interaction through the communication connection between them.
[0031] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0032] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.
[0033] Example Interaction
[0034] Figures 2A to 2C illustrate example interfaces 200A to 200C according to some embodiments of the present disclosure. Interfaces 200A to 200C may be provided, for example, by the electronic device 110 shown in Figure 1. As an example, interfaces 200A to 200C may correspond to a first participant in a session (e.g., user A).
[0035] It should be understood that the conversation interface shown below corresponds to an example conversation between two participants (e.g., a one-on-one chat conversation), and the embodiments of this disclosure can also be applied to conversation scenarios involving multiple participants (e.g., a group chat conversation).
[0036] As shown in Figure 2A, the session interface 200A (also referred to as the first session interface) can correspond to a session between the current user (e.g., user A) and another participant (e.g., user B). As shown, the session interface 200A may include a message area 210 for displaying messages sent and received during the session. Additionally, the session interface 200A may also include a control area 220, which can provide one or more interactive controls associated with the session, such as interactive control 221.
[0037] In some embodiments, the electronic device 110 may receive a first operation from a user on the interactive control 221. For example, the electronic device 110 may receive a click or other appropriate operation from a user on the interactive control 221. Accordingly, in response to receiving the first operation, the electronic device 110 may present an interface 200B as shown in FIG2B.
[0038] In some embodiments, the electronic device 110 can acquire a first image associated with a first participant (e.g., user A) via an interface 200B. Specifically, as shown in FIG2B, the interface 200B may include a shooting control 230. For example, the interface 200B may display a real-time image captured by an image acquisition device and may capture the first image based on the user's triggering of the shooting control 230.
[0039] Additionally, as shown in Figure 2B, interface 200B may also provide an upload control 240. As an example, electronic device 110 may present a set of candidate images based on the user's selection of the upload control 240. Such a set of candidate images may include, for example, an image library local to electronic device 110 or an online image library associated with application 120. It should be understood that the acquisition and use of such candidate images are performed with the user's knowledge and authorization.
[0040] Furthermore, the electronic device 110 can receive the selection of at least one image from the group of candidate images by the current user (i.e., the first participant) as the first image associated with the current user.
[0041] Additionally, the electronic device 110 can also determine whether the image currently provided by the user (e.g., a captured image or an uploaded image) meets preset requirements. Such preset requirements may be related to, for example, the content, quality, and / or size of the image. For instance, preset requirements may include the requirement that the image contain a specific type of object.
[0042] In response to acquiring a first image associated with the current user, electronic device 110 can trigger an interaction request associated with the first image in the session. As will be described below with reference to FIG2C, the interaction request can trigger the generation of an interaction message corresponding to the interaction request in the session.
[0043] In some embodiments, user A (i.e., the first participant) may also initiate an interaction request associated with the first image in other ways. For example, electronic device 110 may receive a user's request to send image content in a session and accordingly provide one or more sending modes associated with the image content. As an example, in a first sending mode, the image content may be sent as an image message in the session. In another example, in response to receiving a selection of a second sending mode, electronic device 110 may trigger an interaction request associated with the image content to generate a corresponding interactive message, rather than a regular image message.
[0044] Furthermore, as shown in FIG2C, the electronic device 110 can present the interactive message 250 generated based on the interactive request in the session interface 200C. As shown, the interactive message 250 can, for example, present preset text content to indicate a media interaction request initiated with the first participant (e.g., user A).
[0045] Alternatively, the interactive message 250 may also present at least a portion of the first image indicated by the interactive request. For example, the interactive message 250 may be presented in a message card style, and the first image may be used to fill at least a portion of the background of the message card.
[0046] Additionally, as shown in Figure 2A, the interactive message 250 may also include a control 260. As an example, the electronic device 110 may also receive a selection of the control 260 to obtain an additional image associated with the first participant, triggering the generation of corresponding dynamic media content. The specific generation process of the dynamic media content will be described in detail below.
[0047] In some embodiments, for interactive messages presented in a session interface associated with the first participant, the electronic device 110 may, for example, not provide control 260 or disable control 260.
[0048] Figures 3A to 3C illustrate example interfaces 300A to 300C according to some embodiments of the present disclosure. Interfaces 300A to 300C may be provided, for example, by the electronic device 110 shown in Figure 1. As an example, interfaces 300A to 300C may correspond to a second participant in a session (e.g., user B).
[0049] As shown in Figure 3A, the electronic device 110 can present interactive messages 320 in the session interface 300A. As an example, interactive messages 320 can be generated based on interactive requests (e.g., from user A) described above with reference to Figures 2A to 2C.
[0050] Additionally, the electronic device 110 can receive preset operations from a second participant (e.g., user B) for the interactive message 320 and accordingly present the interface 300B as shown in FIG3B.
[0051] For example, as shown in FIG3A, interactive message 320 may include control 330. Electronic device 110 may receive user selection of control 330 to present interface 300B. As shown in FIG3B, electronic device 110 may present interactive panel 340, which may, for example, display a first image 350 associated with interactive message 320. As discussed above, the first image 350 may, for example, be associated with a first participant in the session (e.g., user A).
[0052] Additionally, as shown in Figure 3B, the interactive panel 340 may also provide a control 360. As an example, the electronic device 110 may receive a selection of the control 360 and may accordingly acquire a second image associated with the second participant.
[0053] As an example, electronic device 110 may provide an interface similar to interface 200B shown in FIG2B for acquiring a second image associated with a second participant. For example, the second image may include an image captured using an image acquisition device, or it may also include an image uploaded by the second participant.
[0054] In response to acquiring a second image associated with a second participant, electronic device 110 can trigger the generation of dynamic media content based on the first and second images. Accordingly, as shown in FIG2C, electronic device 110 can display the generated dynamic media content 370 in session interface 200C.
[0055] Thus, the electronic device 110 can acquire a second image associated with the second participant based on the second participant's interactive operation (e.g., selection of control 330 and image capture or upload operation), and trigger the generation of dynamic media content based on multiple images (e.g., first image and second image) associated with the participants in the session.
[0056] In some embodiments, as shown in FIG3C, in response to the completion of dynamic media content 370 generation, electronic device 110 may present the generated dynamic media content 370 in session interface 300C. In some embodiments, as shown in FIG3C, after the dynamic media content 370 is generated, electronic device 110 may, for example, replace interactive message 320 in session interface 300A with dynamic media content 370.
[0057] In other embodiments, after the dynamic media content 370 is generated, the electronic device 110 may, for example, stop displaying the interactive message 320. Furthermore, the electronic device 110 may present the dynamic media content 370 in a message area, the display position of which may be independent of the interactive message 320. For example, the interactive message 320 may be moved to a different position due to subsequent message interactions received in the session. Accordingly, after the dynamic media content 370 is generated, the electronic device 110 may, for example, present the generated dynamic media content 370 below the latest message in the message area to prevent the user from missing the generated dynamic media content.
[0058] It should be understood that, although Figure 3C shows the session interface of the second participant, the session interface associated with the first participant (e.g., interface 200C as shown in Figure 2C) can also be updated similarly to present the generated dynamic media content.
[0059] Considering that dynamic media content is generated asynchronously, after the dynamic media content 370 is generated, the electronic device 110 associated with the first or second participant can also send a reminder message related to the dynamic media content 370 to the first or second participant in the session. As an example, the reminder message may include, but is not limited to, graphic prompts, voice prompts, vibration prompts, etc.
[0060] For example, after the dynamic media content 370 is generated, the electronic device 110 can display a notification message on the system desktop, regardless of whether the user is currently accessing application 120 or its session interface. Accordingly, the electronic device 110 can receive the user's selection of the notification message and can jump to the session interface to display the generated dynamic media content. For example, the session interface can be automatically positioned at the location of the generated dynamic media content.
[0061] Additionally or alternatively, before completing the generation of dynamic media content 370, the electronic device 110 may also present dynamic information related to the generation process of the dynamic media content in a session interface. As an example, the electronic device 110 may update the interactive messages presented in the session interfaces of the first and second participants with generation progress information. As an example, the progress information may indicate the completed progress of the generation process (e.g., 50%), the remaining progress of the generation process (e.g., remaining time), etc. It should be understood that the progress information can be presented using any suitable form, examples of which may include, but are not limited to, progress bars, percentage numbers, etc.
[0062] In this way, the embodiments of this disclosure can better help users perceive the generation status and results of dynamic media content, thereby improving the efficiency of content acquisition and interaction.
[0063] The specific process of generating dynamic media content 370 will be further described below. In some embodiments, dynamic media content can be generated by electronic device 110 and / or server 130. For ease of description, server 130 will be used as an example to describe the specific process of generating dynamic media content 370 below.
[0064] For ease of description, the process of fusing and generating dynamic media content 370 will be illustrated below using two images as examples. It should be understood that such a generation process can also be applied to the fusion of images associated with more participants (e.g., three or more). Such images may include still images and / or dynamic video content.
[0065] In some embodiments, server 130 can generate target background content by fusing the first background content of the first image and the second background content of the second image. Specifically, server 130 can perform segmentation processing on the first image and / or the second image to correspond to foreground content and background content in the image.
[0066] Furthermore, server 130 can stitch together the first background content and the second background content to generate intermediate background content. As an example, server 130 can use feature point matching or image stitching algorithms for background alignment. Additionally, server 130 can use pyramid fusion technology to smoothly transition the stitched areas, thereby avoiding abruptness.
[0067] Furthermore, server 130 can also generate target background content by filling in at least one missing area in the intermediate background content. As an example, when the stitched intermediate background content has missing parts, server 130 can use in-painting techniques to fill in the missing areas, ensuring the continuity and naturalness of the target background content. It should be understood that server 130 can utilize any appropriate technique, such as a generative model, to fill in the missing areas.
[0068] Furthermore, server 130 can generate an intermediate image based on the target background content, the first foreground content of the first image, and the second foreground content of the second image. As an example, such an intermediate image can be a static image.
[0069] Additionally, server 130 can generate the dynamic media content based on the intermediate image. For example, server 130 can provide the intermediate image to a generative model to generate dynamic media content, such as video or animated images. As an example, such a generative model may include an image-to-video model.
[0070] In some embodiments, the generative model can also process intermediate images based on cue words to generate dynamic media content. Such cue words may include, for example, a preset first cue word, such as a cue word pre-configured by application 120.
[0071] Alternatively, the prompt may also include a second prompt determined based on input information from the first participant and / or the second participant. As an example, when providing the associated image, the first or second participant may also provide input information for determining the prompt. Such input information may, for example, include generation parameters for generating dynamic media content.
[0072] For example, after uploading the images, the first or second participant can specify that the desired style of the generated dynamic media content is cartoon-style. Alternatively, after uploading the images, the first or second participant can specify that the objects included in the two images perform specific motion actions, such as hugging or shaking hands.
[0073] Additionally, such generation parameters may include any other appropriate parameters suitable for guiding the generative model to generate dynamic media content. The first or second participant can specify such generation parameters, for example, through appropriate interactive methods such as entering text prompts, selecting preset labels, or adjusting parameter sizes.
[0074] Based on the process described above, by utilizing images provided by multiple participants in the session to generate and provide dynamic media content, embodiments of this disclosure can enrich the interaction methods in the session scenario, improve the message interaction rate in the session, and thus enhance the user experience.
[0075] Example process
[0076] Figure 4 illustrates a flowchart of an example process 400 of session interaction according to some embodiments of the present disclosure. Process 400 may be implemented at electronic device 110. Process 400 will now be described with reference to Figure 1.
[0077] As shown in Figure 3, in block 410, electronic device 110 responds to an interaction request from a first participant in the session by sending an interaction message in the session, the interaction message being associated with the first image indicated by the interaction request.
[0078] In box 420, electronic device 110 responds to an interactive action by a second participant in the session to an interactive message, and presents dynamic media content generated based on a first image and a second image in the session interface of the session, wherein the second image is determined based on the interactive action.
[0079] In some embodiments, the interaction request is triggered based on the following process: presenting a first session interface of the session to a first participant; receiving a first operation from the first participant on an interactive control in the first session interface; and in response to receiving the first operation, acquiring a first image associated with the first participant to trigger the interaction request.
[0080] In some embodiments, process 400 further includes: in response to obtaining a first image associated with a first participant, presenting an interactive message generated based on the first image in a first session interface.
[0081] In some embodiments, the second image is determined based on the following process: presenting a second session interface of the session to the second participant; presenting interactive messages in the second session interface; and acquiring a second image associated with the second participant based on a preset operation for the interactive messages.
[0082] In some embodiments, presenting dynamic media content generated based on the first image and the second image in the session interface includes: in response to the completion of dynamic media content generation, updating the interactive message presented in the session interface to the generated dynamic media content.
[0083] In some embodiments, presenting dynamic media content generated based on the first image and the second image in the session interface includes: in response to the completion of dynamic media content generation, triggering the sending of a reminder message associated with the dynamic media content to the first participant and / or the second participant.
[0084] In some embodiments, process 400 further includes: during the generation of dynamic media content, presenting the generation progress information of the dynamic media content in a session interface.
[0085] In some embodiments, the first image and / or the second image include picture content and / or video content.
[0086] In some embodiments, the dynamic media content is also generated based on input information obtained from a first participant and / or a second participant, the input information indicating the generation parameters of the dynamic media content.
[0087] In some embodiments, dynamic media content is generated based on the following process: generating target background content by fusing first background content of a first image and second background content of a second image; generating an intermediate image based on the target background content, first foreground content of the first image, and second foreground content of the second image; and generating dynamic media content based on the intermediate image.
[0088] In some embodiments, generating target background content by fusing the first background content of a first image and the second background content of a second image includes: splicing the first background content and the second background content to generate intermediate background content; and generating target background content by filling at least one empty area in the intermediate background content.
[0089] In some embodiments, generating dynamic media content based on an intermediate image includes providing an intermediate image and cue words to a media generation model to generate dynamic media content.
[0090] In some embodiments, the prompt word includes: a preset first prompt word; or a second prompt word determined based on input information from a first participant and / or a second participant.
[0091] Example devices and equipment
[0092] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 5 shows a schematic structural block diagram of an example apparatus 500 for session interaction according to certain embodiments of this disclosure. Apparatus 500 may be implemented as or included in electronic device 110. The various modules / components in apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.
[0093] As shown in Figure 5, the device 500 includes: a sending module 510 configured to send an interactive message in the session in response to an interaction request from a first participant in the session, the interactive message being associated with a first image indicated by the interaction request; and a presentation module 520 configured to present dynamic media content generated based on the first image and the second image in the session's session interface in response to an interactive operation by a second participant in the session in response to the interactive message, wherein the second image is determined based on the interactive operation.
[0094] In some embodiments, the interaction request is triggered based on the following process: presenting a first session interface of the session to a first participant; receiving a first operation from the first participant on an interactive control in the first session interface; and in response to receiving the first operation, acquiring a first image associated with the first participant to trigger the interaction request.
[0095] In some embodiments, the apparatus 500 further includes an interactive message generation module configured to present an interactive message generated based on the first image in a first session interface in response to acquiring a first image associated with a first participant.
[0096] In some embodiments, the second image is determined based on the following process: presenting a second session interface of the session to the second participant; presenting interactive messages in the second session interface; and acquiring a second image associated with the second participant based on a preset operation for the interactive messages.
[0097] In some embodiments, presenting dynamic media content generated based on the first image and the second image in the session interface includes: in response to the completion of dynamic media content generation, updating the interactive message presented in the session interface to the generated dynamic media content.
[0098] In some embodiments, presenting dynamic media content generated based on the first image and the second image in the session interface includes: in response to the completion of dynamic media content generation, triggering the sending of a reminder message associated with the dynamic media content to the first participant and / or the second participant.
[0099] In some embodiments, the device 500 further includes a generation progress information module, configured to present generation progress information of dynamic media content in a session interface during the generation of dynamic media content.
[0100] In some embodiments, the first image and / or the second image include picture content and / or video content.
[0101] In some embodiments, the dynamic media content is also generated based on input information obtained from the first participant and / or the second participant, prompting the generation parameters of the dynamic media content.
[0102] In some embodiments, dynamic media content is generated based on the following process: generating target background content by fusing first background content of a first image and second background content of a second image; generating an intermediate image based on the target background content, first foreground content of the first image, and second foreground content of the second image; and generating dynamic media content based on the intermediate image.
[0103] In some embodiments, generating target background content by fusing the first background content of a first image and the second background content of a second image includes: splicing the first background content and the second background content to generate intermediate background content; and generating target background content by filling at least one empty area in the intermediate background content.
[0104] In some embodiments, generating dynamic media content based on an intermediate image includes providing an intermediate image and cue words to a media generation model to generate dynamic media content.
[0105] In some embodiments, the prompt word includes: a preset first prompt word; or a second prompt word determined based on input information from a first participant and / or a second participant.
[0106] As shown in Figure 6, the electronic device 600 is in the form of a general-purpose electronic device. Components of the electronic device 600 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 600.
[0107] Electronic device 600 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 600.
[0108] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0109] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0110] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0111] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0112] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0113] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0114] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0115] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0116] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for interaction during a session, comprising: In response to an interaction request from a first participant in the session, an interaction message is sent in the session, the interaction message being associated with a first image indicated by the interaction request; as well as In response to an interactive action by a second participant in the session regarding the interactive message, dynamic media content generated based on the first image and the second image is presented in the session interface of the session, wherein the second image is determined based on the interactive action.
2. The method of claim 1, wherein the interaction request is triggered based on the following process: Present the first session interface of the session to the first participant; Receive the first operation from the first participant on the interactive controls in the first session interface; and In response to receiving the first operation, the first image associated with the first participant is acquired to trigger the interaction request.
3. The method according to claim 2, further comprising: In response to obtaining the first image associated with the first participant, the interactive message generated based on the first image is presented in the first session interface.
4. The method of claim 1, wherein the second image is determined based on the following process: Present the second session interface of the session to the second participant; The interactive message is presented in the second session interface; and Based on a preset operation for the interactive message, the second image associated with the second participant is obtained.
5. The method according to claim 1, wherein presenting dynamic media content generated based on the first image and the second image in the session interface of the session includes: In response to the completion of the dynamic media content generation, the interactive message presented in the session interface is updated to the generated dynamic media content.
6. The method according to claim 5, wherein presenting dynamic media content generated based on the first image and the second image in the session interface of the session includes: In response to the completion of the dynamic media content generation, a reminder message associated with the dynamic media content is sent to the first participant and / or the second participant.
7. The method according to claim 1, further comprising: During the generation of the dynamic media content, the generation progress information of the dynamic media content is displayed in the session interface.
8. The method of claim 1, wherein the dynamic media content is further generated based on input information obtained from the first participant and / or the second participant, the input information indicating generation parameters of the dynamic media content.
9. The method of claim 1, wherein the dynamic media content is generated based on the following process: The target background content is generated by fusing the first background content of the first image and the second background content of the second image; An intermediate image is generated based on the target background content, the first foreground content of the first image, and the second foreground content of the second image; as well as The dynamic media content is generated based on the intermediate image.
10. The method of claim 9, wherein generating the target background content by fusing the first background content of the first image and the second background content of the second image comprises: The first background content and the second background content are combined to generate the intermediate background content; as well as The target background content is generated by filling in at least one empty area in the intermediate background content.
11. The method of claim 9, wherein generating the dynamic media content based on the intermediate image comprises: The intermediate images and prompts are provided to the media generation model to generate the dynamic media content.
12. The method of claim 11, wherein the prompt word comprises: The preset first prompt word; or The second prompt word is determined based on the input information of the first participant and / or the second participant.
13. An apparatus for interaction during a session, comprising: The sending module is configured to send an interactive message in the session in response to an interaction request from a first participant in the session, the interactive message being associated with a first image indicated by the interaction request; as well as The presentation module is configured to, in response to an interactive action by a second participant in the session on the interactive message, present dynamic media content generated based on the first image and the second image in the session interface of the session, wherein the second image is determined based on the interactive action.
14. An electronic device comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 12 when executed by the at least one processing unit.
15. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 1 to 12.