Information interaction method and device, electronic equipment, storage medium and program product
Through the generative model, the problem of single interaction mode in multi-terminal interaction scenarios is solved, and diversified interactive experience and higher correlation are achieved.
Patent Information
- Application Number
- CN202510639417.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-19
AI Technical Summary
In multi-terminal interaction scenarios, the interaction method is single and the effect is monotonous, which affects the information interaction effect.
Through the generative model, multi-end media data is processed based on the interaction prompt information, and a third media data that can present multi-end interaction methods can be generated.
It enhances the interaction participation and relevance between multiple ends, enriches the interactive form and display effect, and improves the interactive experience.
Smart Images

Figure CN120508226A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of computer processing technology, and more particularly to an information interaction method, apparatus, electronic device, storage medium, and program product. Background Art
[0002] With the continuous development of network technology, various multi-terminal interaction scenarios are emerging. In multi-terminal interaction scenarios, the scene functions and the interaction methods with objects are often enriched through various interactive services associated with them. For example, the interaction between the live broadcast object and the live broadcast viewing object can be realized through the live broadcast platform.
[0003] In related technologies, the interaction method between multiple terminals in a multi-terminal interaction scenario is relatively simple. Usually, the interaction is carried out by triggering preset interactive elements and displaying them in a preset manner. The presentation effect is relatively fixed, which makes the interaction effect relatively monotonous, thereby affecting the information interaction effect. Summary of the Invention
[0004] The present disclosure provides an information interaction method, device, electronic device, storage medium, and program product to realize, in a multi-terminal interaction scenario, processing of media data from multiple terminals based on interaction prompt information through a generative model to generate third media data that can present the effect of the media data from multiple terminals interacting according to a preset interaction method.
[0005] In a first aspect, an embodiment of the present disclosure provides an information interaction method, the method comprising:
[0006] In response to an interaction request initiated in a multi-terminal interaction scenario, obtaining first media data of a first terminal initiating the interaction request and second media data of a second terminal participating in the interaction;
[0007] Obtaining first prompt information corresponding to the interaction request; wherein the first prompt information includes description data of an interaction method for at least a portion of the first media data and at least a portion of the second media data;
[0008] Third media data is generated according to the first prompt information, the first media data, the second media data, and a media data generation model, and the third media data is displayed on the first end and / or the second end.
[0009] In a second aspect, an embodiment of the present disclosure further provides an information interaction device, the device comprising:
[0010] A media data acquisition module, configured to respond to an interaction request initiated in a multi-terminal interaction scenario and acquire first media data of a first terminal initiating the interaction request and second media data of a second terminal participating in the interaction;
[0011] a prompt information acquisition module, configured to acquire first prompt information corresponding to the interaction request; wherein the first prompt information includes description data of an interaction method for at least a portion of the first media data and at least a portion of the second media data;
[0012] The media data generation module is configured to generate third media data based on the first prompt information, the first media data, the second media data, and a media data generation model, and display the third media data on the first end and / or the second end.
[0013] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:
[0014] one or more processors;
[0015] a storage device for storing one or more programs,
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the information interaction method as described in any one of the embodiments of the present disclosure.
[0017] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to execute the information interaction method as described in any one of the embodiments of the present disclosure.
[0018] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, including a computer program, which, when executed by a processor, implements the information interaction method as described in any one of the embodiments of the present disclosure.
[0019] The technical solution of the embodiment of the present disclosure supports obtaining media data of both parties of the interaction by responding to an interaction request initiated in a multi-terminal interaction scenario, obtaining first media data of the first end initiating the interaction request and second media data of the second end participating in the interaction, thereby providing a data basis for the subsequent generation of third media data, and achieving the effect of responding to interaction requests in a multi-terminal interaction scenario; further, obtaining first prompt information corresponding to the interaction request provides a generation basis for the subsequent generation of third media data, and the first prompt information corresponding to the interaction request reflects the degree of association between the first prompt information and the first end initiating the interaction request, achieving the effect of supporting customized editing of the first prompt information, enhancing the interactive participation between multiple terminals, and because the first prompt information includes the interactive method of at least part of the first media data and at least part of the second media data The description data enables the user to quickly know how to process the first media data and the second media data when subsequently generating the third media data, thereby improving the generation efficiency of the third media data, and the clear interaction method makes the interaction method in the generated third media data more standardized; further, the third media data is generated according to the first prompt information, the first media data, the second media data and the media data generation model, and the third media data is displayed on the first end and / or the second end, thereby realizing the effect of processing the first media data and the second media data according to the first prompt information through the media data generation model to generate the third media data, and displaying the third media data, thereby improving the degree of correlation between the third media data and the first end and the second end, enriching the presentation effect of the third media data, and displaying the third media data, providing a richer and more diverse interactive experience for multiple ends participating in the interaction. The technical solution provided by the embodiments of the present disclosure solves the technical problems in the related art of relatively single interaction methods and monotonous interaction effects. It realizes that in a multi-terminal interaction scenario, media data from multiple terminals are processed by a generative model based on interaction prompt information to generate third media data that can present the effect of the media data from multiple terminals interacting according to a preset interaction method. This enhances the diversity of the third media data, enhances the correlation between the third media data and the multiple terminals participating in the interaction and the multi-terminal interaction scenario, improves the authenticity of the third media data, enriches the interactive participation forms in the multi-terminal interaction scenario and the display effect of the third media data, and thus improves the multi-terminal interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0021] Figure 1 A flowchart of an information interaction method provided in an embodiment of the present disclosure;
[0022] Figure 2 A schematic diagram of an interface including a prompt information entry area provided in an embodiment of the present disclosure;
[0023] Figure 3 A schematic diagram of a display interface including optional prompt information provided in an embodiment of the present disclosure;
[0024] Figure 4 A flowchart of another information interaction method provided by an embodiment of the present disclosure;
[0025] Figure 5 A schematic diagram of an information interaction process provided by an embodiment of the present disclosure;
[0026] Figure 6 A schematic diagram of the structure of an information interaction device provided by an embodiment of the present disclosure;
[0027] Figure 7 A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure for implementing the embodiment of the present disclosure is provided. DETAILED DESCRIPTION
[0028] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0029] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0030] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0031] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0032] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0033] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0034] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0035] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0036] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0037] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0038] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0039] Figure 1This is a flow chart of an information interaction method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to scenarios in which information interaction is performed based on media data from multiple terminals in a multi-terminal interaction scenario. The method can be executed by an information interaction device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, which can be a mobile terminal, a PC, or a server.
[0040] like Figure 1 As shown, the method of this embodiment may specifically include:
[0041] S110: In response to an interaction request initiated in a multi-terminal interaction scenario, obtain first media data of a first terminal initiating the interaction request and second media data of a second terminal participating in the interaction.
[0042] Typically, for application software that supports multi-terminal network communication, when the current end can communicate with at least one end based on the application software, and during the network communication process, information can also be exchanged between the multiple ends, then the scenario in which multiple ends participate in the interaction can be referred to as a multi-terminal interaction scenario. In the disclosed embodiment, the multi-terminal interaction scenario can be understood as a scenario in which network communication and / or information interaction is conducted between any two or more terminal devices. The multi-terminal interaction scenario may include a variety of scenarios that support multi-terminal interaction, optionally including at least a live broadcast scenario, a video call scenario, and / or a video conferencing scenario.
[0043] Among them, the interaction request can be understood as an instruction for requesting to perform the operation of interacting with the other end in ongoing communication. The interaction request can include multiple initiation methods. Optionally, when a trigger operation for a preset interactive control is detected, an interaction request can be initiated; or, when it is detected that the received audio information includes a trigger word associated with initiating interaction, an interaction request can be initiated; or, when it is detected that the body movement of an object of a preset type within the field of view is consistent with the preset interaction trigger action, an interaction request can be initiated; or, when an input interaction initiation instruction is received, an interaction request can be initiated.
[0044] Among them, the first end can be an interaction initiating end, that is, the device end to which the initiator of the interactive operation belongs. The first end can be an electronic device pre-installed with application software that supports communication interaction and / or information interaction functions. Optionally, the first end includes at least one of a mobile terminal, a PC, a tablet computer, and a virtual reality device. The first media data can be media data obtained based on the first end for generating data with an interactive effect (third media data). The first media data can include data of any media type. Optionally, the first media data includes data of at least one media type of image, video, text, and audio. The first media data can be media data obtained by real-time shooting based on a camera device set on the first end, or it can also be media data obtained from a target storage space (such as a material library of the application software, or a local terminal album, etc.), or it can also be media data uploaded after taking a screenshot of the communication interface including the first end and the second end, or it can also be template media data pre-set in the application software.
[0045] Among them, the second end can be an interactive receiving end or an interactive participating end, that is, the device end to which the recipient of the interactive operation belongs. Optionally, the second end includes at least one of a mobile terminal, a PC end, a tablet computer and a virtual reality device. The second media data can be media data obtained at the second end for generating data with an interactive effect (third media data). The second media data can include data of any media type. Optionally, the second media data includes data of at least one media type of image, video, text and audio. The second media data can be media data obtained by real-time shooting based on the camera device set on the first end, or it can also be media data obtained from the target storage space (such as the material library of the application software, or the local terminal album, etc.), or it can also be media data uploaded after taking a screenshot of the communication interface including the first end and the second end, or it can also be template media data pre-set in the application software.
[0046] It should be noted that in a multi-terminal interaction scenario, the information interaction process provided by the embodiments of the present disclosure needs to be executed when the multi-terminals participating in the interaction authorize the information interaction operation, that is, the acquisition operation of the first media data of the first terminal and the second media data of the second terminal can be executed when the information interaction operation is authorized. Generally speaking, the authorization of the information interaction operation may include at least two situations, optionally including responding to a trigger operation for an interaction restriction setting item; or responding to a request feedback operation for the interaction request when the second terminal receives an interaction request, etc.
[0047] Optionally, in response to a first trigger operation in a multi-terminal interaction scenario, an interactive information setting page is displayed, wherein the interactive information setting page includes an interactive restriction setting item; in response to a second trigger operation for the interactive restriction setting item, a restriction attribute value corresponding to the interactive restriction setting item is obtained, and the restriction attribute value is displayed.
[0048] Among them, the first trigger operation can be understood as an operation that triggers the display of an interactive information setting page corresponding to a multi-terminal interaction scenario. The interactive information setting page can be understood as a page for providing interactive operations with interactive information items. The interactive restriction setting item can be understood as an operation entry for setting interactive restriction information. Optionally, the interactive restriction setting item may include a state switching control and / or a state selection control. The first trigger operation may be an edit trigger operation for the interactive restriction setting item. The first trigger operation may include a state switching operation and / or a state selection operation. The restriction attribute value is used to characterize the restriction of the interactive restriction setting item on the multi-terminal interaction scenario. Optionally, the restriction attribute value may characterize whether the restriction state corresponding to the interactive restriction setting item is enabled. For example, assuming that the restriction state corresponding to the interactive restriction setting item is enabled, it means that the current end does not accept interactive requests with any end; assuming that the restriction state corresponding to the interactive restriction setting item is disabled, it means that the current end can accept interactive requests with any end.
[0049] In a specific implementation, in response to a first trigger operation input in a multi-terminal interaction scenario, an interactive information setting page is displayed, and the interactive information setting page includes an interactive restriction setting item. By inputting a state selection operation for the displayed interactive restriction setting item, the interactive restriction information associated with the multi-terminal interaction scenario can be set. Furthermore, in the case of receiving a state display operation for the interactive restriction setting item, at least two candidate states are displayed based on the interactive restriction setting item, and the at least two candidate states include an on state and an off state. Furthermore, in the case of receiving a state selection operation for the off state, the restriction state corresponding to the interactive restriction setting item, which is off, can be used as the restriction attribute value, and the restriction attribute value is displayed in the corresponding restriction information display area.
[0050] Optionally, upon receiving an interaction request initiated in a multi-end interaction scenario, a request response page is displayed, including options for approving the request and rejecting the request. Furthermore, upon receiving a triggering action for selecting the option to approve the request, first media data from the first end initiating the interaction request and second media data from the second end participating in the interaction are retrieved.
[0051] It should be noted that, for the second end, if the second end authorizes the interactive operation, the second end's second media data can also be obtained by the first end that initiates the interactive request. Since the initiation of the interactive request is usually during the communication between the first end and the second end, the display interface of the first end will include the screen content displayed by the second end. Furthermore, if the second end authorizes the interactive operation, the first end can take a screenshot of the screen content displayed on its display interface, and the screenshot data containing the screen content displayed by the second end is used as the second media data of the second end.
[0052] For example, assuming the multi-end interaction scenario is a video call scenario, the first and second ends are engaged in video communication, the first media data is first image data, and the second media data is second image data. During the video communication between the first and second ends, upon receiving a trigger operation from the first end for an interactive control input in the video communication interface, an interaction request can be initiated to the second end. Furthermore, upon receiving the interaction request, the second end can display a request response page based on the second end's display interface, the request response page including an option to approve the request and an option to reject the request. Furthermore, upon receiving a trigger operation from the second end for selecting the option to approve the request, the first and second ends can begin the interaction operation. Furthermore, upon receiving an image upload operation from the first end, an image selection page can be displayed, the image selection page including at least one candidate image. Subsequently, upon receiving a selection operation for any candidate image, the selected candidate image can be used as the first image data of the first end. For the second end, when receiving the image screenshot operation input by the first end, the real-time communication stream containing the second end can be pulled, and the pulled real-time communication stream can be framed, and the image frame obtained by the frame capture can be used as the second image data of the second end.
[0053] S120. Obtain first prompt information corresponding to the interaction request; wherein the first prompt information includes description data of an interaction method between at least a portion of the first media data and at least a portion of the second media data.
[0054] In which, the first prompt information can be used to prompt the expected interaction mode contained in the generated third media data, and the expected interaction mode is used to characterize the interaction between at least part of the data in the first media data and at least part of the data in the second media data. The first prompt information is used to guide the media data generation model on how to process the first media data and the second media data to generate the third media data containing the expected interaction mode. At least part of the data of the first media data can be data containing interactive objects in the first media data. At least part of the data includes all media data of the first media data or part of the media data of the first media data containing interactive objects. At least part of the data of the second media data includes data containing interactive objects in the second media data. At least part of the data includes all media data of the second media data or part of the media data of the second media data containing interactive objects. In which, the interactive object can be a main object and / or at least part of the body part of the main object.
[0055] The interaction method can be understood as the interactive operation performed on at least a portion of the first media data and at least a portion of the second media data in the ultimately generated third media data. For example, assuming that at least a portion of the first media data is media data containing a first interactive object, and at least a portion of the second media data is media data containing a second interactive object, the interaction method for the first interactive object and the second interactive object can include at least one of "handshake," "hug," "eye contact," and "take a photo together."
[0056] In the embodiment of the present disclosure, there are at least two ways to obtain the first prompt information, optionally including prompt information entry and / or prompt information selection. The following will describe these two ways of obtaining the first prompt information respectively.
[0057] It should be noted that the editing operation for the first prompt information can be performed based on the first end or the second end. In other words, the first prompt information can be determined based on the prompt information editing operation inputted on the first end or the prompt information editing operation inputted on the second end.
[0058] Optionally, obtaining the first prompt information corresponding to the interaction request includes: displaying a prompt information entry area, receiving an information entry operation through the prompt information entry area, and determining the first prompt information corresponding to the interaction request according to the information entry operation.
[0059] Among them, the prompt information entry area can be understood as a specific area presented in the display interface, in which relevant prompt information can be entered. The display form of the prompt information entry area in the display interface includes at least one of a pop-up window display, a pull-up or pull-down display panel, a sidebar display panel, and using at least part of the display interface as the prompt information entry area. The information entry operation can be understood as the operation of editing the prompt information in the prompt information entry area. Optionally, the information entry operation includes a keyword filling operation and / or an interactive demand input operation.
[0060] In a specific implementation, when an interaction request initiated in a multi-terminal interaction scenario is received, a prompt information entry area can be displayed based on the display interface of the first end. Furthermore, the first end can input text content containing interaction requirements into the prompt information entry area, and upload the text content input into the prompt information entry area when a trigger operation is received for the information upload control contained in the prompt information entry area. Furthermore, the uploaded text content can be parsed to parse out at least part of the first media data involved in the first prompt information, at least part of the second media data, and the interaction method from the text content, and fill the parsed information into a preset prompt information template to obtain the first prompt information corresponding to the interaction request.
[0061] For example, Figure 2 This is a schematic diagram of an interface including a prompt information entry area provided by an embodiment of the present disclosure. Figure 2 As shown, area 21 is the prompt information input area, and prompt information can be input in area 21. Further, when a trigger operation for the "confirm" control is received, the first prompt information can be generated according to the prompt information input in area 21.
[0062] Optionally, obtaining the first prompt information corresponding to the interaction request includes: displaying multiple optional prompt information, receiving an information selection operation for the optional prompt information, and determining the first prompt information corresponding to the interaction request from the multiple optional prompt information according to the information selection operation.
[0063] The optional prompt information may be selectable prompt information. In the embodiment of the present disclosure, the optional prompt information may be a complete prompt information, that is, prompt information containing at least part of the data in the first media data, at least part of the data in the second media data, and the interaction method; different optional prompt information may include different interaction methods and / or different media data at both ends of the interaction method. Exemplarily, the optional prompt information may be "object A hugs object B", "object A and object B take a face photo together", "object A gives a gift to object B", etc. Among them, "object A" corresponds to at least part of the data in the first media data, "object B" and "object B's face" correspond to at least part of the data in the second media data, and "hug", "take a photo together", and "give a gift" correspond to the interaction method. Alternatively, the optional prompt information may also be prompt information containing part of the prompt information content, that is, prompt information containing only the interaction method, prompt information containing only at least part of the data in the first media data, and prompt information containing only at least part of the data in the second media data.
[0064] As an optional implementation of an embodiment of the present disclosure, when the optional prompt information includes complete prompt information, multiple optional prompt information are displayed. When an information selection operation for the displayed optional prompt information is received, the selected optional prompt information can be used as the first prompt information corresponding to the interaction request.
[0065] As another optional implementation of the embodiment of the present disclosure, when the optional prompt information includes partial prompt information, multiple optional prompt information are displayed. When an information selection operation is received for the displayed optional prompt information, the selected optional prompt information can be spliced according to a preset prompt information template, and the spliced prompt information can be used as the first prompt information corresponding to the interaction request.
[0066] For example, Figure 3 This is a schematic diagram of a display interface including optional prompt information provided by an embodiment of the present disclosure. Figure 3 As shown, area 31 is a display area for optional prompt information, which includes 8 optional prompt information, namely prompt information A, prompt information B, prompt information C, prompt information D, prompt information E, prompt information F, prompt information G, and prompt information H. When an information selection operation for prompt information F is received, prompt information F can be used as the first prompt information.
[0067] It should be noted that the benefits of using the above two methods to obtain the first prompt information are: enabling the interactive end participating in the interaction to customize the first prompt information, enhancing the flexibility of interactive operations in multi-end interactive scenarios, and thus enhancing the customized participation of the end-to-end in the third media data, thereby improving the interactive participation experience between multiple ends.
[0068] S120: Generate third media data according to the first prompt information, the first media data, the second media data, and the media data generation model, and display the third media data on the first end and / or the second end.
[0069] The media data generation model may be a generative model that processes the input first and second media data based on the input first prompt information to generate third media data. Based on the first media data, the second media data, and the first prompt information, the media data generation model may generate third media data representing an interaction between at least a portion of the first media data and at least a portion of the second media data according to the interaction method described in the first prompt information. The third media data may be media data representing an interaction between at least a portion of the first media data and at least a portion of the second media data according to the interaction method described in the first prompt information. The third media data may include any type of media data capable of representing the interaction between the first and second media data. Optionally, the third media data includes static image data and / or sequence frame data. The sequence frame data includes at least one of video data, dynamic image data, and animation data. For example, assume that the first media data includes first image data, at least a portion of which includes object A; the second media data includes second image data, at least a portion of which includes object B; and the interaction method is a hug. Furthermore, by inputting the first prompt information, the first image data, and the second image data into a media data generation model, video data of the object A and the object B hugging each other can be generated.
[0070] In the embodiment of the present disclosure, there may be at least two ways to generate the third media data, and these two ways will be described below respectively.
[0071] Optionally, generating third media data based on the first prompt information, the first media data, the second media data and the media data generation model includes: inputting the first prompt information, the first media data and the second media data into the media data generation model, processing the first media data and the second media data according to the first prompt information based on the media data generation model, and outputting the third media data.
[0072] In a specific implementation, after obtaining first prompt information, first media data from the first end, and second media data from the second end, the first prompt information, the first media data, and the second media data can be input into a media data generation model. Furthermore, based on the first prompt information, the media data generation model extracts at least a portion of the first media data and at least a portion of the second media data, and generates third media data in which at least a portion of the first media data and at least a portion of the second media data interact according to the interaction method described in the first prompt information.
[0073] Optionally, generating third media data based on the first prompt information, the first media data, the second media data and the media data generation model includes: generating fourth media data based on the first media data and the second media data, and inputting the first prompt information and the fourth media data into the media data generation model to obtain the third media data.
[0074] The fourth media data may be composite media data comprising the data content of the first media data and the data content of the second media data. It should be noted that the fourth media data may be media data comprising at least a portion of the first media data and at least a portion of the second media data, without stylizing the data. Alternatively, the fourth media data may be media data comprising at least a portion of the first media data and at least a portion of the second media data, with stylization being performed on at least a portion of the first media data and / or at least a portion of the second media data.
[0075] In an embodiment of the present disclosure, the fourth media data may be generated in a variety of ways, optionally including: splicing at least part of the first media data and at least part of the second media data to obtain the fourth media data; or fusing at least part of the first media data and at least part of the second media data to obtain the fourth media data; or generating the fourth media data based on the second prompt information, the first media data, the second media data and an image generation model, etc.
[0076] As an optional implementation of an embodiment of the present disclosure, upon obtaining the first prompt information, the first media data, and the second media data, at least a portion of the first media data and at least a portion of the second media data may be concatenated to obtain fourth media data. Furthermore, the first prompt information and the fourth media data may be input into a media data generation model. Furthermore, based on the media data generation model, the fourth media data is processed according to the first prompt information to obtain third media data in which at least a portion of the first media data and at least a portion of the second media data interact according to the interaction method described in the first prompt information.
[0077] In the embodiment of the present disclosure, after the third media data is obtained, the third media data may be displayed, or the third media data may be stored in a target storage space so that the third media data can be viewed and / or shared.
[0078] In an embodiment of the present disclosure, when displaying the third media data, the third media data may be displayed only on the display interface of the first end, or only on the display interface of the second end, or on the display interfaces of both the first end and the second end.
[0079] Optionally, the multi-terminal interaction scenario includes a live broadcast scenario; presenting the third media data on the first terminal and / or the second terminal includes rendering the third media data in a preset manner within the live broadcast room where the first terminal and the second terminal are located. This arrangement has the advantage of presenting the third media data, providing a richer and more diverse interactive experience for the multiple terminals participating in the interaction, and enriching the visual content and live broadcast effect of the live broadcast room.
[0080] In an embodiment of the present disclosure, in a live broadcast scenario, the first end can be used as the audience end, and the second end can be used as the host end. The preset method can be a pre-set method for displaying third media data generated based on an interactive request. Optionally, the preset method includes rendering the third media data according to a preset rendering format at a preset location in the live broadcast room, and / or achieving a natural transition from the live broadcast room display to the third media data according to a preset special effect transition method.
[0081] In a specific implementation, when the multi-terminal interaction scenario is a live broadcast scenario, after generating the third media data, the third media data can be rendered to a preset position in the live broadcast room where the first end and the second end are located according to a preset rendering format, so that both the first end and the second end can see the display effect of the third media data in the live broadcast room.
[0082] The technical solution of the embodiment of the present disclosure supports obtaining media data of both parties of the interaction by responding to an interaction request initiated in a multi-terminal interaction scenario, obtaining first media data of the first end initiating the interaction request and second media data of the second end participating in the interaction, thereby providing a data basis for the subsequent generation of third media data, and achieving the effect of responding to interaction requests in a multi-terminal interaction scenario; further, obtaining first prompt information corresponding to the interaction request provides a generation basis for the subsequent generation of third media data, and the first prompt information corresponding to the interaction request reflects the degree of association between the first prompt information and the first end initiating the interaction request, achieving the effect of supporting customized editing of the first prompt information, enhancing the interactive participation of multiple terminals, and because the first prompt information includes a description of the interaction method of at least part of the first media data and at least part of the second media data. The above data enables the user to quickly know how to process the first media data and the second media data when subsequently generating the third media data, thereby improving the generation efficiency of the third media data, and the clear interaction method makes the interaction method in the generated third media data more standardized; further, the third media data is generated according to the first prompt information, the first media data, the second media data and the media data generation model, and the third media data is displayed on the first end and / or the second end, thereby realizing the effect of processing the first media data and the second media data according to the first prompt information through the media data generation model to generate the third media data, and displaying the third media data, thereby improving the correlation between the third media data and the first end and the second end, enriching the presentation effect of the third media data, and displaying the third media data, providing a richer and more diverse interactive experience for multiple ends participating in the interaction. The technical solution provided by the embodiments of the present disclosure solves the problems of relatively single interaction methods and monotonous interaction effects in related technologies. It realizes that in a multi-terminal interaction scenario, media data from multiple terminals are processed by a generative model based on interaction prompt information to generate third media data that can present the effect of media data from multiple terminals interacting according to a preset interaction method. This enhances the diversity of the third media data, enhances the correlation between the third media data and the multiple terminals participating in the interaction and the multi-terminal interaction scenario, improves the authenticity of the third media data, enriches the interactive participation forms in the multi-terminal interaction scenario and the display effects of the third media data, and thus improves the interactive experience between terminals.
[0083] Figure 4This is a flow chart of another information interaction method provided by the embodiment of the present disclosure. The technical solution of this embodiment is based on the above embodiment. When the third media data includes video data, the first media data includes first image data, and the second media data includes second image data, the generation process of video data is further refined. For specific implementation methods, please refer to the description of this embodiment. Among them, the technical features that are the same or similar to the above embodiments are not repeated here. Figure 4 As shown, the method of this embodiment may specifically include:
[0084] S210. In response to an interaction request initiated in a multi-terminal interaction scenario, obtain first media data of the first end initiating the interaction request and second media data of the second end participating in the interaction; wherein the first media data includes first image data, and the second media data includes second image data.
[0085] S220. Obtain first prompt information corresponding to the interaction request; wherein the first prompt information includes description data of an interaction method between at least a portion of the first image data and at least a portion of the second image data.
[0086] S220. Generate third image data based on the first image data and the second image data, generate video data based on the first prompt information, the third image data and the media data generation model, and display the video data on the first end and / or the second end; wherein the third image data includes at least part of the data in the first image data and at least part of the data in the second image data.
[0087] The first image data may be image data corresponding to the first end to be interacted with the second end. Optionally, the first image data includes image data captured at the first end and / or image data uploaded. The second image data may be image data corresponding to the second end to be interacted with the first end. Optionally, the second image data includes at least one of image data captured at the second end, image data uploaded at the second end, and screenshot data of the image data of the second end displayed on the first end. The third image data may be composite image data including at least a portion of the first image data and at least a portion of the second image data.
[0088] It should be noted that both the first image data and the second image data include multiple acquisition methods, which supports multiple image data uploading methods in multi-terminal interaction scenarios, enhances the flexibility of interactive operations, and enriches the video content of the video data, enhancing the diversity of the video data.
[0089] Exemplarily, in the case where the multi-end interaction scenario is a live broadcast scenario, the first end is the audience end, and the first image data may include image data captured by a camera device provided on the audience end, and / or image data selected and uploaded from the target storage space in response to an image upload operation of a live broadcast viewing account. The second end may be the host end, and the second image data may include image data captured by a camera device provided on the host end, and / or image data selected and uploaded from the target storage space in response to an image upload operation of the host (for example, life pictures / portraits uploaded by the host for users to interact with live broadcast viewing), and / or live broadcast room screen content captured and displayed on the audience end in response to a screenshot operation of a live broadcast viewing account, wherein the live broadcast room screen content includes the host end. The third image data may be image data including the audience end and the host end generated based on the first image data corresponding to the audience end and the second image data corresponding to the host end.
[0090] In the embodiments of the present disclosure, the third image data may be generated in a variety of ways, including: splicing at least a portion of the first image data with at least a portion of the second image data to obtain the third image data; or fusing at least a portion of the first image data with at least a portion of the second image data to obtain the third image data; or generating the third image data based on the second prompt information, the first image data, the second image data, and an image generation model, etc. These generation methods are described below.
[0091] The first method is to generate third image data according to the first image data and the second image data, including: splicing the first image data and the second image data, and using the spliced image data as the third image data.
[0092] The second method is to generate third image data based on first and second image data, including obtaining second prompt information and generating third image data based on the second prompt information, the first and second image data, and the image generation model. The advantage of this method is that the image generation model processes the first and second image data based on the second prompt information to generate the third image data that meets the image generation requirements, thereby meeting the customization requirements of the generated image and enhancing the content richness of the generated image. In addition, the use of the second prompt information provides clear guidance for the image generation process, making the generated third image data more in line with expectations, and achieving precise adjustment of the image content, image style, and other aspects of the third image data through the second prompt information.
[0093] The second prompt information can be used to indicate the data processing method to be performed by the image generation model and / or the expected output image content. The second prompt information includes descriptive data for the data processing method of the first image data and the second image data and / or the expected display information of the generated image. It should be noted that the second prompt information can be prompt information pre-set during the application software development phase, which clearly specifies how to perform data processing on the first image data and the second image data and / or the expected display information of the generated image. And / or, the second prompt information can also be prompt information customized and edited according to image generation requirements. The editing method of the second prompt information may include: prompt information entry and / or prompt information selection, etc. The data processing method for the first image data and the second image data may include: object segmentation processing, stylization processing, special effects addition processing, and object deformation processing. The generated image can be the image output by the image generation model. The expected display information can be the expected image content of the image output by the image generation model.
[0094] As an optional implementation of an embodiment of the present disclosure, multiple optional prompt information is displayed, an information selection operation is received for the optional prompt information, and second prompt information is determined from the multiple optional prompt information based on the information selection operation. Furthermore, the second prompt information, the first image data, and the second image data can be input into an image generation model, so that the image generation model processes the first image data and the second image data based on the second prompt information to generate third image data containing the intended display information.
[0095] The third method generates third image data based on first and second image data, including: performing image segmentation processing on the first image data to obtain object image data including first object data, and generating the third image data based on the object image data and the second image data. This arrangement has the advantages of: performing object segmentation on the first image data avoids interference from background data during the image generation process, thereby improving the generation of the third image data and the video data; and, by performing image segmentation processing before image generation, reducing network computational effort and the load on the third media data, thereby improving the efficiency of generating the third media data.
[0096] The first object data may be a specific object extracted from the first image data, and the specific object may be an object that interacts in subsequently generated video data. The first object data may be object data filtered out from the foreground object data contained in the first image data. The number of first object data may be one or more. The first object data may include any type of object contained in the first image data. Optionally, the first object data includes at least one of a person, at least part of a body part associated with a person, a pet, a plant, and a building. The object image data may be image data obtained by segmentation that includes the first object data. Exemplarily, the object image data may be an image of a person cut out from the first image data.
[0097] In an embodiment of the present disclosure, there may be multiple ways to determine the first object data, optionally including: in response to an object selection operation input for the first image data, determining the first object data based on the object selection operation; and / or, determining object data in the first image data that matches a preset object type as the first object data, etc.
[0098] In an embodiment of the present disclosure, there may be multiple ways of generating the third image data based on the object image data and the second image data, optionally including: splicing the object image data and the second image data, and using the spliced image data as the third image data; and / or splicing the object image data and the second image data to obtain spliced image data, generating fused image data based on the spliced image data, and generating the third image data based on the fused image data.
[0099] As an optional implementation of an embodiment of the present disclosure, when an object selection operation is received for first image data, the selected object can be used as first object data, and image segmentation processing can be performed on the first image data based on the first object data to obtain object image data containing the first object data. Furthermore, the object image data can be spliced with a second image, and the resulting spliced image data can be used as third image data.
[0100] As another optional implementation of the embodiment of the present disclosure, at least one object type can be pre-set, and then, when the first image data is obtained, the object type contained in the first image data is obtained, and type matching is performed with the object type contained in the first image data based on the preset at least one object type. Further, when an object type that matches the preset object type is determined, the object data under the object type in the first image data can be used as the first object data. Further, image segmentation processing can be performed on the first image data based on the first object data to obtain object image data containing the first object data. Further, the object image data and the second image data are spliced together to obtain spliced image data, and the spliced image data obtained at this time is spliced together by two image data with different backgrounds. Furthermore, the stitched image data can be fused according to a preset image fusion method to obtain fused image data. Since the obtained fused image data is formed by fusing two images with different backgrounds into image data with the same background, there may be a mismatch in the ratio between the foreground object data and the background data in the fused image data. Furthermore, since one of the object data in the fused image data is segmented from the first image data, after image fusion, the edge clarity of the object in the object data may be low. Furthermore, the second object data in the fused image data can be replaced with the first object data in the first image data to obtain third image data; and / or the fused image data can be expanded to obtain the third image data. The preset image fusion method may include: processing the stitched image data according to an image fusion algorithm to obtain fused image data; or inputting the stitched image data into an image fusion model to process the stitched image data based on the image fusion model to obtain fused image data.
[0101] The fourth method involves generating third image data based on the first and second image data, including generating fused image data based on the first and second image data and an image fusion model, and generating third image data based on the fused image data. This arrangement provides the advantage of achieving image fusion of the first and second image data, enriching the image content of the third image data, and, by generating the fused image data using the image fusion model, improving the accuracy of fused image generation while increasing image generation efficiency, thereby providing a data foundation for subsequent generation of high-quality third media data.
[0102] The fused image data includes second object data corresponding to the first object data in the first image data. The fused image data may be image data obtained by fusing at least part of the first image data with at least part of the second image data. The fused image data may include background data and foreground object data; the background data may be background data in the first image data or the second image data; the foreground object data may include first object data in the first image data and third object data in the second image data; or the foreground object data may only include first object data in the first image data. In other words, the fused image data may be a fusion of the background data of the first image data and the second image data, so that the two background data are merged into one background data.
[0103] The image fusion model can be understood as a deep learning model that fuses images containing at least two types of image content and generates a fused image containing one type of image content. The image fusion model can be a neural network model with any model structure, optionally including a generative adversarial network and / or a diffusion model.
[0104] In an embodiment of the present disclosure, there are at least two ways to generate fused image data: directly inputting the first image data and the second image data into an image fusion model to obtain fused image data; or determining stitched image data corresponding to the first image data and the second image data, and inputting the stitched image data into an image fusion model to obtain fused image data.
[0105] Optionally, generating fused image data based on the first image data, the second image data, and the image fusion model includes: determining stitched image data based on the first image data and the second image data, and generating the fused image data based on the stitched image data and the image fusion model. This arrangement has the advantage of reducing the amount of data processed by the image fusion model, thereby improving the efficiency of generating the fused image data. Furthermore, by stitching the images before fusing them, the resulting fused image data retains the content details of the first and second image data to the greatest extent possible, thereby improving the quality of the fused image data.
[0106] The spliced image data may be a piece of image data obtained by splicing two pieces of image data, and the image data includes the image content of the first image data and the image content of the second image data.
[0107] In an embodiment of the present disclosure, determining the stitched image data based on the first image data and the second image data includes: stitching the first image data and the second image data to obtain the stitched image data; or performing object segmentation on the first image data to obtain object image data containing the first object data, and stitching the object image data and the second image data to obtain the stitched image data.
[0108] As an optional implementation of the first embodiment of the present disclosure, the first image data and the second image data may be spliced to obtain spliced image data. Further, the spliced image data is input into an image fusion model to perform image fusion on the spliced image data based on the image fusion model to obtain fused image data.
[0109] As another optional implementation of the disclosed embodiment, the first image data is segmented to obtain object image data containing the first object data, and the object image data and the second image data are spliced together to obtain spliced image data. Furthermore, the spliced image data is input into an image fusion model to perform image fusion on the spliced image data based on the image fusion model to obtain fused image data.
[0110] In the disclosed embodiments, since the resulting fused image data is formed by fusing two images with different backgrounds into image data with the same background, the ratio between the foreground object data and the background data in the fused image data may be inconsistent. Furthermore, since one of the object data in the fused image data is segmented from the first image data, the edge clarity of the object data may be low after image fusion. Furthermore, once the fused image data is obtained, methods for generating third image data from the fused image data may include object data replacement and / or background data expansion.
[0111] Optionally, generating the third image data based on the fused image data includes: replacing the second object data in the fused image data with the first object data in the first image data to obtain the third image data; and / or performing expansion processing on the background data in the fused image data to obtain the third image data. This arrangement has the advantage of improving the image quality of the third image data through object replacement and / or background expansion, and maximally preserving the object features of the first object data in the first image data through object replacement, thereby improving the correlation between the third media data and the first media data, and enhancing the quality and effectiveness of generating the third media data.
[0112] It should be noted that, when generating the third image data based on the fused image data, the fused image data may be processed using at least one of the above-mentioned methods.
[0113] The second object data is the object data corresponding to the first object data contained in the fused image data, that is, the object data obtained by image fusion of the first object data segmented from the first image data.
[0114] Among them, the expansion processing of the background data can be understood as expanding the boundary of the fused image data outward and filling the expanded boundary area based on the image content of the background data. In other words, it can be to generate picture content outside the fused image data, and the picture content can be generated based on the image content of the background data to maintain consistency with the fused image data in style, semantics and logic. The expansion processing of the background data can include a variety of processing methods, optionally including processing the fused image data based on a diffusion model; and / or, processing the fused image data based on a generative adversarial network; and / or, obtaining third prompt information, and generating third image data based on the image generation model, the fused image data and the third prompt information.
[0115] As an optional implementation of an embodiment of the present disclosure, when the fused image data is obtained, the first object data in the first image data can be obtained, and the second object data in the fused image data can be replaced with the first object data, and the replaced fused image data can be used as the third image data.
[0116] As another optional implementation of the embodiment of the present disclosure, when fused image data is obtained, the fused image data can be input into a diffusion model to perform expansion processing on background data in the fused image data based on the diffusion model and output third image data.
[0117] As another optional implementation of the embodiment of the present disclosure, when the fused image data is obtained, the first object data in the first image data is obtained, and the second object data in the fused image data is replaced with the first object data, and the replaced fused image data is used as the fourth image data to be expanded; further, the fourth image data can be input into the diffusion model to perform expansion processing on the background data in the fourth image data based on the diffusion model, and output the third image data.
[0118] In an embodiment of the present disclosure, after obtaining the third image data, the first prompt information and the third image data may be input into a media data generation model to generate video data based on the first prompt information and the third image data through the media data generation model.
[0119] The technical solution of the embodiment of the present disclosure is to obtain, in response to an interaction request initiated in a multi-terminal interaction scenario, first media data of a first terminal initiating the interaction request and second media data of a second terminal participating in the interaction; wherein the first media data includes first image data and the second media data includes second image data; further, obtain first prompt information corresponding to the interaction request; wherein the first prompt information includes description data of an interaction method for at least a portion of the first image data and at least a portion of the second image data; further, generate third image data based on the first image data and the second image data, and generate video data based on the first prompt information, the third image data, and a media data generation model, so as to realize, in a multi-terminal interaction scenario, processing the image data of multiple terminals according to the interaction prompt information through a generative model to generate video data that can present the effect of the image data of multiple terminals interacting according to a preset interaction method, thereby achieving the goal of ensuring the generation quality of the video data while improving the efficiency of video data generation, thereby improving the authenticity of the video data, enriching the interactive participation forms and the display effect of the video data in the multi-terminal interaction scenario, and thereby improving the end-to-end interaction experience.
[0120] Figure 5 This is a schematic diagram of an information interaction process provided by an embodiment of the present disclosure. This embodiment of the present disclosure is a preferred embodiment of the above disclosed embodiments. Assuming that the multi-end interaction scenario is a live broadcast scenario, the first end is the audience end, the first media data is the first image data, the second end is the anchor end, the second media data is the second image data, and the third media data is the video data. Next, you can combine Figure 5 The implementation process of the information interaction method provided in the embodiment of the present disclosure is described below:
[0121] First, while the audience is watching the live broadcast, the audience can initiate an interactive request to the host in the live broadcast room. Furthermore, when the host receives the interactive request and responds that it agrees to the interaction, the audience can obtain the first image data in real time based on the camera device set on the audience, or select the locally stored image data to upload as the first image data; for the host, the host can select the locally stored life pictures or portraits that can be used for audience interaction as the second image data. Furthermore, the background of the first image data is segmented, and the redundant background is removed to obtain the object image data containing the first object data. Afterwards, the object image data and the second image data are spliced to obtain spliced image data. Furthermore, the deep learning model is called to process the spliced image data to respond to the interactive request and generate video data.
[0122] The video data generation process may include: obtaining prompt information, wherein the prompt information includes interactive prompt information for generating video data and stylized prompt information for performing image processing on the spliced image data. Furthermore, the stylized prompt information and the spliced image data may be input into an image fusion model to obtain fused image data. Furthermore, the fused image data and the first image data may be input into a face replacement model to replace the second object data corresponding to the first object data in the fused image data with the first object data in the first image data based on the face replacement model, and to ensure that the style type of the replaced first object data is consistent with the style type of the second object data. Furthermore, the background data in the replaced fused image data may be expanded based on an image expansion model, and the fused image data after the background expansion is used as the third image data. Furthermore, the third image data and the interactive prompt information may be input into a video generation model to process the third image data based on the interactive prompt information based on the video generation model to generate video data.
[0123] The technical solution provided by the disclosed embodiments supports both communication stream (or camera stream) frame capture and local image upload for acquiring the first and second image data. Real-time communication stream frame capture provides greater randomness, resulting in unexpected results for the final generated video data, while local image upload provides a backup for the generated results.
[0124] When obtaining the first image data and the second image data, directly stitching the two images together often fails to achieve the desired effect, and the input needs to be preprocessed. In order to avoid mutual interference between the backgrounds of the two images or background fragmentation during the fusion process, the locally collected first image data needs to be segmented for background stitching, and then the model's image expansion capability is used to naturally fill in the background. The intelligent editing and fusion capabilities of the image fusion model are used to fuse two independent images into one scene. Different actions and backgrounds can be customized based on the usage scenario and image generation requirements. The facial features of the first object data in the first image data are not retained, and the fused image uses facial feature replacement capabilities to replace objects to avoid low correlation between the fused image data generated by the link and the input first image data.
[0125] Image fusion and video data generation support customized editing of prompt information to adapt to different scenarios.
[0126] After generating video data, it supports saving the video data locally for viewing and sharing, and also supports interactive preview of real-time communication streams. For better visual effects, special effects transition capabilities can be used to complete a natural transition from the real-time link to the generated video. Other real-time special effects can also be superimposed on the video data and pushed to the remote end, greatly enriching the effect and display form of the final generated video data.
[0127] Regarding interactive link design, in live streaming scenarios, by supporting local image upload or real-time camera stream acquisition methods, the viewer's first image data is subjected to background culling and then used as video data input for remote interaction with the remote host. This interactive method meets the basic needs of close contact between viewers and the host at a relatively low cost, and can also be promoted through the host's communication broadcast, making it highly appealing. In video call scenarios, due to distance and other objective factors, the callers are often in different locations. In addition to supporting image or countdown frame gameplay, more open prompt information options can be used to generate video data that brings the callers closer.
[0128] In order to achieve better video generation effects, the input image can be pre-beautified to brighten the picture, background cutout can be used to avoid interference from complex backgrounds, and a face replacement model can be used to obtain a more consistent fusion effect.
[0129] The video generation process is relatively time-consuming. In order to reduce the response waiting time, the input first image data and the second image data can be downsampled, and the resolution of image fusion and video generation can be reduced to reduce the network computing amount and reduce the service load. On this basis, further, super-resolution and image quality restoration algorithms can be used to process the generated video data to optimize the display effect of the video data. The parallelism of interactive requests can be increased, and the user's waiting time expectation can be provided by returning multiple images or video data for the user to choose from. In addition, in the live broadcast scenario, if the expected waiting time of the interactive participant is relatively long (for example, exceeding the preset time threshold), the satisfaction of the interactive terminal can be increased by requesting multiple additional requests (generating an additional video data). In the real-time call scenario, a waiting bar and interactive mini-games can be added during the waiting stage to reduce the perception of waiting time and enhance the information interaction experience.
[0130] Figure 6 This is a structural diagram of an information interaction device provided by an embodiment of the present disclosure, such as Figure 6 As shown, the device includes: a media data acquisition module 310, a prompt information acquisition module 320 and a media data generation module 330.
[0131] Among them, the media data acquisition module 310 is used to respond to an interaction request initiated in a multi-terminal interaction scenario, obtain the first media data of the first end that initiates the interaction request and the second media data of the second end that participates in the interaction; the prompt information acquisition module 320 is used to obtain first prompt information corresponding to the interaction request; wherein, the first prompt information includes description data of the interaction method of at least part of the first media data and at least part of the second media data; the media data generation module 330 is used to generate third media data based on the first prompt information, the first media data, the second media data and the media data generation model, and display the third media data at the first end and / or the second end.
[0132] The technical solution of the embodiment of the present disclosure responds to the interaction request initiated in the multi-terminal interaction scenario through the media data acquisition module 310, obtains the first media data of the first end that initiates the interaction request and the second media data of the second end that participates in the interaction, supports the acquisition of media data of both parties of the interaction, and thus provides a data basis for the subsequent generation of third media data, and realizes the effect of responding to the interaction request in the multi-terminal interaction scenario; further, the first prompt information corresponding to the interaction request is obtained through the prompt information acquisition module 320, which provides a generation basis for the subsequent generation of the third media data, and the first prompt information corresponding to the interaction request reflects the degree of association between the first prompt information and the first end that initiates the interaction request, and realizes the effect of supporting users to customize the first prompt information, thereby enhancing the user's interactive participation, and because the first prompt information includes at least part of the data of the first media data and at least part of the data of the second media data The description data of the interaction method enables the user to quickly know how to process the first media data and the second media data when subsequently generating the third media data, thereby improving the generation efficiency of the third media data, and the clear interaction method makes the interaction method in the generated third media data more standardized; further, the media data generation module 330 generates the third media data according to the first prompt information, the first media data, the second media data and the media data generation model, and displays the third media data on the first end and / or the second end, thereby realizing the effect of processing the first media data and the second media data according to the first prompt information through the media data generation model to generate the third media data, and displaying the third media data, thereby improving the degree of correlation between the third media data and the first end and the second end, enriching the presentation effect of the third media data, and displaying the third media data, providing a richer and more diverse interactive experience for multiple ends participating in the interaction. The technical solution provided by the embodiments of the present disclosure solves the problems of relatively single interaction methods and monotonous interaction effects in related technologies. It realizes that in a multi-terminal interaction scenario, media data from multiple terminals are processed by a generative model based on interaction prompt information to generate third media data that can present the effect of media data from multiple terminals interacting according to a preset interaction method. This enhances the diversity of the third media data, enhances the correlation between the third media data and the multiple terminals participating in the interaction and the multi-terminal interaction scenario, improves the authenticity of the third media data, enriches the interactive participation forms in the multi-terminal interaction scenario and the display effect of the third media data, and thus improves the user's interactive experience.
[0133] Based on any optional technical solution in the embodiment of the present disclosure, the third media data includes video data; the first media data includes first image data; the second media data includes second image data;
[0134] The media data generation module 330 includes an interactive video generation submodule. The interactive video generation submodule is configured to generate third image data based on the first image data and the second image data, and to generate video data based on the first prompt information, the third image data, and a video generation model; the third image data includes at least a portion of the first media data and at least a portion of the second media data.
[0135] Based on any optional technical solution in the embodiments of the present disclosure, the interactive video generation submodule includes: a first image data generation unit. The first image data generation unit is configured to obtain second prompt information and generate third image data based on the second prompt information, the first image data, the second image data, and an image generation model; wherein the second prompt information includes descriptive data regarding a data processing method for the first media data and the second media data and / or expected display information of the generated image.
[0136] Based on any optional technical solution in the embodiments of the present disclosure, the interactive video generation submodule includes: a second image data generation unit. The second image data generation unit is configured to perform image segmentation processing on the first image data to obtain object image data including first object data, and to generate third image data based on the object image data and the second image data.
[0137] Based on any optional technical solution in the embodiments of the present disclosure, the interactive video generation submodule includes: a third image data generation unit. The third image data generation unit is configured to generate fused image data based on the first image data, the second image data, and an image fusion model, and to generate third image data based on the fused image data; wherein the fused image data includes second object data corresponding to the first object data in the first image data.
[0138] Based on any optional technical solution in the embodiments of the present disclosure, the third image data generating unit includes: an image data generating subunit. The image data generating subunit is configured to replace the second object data in the fused image data with the first object data in the first image data to obtain third image data; and / or perform expansion processing on the background data in the fused image data to obtain the third image data.
[0139] Based on any optional technical solution in the embodiments of the present disclosure, the third image data generation unit includes a fused image data generation subunit, wherein the fused image data generation subunit is configured to determine stitched image data based on the first image data and the second image data, and generate the fused image data based on the stitched image data and an image fusion model.
[0140] Based on any optional technical solution in the embodiments of the present disclosure, the first image data includes image data taken at the first end and / or image data uploaded; the second image data includes at least one of image data taken at the second end, image data uploaded at the second end, and screenshot data of the image data of the second end displayed at the first end.
[0141] Based on any optional technical solution in the embodiments of the present disclosure, the prompt information acquisition module includes: a prompt information entry unit and / or a prompt information selection unit. The prompt information entry unit is configured to display a prompt information entry area, receive an information entry operation through the prompt information entry area, and determine the first prompt information corresponding to the interaction request based on the information entry operation; and / or the prompt information selection unit is configured to display multiple optional prompt information, receive an information selection operation for the optional prompt information, and determine the first prompt information corresponding to the interaction request from the multiple optional prompt information based on the information selection operation.
[0142] Based on any optional technical solution in the embodiments of the present disclosure, the multi-terminal interaction scenario includes a live broadcast scenario; and the third media data generation module includes a media data display unit. The media data display unit is configured to render the third media data in a preset manner in the live broadcast room where the first and second terminals are located.
[0143] Based on any optional technical solution in the embodiments of the present disclosure, the multi-terminal interaction scenario includes at least a live broadcast scenario and / or a video call scenario.
[0144] The information interaction device provided in the embodiments of the present disclosure can execute the information interaction method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects for executing the information interaction method.
[0145] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.
[0146] Reference below Figure 7, which shows a schematic structural diagram of an electronic device (e.g., a terminal device or a server) 600 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0147] like Figure 7 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0148] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 7 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0149] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0150] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0151] The electronic device provided by the embodiment of the present disclosure and the information interaction method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in the embodiment of the present disclosure, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0152] An embodiment of the present disclosure provides a computer storage medium on which a computer program is stored. When the program is executed by a processor, the information interaction method provided by the above embodiment is implemented.
[0153] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0154] According to one or more embodiments of the present disclosure, [Example 1] provides an information interaction method, including: in response to an interaction request initiated in a multi-terminal interaction scenario, obtaining first media data of the first end that initiates the interaction request and second media data of the second end that participates in the interaction; obtaining first prompt information corresponding to the interaction request; wherein the first prompt information includes description data of the interaction method of at least part of the first media data and at least part of the second media data; generating third media data based on the first prompt information, the first media data, the second media data and a media data generation model, and displaying the third media data on the first end and / or the second end.
[0155] According to one or more embodiments of the present disclosure, [Example 2] provides the method of Example 1, further including: optionally, the third media data includes video data; the first media data includes first image data; the second media data includes second image data; the generating of the third media data based on the interactive prompt information, the first media data, the second media data and the video generation model includes: generating the third image data based on the first image data and the second image data, and generating video data based on the first prompt information, the third image data and the media data generation model; wherein the third image data includes at least part of the data in the first image data and at least part of the data in the second image data.
[0156] According to one or more embodiments of the present disclosure, [Example Three] provides the method of Example Two, further including: optionally, generating third image data based on the first image data and the second image data includes: obtaining second prompt information, and generating third image data based on the second prompt information, the first image data, the second image data and an image generation model; wherein the second prompt information includes descriptive data of the data processing method of the first media data and the second media data and / or the expected display information of the generated image.
[0157] According to one or more embodiments of the present disclosure, [Example 4] provides the method of Example 2, further including: optionally, generating third image data based on the first image data and the second image data includes: performing image segmentation processing on the first image data to obtain object image data including first object data, and generating third image data based on the object image data and the second image data.
[0158] According to one or more embodiments of the present disclosure, [Example 5] provides the method of Example 2, further including: optionally, generating third image data based on the first image data and the second image data includes: generating fused image data based on the first image data, the second image data and an image fusion model, and generating third image data based on the fused image data; wherein the fused image data includes second object data corresponding to the first object data in the first image data.
[0159] According to one or more embodiments of the present disclosure, [Example Six] provides the method of Example Five, further including: optionally, generating third image data based on the fused image data includes: replacing the second object data in the fused image data with the first object data in the first image data to obtain third image data; and / or, performing expansion processing on the background data in the fused image data to obtain third image data.
[0160] According to one or more embodiments of the present disclosure, [Example 7] provides the method of Example 5, further including: optionally, performing image fusion based on the first image data, the second image data and the image fusion model to obtain fused image data includes: determining spliced image data based on the first image data and the second image data, and generating fused image data based on the spliced image data and the image fusion model.
[0161] According to one or more embodiments of the present disclosure, [Example 8] provides the method of Example 2, further including: optionally, the first image data includes image data captured at the first end and / or image data uploaded; the second image data includes at least one item of image data captured at the second end, image data uploaded at the second end, and screenshot data of the image data of the second end displayed in the first end.
[0162] According to one or more embodiments of the present disclosure, [Example 9] provides the method of Example 1, further including: optionally, obtaining the first prompt information corresponding to the interaction request includes: displaying a prompt information entry area, receiving an information entry operation through the prompt information entry area, and determining the first prompt information corresponding to the interaction request based on the information entry operation; and / or displaying multiple optional prompt information, receiving an information selection operation for the optional prompt information, and determining the first prompt information corresponding to the interaction request from the multiple optional prompt information based on the information selection operation.
[0163] According to one or more embodiments of the present disclosure, [Example 10] provides the method of Example 1, further including: optionally, the multi-terminal interaction scenario includes a live broadcast scenario; the display of the third media data at the first terminal and / or the second terminal includes: rendering the third media data in a preset manner in the live broadcast room where the first terminal and the second terminal are located.
[0164] According to one or more embodiments of the present disclosure, [Example 11] provides the method of Example 1, further including: optionally, the multi-terminal interaction scenario includes at least a live broadcast scenario and / or a video call scenario.
[0165] According to one or more embodiments of the present disclosure, [Example 12] provides an information interaction device, including: a media data acquisition module, used to respond to an interaction request initiated in a multi-terminal interaction scenario, to obtain first media data of the first end that initiates the interaction request and second media data of the second end that participates in the interaction; a prompt information acquisition module, used to obtain first prompt information corresponding to the interaction request; wherein, the first prompt information includes description data of the interaction method of at least part of the first media data and at least part of the second media data; a media data generation module, used to generate third media data based on the first prompt information, the first media data, the second media data and the media data generation model, and display the third media data on the first end and / or the second end.
[0166] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0167] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0168] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: responds to an interaction request initiated in a multi-terminal interaction scenario, obtains the first media data of the first end that initiates the interaction request and the second media data of the second end that participates in the interaction; obtains first prompt information corresponding to the interaction request; wherein the first prompt information includes description data of the interaction method of at least part of the first media data and at least part of the second media data; generates third media data according to the first prompt information, the first media data, the second media data and the media data generation model, and displays the third media data on the first end and / or the second end.
[0169] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0170] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0171] The units described in the embodiments of the present disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, a media data acquisition module may also be described as a "module that acquires first media data from a first end initiating an interaction request and second media data from a second end participating in the interaction."
[0172] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0173] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0174] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0175] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0176] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An information interaction method, characterized in that: include: In response to an interaction request initiated in a multi-terminal interaction scenario, obtaining first media data of a first terminal initiating the interaction request and second media data of a second terminal participating in the interaction; Obtaining first prompt information corresponding to the interaction request; wherein the first prompt information includes description data of an interaction method for at least a portion of the first media data and at least a portion of the second media data; Third media data is generated according to the first prompt information, the first media data, the second media data, and a media data generation model, and the third media data is displayed on the first end and / or the second end.
2. The information interaction method according to claim 1, characterized in that: The third media data includes video data; the first media data includes first image data; and the second media data includes second image data; The generating the third media data according to the first prompt information, the first media data, the second media data, and the media data generation model includes: Generate third image data based on the first image data and the second image data, and generate video data based on the first prompt information, the third image data and the media data generation model; wherein the third image data includes at least part of the data in the first image data and at least part of the data in the second image data.
3. The information interaction method according to claim 2, characterized in that: Generating third image data according to the first image data and the second image data includes: Obtain second prompt information, and generate third image data based on the second prompt information, the first image data, the second image data, and an image generation model; wherein the second prompt information includes descriptive data of the data processing method of the first media data and the second media data and / or the expected display information of the generated image.
4. The information interaction method according to claim 2, characterized in that: Generating third image data according to the first image data and the second image data includes: Image segmentation processing is performed on the first image data to obtain object image data including first object data, and third image data is generated based on the object image data and the second image data.
5. The information interaction method according to claim 2, characterized in that: Generating third image data according to the first image data and the second image data includes: Fused image data is generated according to the first image data, the second image data and an image fusion model, and third image data is generated according to the fused image data; wherein the fused image data includes second object data corresponding to the first object data in the first image data.
6. The information interaction method according to claim 5, characterized in that: Generating third image data according to the fused image data includes: replacing the second object data in the fused image data with the first object data in the first image data to obtain third image data; and / or, The background data in the fused image data is expanded to obtain third image data.
7. The information interaction method according to claim 5, characterized in that: Generating fused image data according to the first image data, the second image data, and an image fusion model includes: Stitched image data is determined according to the first image data and the second image data, and fused image data is generated according to the stitched image data and an image fusion model.
8. The information interaction method according to claim 2, characterized in that: The first image data includes image data taken at the first end and / or image data uploaded; the second image data includes at least one of image data taken at the second end, image data uploaded at the second end, and screenshot data of the image data of the second end displayed at the first end.
9. The information interaction method according to claim 1, characterized in that: The obtaining of first prompt information corresponding to the interaction request includes: displaying a prompt information entry area, receiving an information entry operation through the prompt information entry area, and determining first prompt information corresponding to the interaction request according to the information entry operation; and / or, Displaying a plurality of optional prompt information, receiving an information selection operation for the optional prompt information, and determining a first prompt information corresponding to the interaction request from the plurality of optional prompt information according to the information selection operation.
10. The information interaction method according to claim 1, characterized in that: The multi-terminal interaction scenario includes a live broadcast scenario; and presenting the third media data on the first terminal and / or the second terminal includes: The third media data is rendered in a preset manner in the live broadcast room where the first end and the second end are located.
11. The information interaction method according to claim 1, characterized in that: The multi-terminal interaction scenario includes at least a live broadcast scenario and / or a video call scenario.
12. An information interaction device, characterized in that: include: A media data acquisition module, configured to respond to an interaction request initiated in a multi-terminal interaction scenario and acquire first media data of a first terminal initiating the interaction request and second media data of a second terminal participating in the interaction; a prompt information acquisition module, configured to acquire first prompt information corresponding to the interaction request; wherein the first prompt information includes description data of an interaction method for at least a portion of the first media data and at least a portion of the second media data; The media data generation module is configured to generate third media data based on the first prompt information, the first media data, the second media data, and a media data generation model, and display the third media data on the first end and / or the second end.
13. An electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the information interaction method according to any one of claims 1 to 10.
14. A storage medium containing computer-executable instructions, characterized in that: When the computer executable instructions are executed by a computer processor, they are used to execute the information interaction method according to any one of claims 1 to 10.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the information interaction method according to any one of claims 1 to 10.