Interactive media generation method and device, equipment, storage medium and program product
By displaying media generation controls and recommended media within the interface, users can quickly obtain generation information and generate output media, solving the problems of low efficiency and poor effect in existing technologies and achieving efficient media generation.
Patent Information
- Application Number
- CN202510727426.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-12
AI Technical Summary
In the prior art, media generation efficiency is low and generation effect is poor, mainly because users need to input accurate prompt words and perform repeated debugging to achieve the expected effect.
An interactive media generation method is provided. By simultaneously displaying media generation controls and recommended media in an interface, users can select recommended media to obtain generation information, and generate output media based on the currently configured functional mode through the media generation controls.
It achieves a rapid improvement in the efficiency of media generation, while ensuring the quality and effect of output media and improving the efficiency of human-computer interaction.
Smart Images

Figure CN120640096A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of artificial intelligence technology, and in particular to an interactive media generation method, apparatus, device, storage medium, and program product. Background Art
[0002] Currently, with the development of artificial intelligence technology, the production of media data such as pictures and videos based on artificial intelligence generated content (AIGC) technology has become a new way of media creation. Users can describe their needs by entering prompt words, and then the image processing model will generate corresponding media data based on the prompt words, realizing fast and efficient media generation.
[0003] In the prior art, the generation effect of media data is mainly controlled by the prompt words input by the user. Therefore, the user must enter accurate and effective prompt words, or perform repeated debugging to achieve the expected generation effect. As a result, the solutions in the prior art have the problems of low media generation efficiency and poor generation effect. Summary of the Invention
[0004] The embodiments of the present disclosure provide an interactive media generation method, apparatus, device, storage medium, and program product to overcome the problems of low media generation efficiency and poor generation effect.
[0005] In a first aspect, an embodiment of the present disclosure provides an interactive media generation method, comprising:
[0006] A first interactive interface is displayed, wherein the first interactive interface is configured to simultaneously display a media generation control and at least two recommended media, wherein the media generation control is used to configure a functional mode and receive generation information, the functional mode is used to characterize a functional type of a media generation function, and the generation information is used to determine the content of the output media; in response to a first trigger operation for a target recommended media, generation information corresponding to the target recommended media is obtained, and the generation information is input into the media generation control; in response to a second trigger operation for the media generation control, the output media is generated based on the target functional mode currently configured for the media generation control and the generation information.
[0007] In a second aspect, an embodiment of the present disclosure provides an interactive media generating device, including:
[0008] a display module configured to display a first interactive interface configured to simultaneously display a media generation control and at least two recommended media, wherein the media generation control is configured to configure a functional mode and receive generation information, the functional mode being used to characterize a functional type of a media generation function, and the generation information being used to determine content of output media;
[0009] A first interaction module is configured to obtain generation information corresponding to a target recommended media in response to a first triggering operation on the target recommended media, and input the generation information into the media generation control;
[0010] The second interaction module is configured to generate the output media in response to a second trigger operation on the media generation control based on the target functional mode currently configured by the media generation control and the generation information.
[0011] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;
[0012] The memory stores computer-executable instructions;
[0013] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the interactive media generation method as described in the first aspect and various possible designs of the first aspect.
[0014] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the interactive media generation method described in the first aspect and various possible designs of the first aspect is implemented.
[0015] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the interactive media generation method described in the first aspect and various possible designs of the first aspect.
[0016] The interactive media generation method, apparatus, device, storage medium, and program product provided in this embodiment display a first interactive interface configured to simultaneously display a media generation control and at least two recommended media, wherein the media generation control is used to configure a functional mode and receive generation information, the functional mode being used to characterize the functional type of the media generation function, and the generation information being used to determine the content of the output media; in response to a first trigger operation on a target recommended media, generation information corresponding to the target recommended media is obtained and input into the media generation control; and in response to a second trigger operation on the media generation control, the output media is generated based on the target functional mode currently configured in the media generation control and the generation information. By simultaneously displaying the media generation control and recommended media in the first interactive interface, a user can obtain generation information corresponding to the recommended media by selecting the recommended media and input it into the media generation control. The corresponding output media is then generated based on the target functional mode currently configured in the media generation control, thereby achieving rapid media generation based on the recommended media, improving the efficiency of media generation, and ensuring the quality and effect of the generated output media. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0018] Figure 1 A schematic diagram of an application scenario for generating media in the prior art;
[0019] Figure 2 Schematic diagram of the process of the interactive media generation method provided by the embodiment of the present disclosure Figure 1 ;
[0020] Figure 3 A schematic diagram of a first interactive interface provided by an embodiment of the present disclosure;
[0021] Figure 4 A schematic diagram of another first interactive interface provided by an embodiment of the present disclosure;
[0022] Figure 5 This is a flowchart of a specific implementation method of step S1001;
[0023] Figure 6 Schematic diagram of the interactive media generation method provided in the embodiment of the present disclosure Figure 2 ;
[0024] Figure 7 A schematic diagram of a flow generation interface provided by an embodiment of the present disclosure;
[0025] Figure 8 A schematic diagram of another generation flow interface provided by an embodiment of the present disclosure;
[0026] Figure 9 A structural block diagram of an interactive media generating device provided in an embodiment of the present disclosure;
[0027] Figure 10 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure;
[0028] Figure 11 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0030] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0031] The following explains the application scenarios of the embodiments of the present disclosure:
[0032] The interactive media generation method provided by the embodiment of the present disclosure can be applied to applications (APPs) with media generation functions, such as image generation applications, video editing applications, AI assistant applications, etc. More specifically, it can be applied to media creation scenarios based on AIGC technology, including image generation, video generation, music generation, etc. The execution subject of this embodiment can be a terminal device running the above-mentioned application with media generation function, or a server that deploys the server corresponding to the above-mentioned application, or other electronic devices that perform similar functions. When the execution subject is a terminal device, the terminal device executes the method provided by this embodiment by running the above-mentioned application; when the execution subject is a server, the server of the above-mentioned application with media generation function can be partially or completely run on the server, and the method provided by this embodiment is executed on the server side, while the terminal device runs the client of the application, or runs a browser, wherein the communication between the server and the terminal device is based on the client-server (CS) architecture, or the communication between the server and the terminal device is based on the browser-server (BS) architecture, so that the terminal device can obtain the execution result of the method provided by this embodiment and display it as needed.
[0033] Among them, in some embodiments, the terminal device or server can implement the interactive media generation method provided by the embodiment of the present disclosure by running various computer executable instructions or computer programs. For example, computer executable instructions can be program-level commands, machine instructions or software instructions. The computer program can be a native program or software module in the operating system; it can be a local application, that is, a program that needs to be installed in the operating system to run, or it can be a small program embedded in any APP, that is, a program that runs based on a browser environment. In summary, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form, and the specific implementation form can be configured as needed. Furthermore, in the process of implementing the interactive media generation method provided by the embodiment of the present disclosure, the terminal device can execute the method by running a computer executable instruction or computer program set locally, or it can execute the method by calling a computer executable instruction or computer program set in an external server. In some embodiments, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud storage, cloud communications, cloud databases, cloud computing, cloud functions, network services, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Among them, cloud services can be interactive processing services for terminal devices to call.
[0034] Figure 1 This is a schematic diagram of an application scenario of generating media in the prior art, referring to Figure 1 As shown in FIG, a terminal device, for example, runs a target application with media generation capabilities via a browser. The target application's interactive interface includes an interactive component and a media presentation component. The interactive component, for example, a dialog window, receives user input for a request, such as, for example, "Draw me a picture of a car." After the user enters the request into the interactive component, the terminal device interprets the request by invoking an image processing model based on AIGC technology. The model generates multiple images matching the request, such as images P1, P2, P3, and P4, and displays them within the media presentation component. The user can then select one of the images displayed within the media presentation component to generate the final output image, or further process the image to generate the final output image, completing the image generation process. In this application scenario, in addition to generating images, other types of media data can also be generated based on the above process. For example, the target application can also generate a video based on the user's input request.
[0035] In the existing technology, in the application scenario of media creation based on AIGC technology, the generation effect of media data is mainly controlled by the prompt words input by the user. Therefore, the user is required to enter accurate and effective prompt words, or perform repeated debugging to achieve the expected generation effect. As a result, the solutions in the existing technology have the problems of low media generation efficiency and poor generation effect.
[0036] The embodiments of the present disclosure provide an interactive media generation method to solve the above problems.
[0037] refer to Figure 2 , Figure 2 Schematic diagram of the process of the interactive media generation method provided by the embodiment of the present disclosure Figure 1 The method of this embodiment can be applied in a terminal device or a server. In the case where a terminal device executes the method provided by this embodiment, in one possible implementation, the terminal device can implement the interactive media generation method provided by this embodiment by executing a locally deployed program code. In another possible implementation, the server can be used to deploy functional services implemented based on the interactive media generation method provided by this embodiment, and the terminal device can access the above server through (a browser, a client) and call the corresponding functional services to implement the interactive media generation method provided by this embodiment. Exemplarily, the interactive media generation method provided by this embodiment includes:
[0038] Step S101: Display a first interactive interface, which is configured to simultaneously display a media generation control and at least two recommended media, wherein the media generation control is used to configure a functional mode and receive generation information, the functional mode is used to characterize a functional type of a media generation function, and the generation information is used to determine the content of the output media.
[0039] Step S102: In response to a first triggering operation on a target recommended media, generation information corresponding to the target recommended media is obtained, and the generation information is input into a media generation control.
[0040] Step S103: In response to the second triggering operation on the media generation control, output media is generated based on the target functional mode and generation information currently configured by the media generation control.
[0041] In this embodiment, the provided interactive media generation method is introduced with a terminal device as the execution subject. Exemplarily, the terminal device runs a target application with a media generation function, wherein the media in the media generation function can be a picture or a video. This embodiment is introduced by taking the case of generating pictures as an example. Exemplarily, after the target application is started, a first interactive interface is displayed, and the first interactive interface is, for example, a main interface or a default display interface. The first interactive interface simultaneously displays a media generation control and at least two recommended media, wherein the media generation control is used to configure the functional mode and receive generation information, and the functional mode is used to characterize the functional type of the media generation function. The functional type of the media generation function includes, for example: video generation, picture generation, digital human generation, video / image action imitation, music generation, etc. Different media generation functions will generate output media of corresponding content and type. The specific implementation principles of the above-mentioned media generation functions will not be repeated here. The generation information is used to determine the content of the output media. In one possible implementation, the generation information includes text used to describe the content of the output media, that is, prompt words. The media generation model based on AIGC technology will generate content based on the prompt words and generate output media, thereby realizing media generation functions such as video generation, picture generation, digital human generation, and music generation; in another possible implementation, the generation information includes pictures, videos or music used to describe the content of the output media, that is, reference media. The media generation model based on AIGC technology will use the reference media as a reference to generate content and generate output media, thereby realizing media generation functions such as video generation, picture generation, digital human generation, video / image action imitation, and music generation.
[0042] Furthermore, recommended media is output media generated based on AIGC technology and uploaded by other users and recommended to users by the target application, such as videos and pictures generated based on AIGC technology. Among them, the recommended media in the first interactive interface is issued by the server and displayed based on the feed stream, and the media issuance, recommendation process and principle based on the feed stream are not repeated here. In this embodiment, multiple recommended media (recommended media exceeding one screen) in the first interactive interface can be switched in response to user operations, such as scrolling up and down, turning pages left and right, etc. The media generation control is permanently displayed at a preset position in the first interactive interface, thereby realizing the simultaneous display of the media generation control and at least two recommended media in the first interactive interface.
[0043] For example, in a possible implementation, after step S101, the following steps are further included:
[0044] Step S101A: In response to a display switching operation on the first interactive interface, the recommended media in the first interactive interface is switched to be displayed, and at the same time, a media generation control is continuously displayed at a preset position of the first interactive interface.
[0045] Among them, the display switching operation is an operation for controlling the first interactive interface to switch the display, such as sliding page operations, dragging, clicking the scroll bar operations, page turning operations, etc. The terminal device switches the display of the recommended media in the first interactive interface by responding to the display switching operation. At the same time, the preset positions of the first interactive interface, such as the lower and upper parts of the interface, continue to display the above-mentioned media generation controls in a permanent manner, thereby realizing the rapid calling up and input of the media generation controls.
[0046] Furthermore, with respect to step S101A, before the recommended media within the first interactive interface is switched to display in response to a display switching operation on the first interactive interface, a media generated control in a first state is displayed at a first preset position within the first interactive interface, for example, at the top of the first interactive interface. After the recommended media within the first interactive interface begins to be switched to display, the terminal device displays the media generated control in a second state at a second preset position within the first interactive interface. The first state, for example, is a full-size media generated control with additional sub-controls displayed above the first-size media generated control. The second state, for example, is a miniature or hidden state of the media generated control, which is smaller than the first-size control and can increase the effective display area of the interface during the switching process. After the recommended media within the first interactive interface ceases to be switched to display, the media generated control is restored to display at the second preset position within the first interactive interface, either manually or automatically. In this case, the media generated control displayed at the second preset position within the first interactive interface can be based on the first-size control or a third additional control.
[0047] Specifically, for example, after stopping switching to display the recommended media in the first interactive interface, the terminal device still displays the media generation control in the second form (miniature form) in the first interactive interface. Afterwards, when the user applies a trigger operation, such as the first trigger operation for the recommended media, the terminal device displays the media generation control in the first state or the third state at the second preset position of the first interactive interface, so that in subsequent steps, the generation information of the recommended media can be obtained and input into the media generation control for display, and it is convenient to modify and output media generation based on the media generation control, thereby improving interaction efficiency.
[0048] Afterwards, in response to a first trigger operation on the target recommended media, such as a click operation, the terminal device obtains the generation information corresponding to the target recommended media. The generation information corresponding to the target recommended media includes one or more of the prompt words, reference images, or reference videos used when generating the target recommended media. Afterwards, the terminal device inputs the generation information into the media generation control to complete the automatic input of the initial generation information. Thereafter, the user can further modify the above generation information, or add additional information based on the above generation information. The generation information corresponding to the target recommended media includes the prompt word t1. The user further inputs the reference image p1 through the media generation control to improve the above generation information, or directly perform subsequent media generation based on the above initial generation information. Finally, by responding to a second trigger operation on the media generation control, the function of the media generation control is triggered. The media generation control calls the corresponding media processing model according to the currently configured target function mode to process the above generation information, thereby generating the corresponding output media.
[0049] Figure 3 This is a schematic diagram of a first interactive interface provided by an embodiment of the present disclosure. Figure 3 To further introduce the above process, Figure 3 As shown, the first interactive interface displays multiple recommended media, including, for example, recommended picture M1, recommended picture M2, recommended picture M3, recommended picture M4, recommended picture M5, recommended picture M6, recommended video V1, recommended video V2, etc., wherein the above-mentioned recommended media are displayed in multiple pages. At the same time, a resident media generation control is also provided at the bottom of the first interactive interface. When the recommended media in the first interactive interface are scrolled, the media generation control is permanently displayed above or below the first interactive interface, so that the user can input the generated information into the media generation control and display it immediately after applying the first trigger operation to any target recommended media.
[0050] Through the above method, users can "collect inspiration" by browsing the recommended media in the first interactive interface. At the same time, after the user finds "inspiration" (i.e., determines the target recommended media), the corresponding generation information can be input into the media generation control for processing as soon as possible. Through the above interaction scheme, the human-computer interaction efficiency in the media creation scenario can be effectively improved.
[0051] Furthermore, in a possible implementation, the generated information includes prompt words used to generate output media, and / or reference media; the media generation control includes a text input area for receiving prompt words, and / or a reference media input area for receiving reference media, wherein the reference media includes at least one frame of picture and / or at least one video; the step of inputting the generated information into the media generation control in step S102 is specifically implemented by: inputting the target prompt words used when generating the target recommended media into the text input area, and / or inputting the target reference media used when generating the target recommended media into the reference media input area, and a thumbnail of the target reference media can be displayed in the reference media input area.
[0052] Figure 4 A schematic diagram of another first interactive interface provided by an embodiment of the present disclosure, such as Figure 4 As shown, in Figure 3 Based on the illustrated embodiment, the media generation control within the first interactive interface includes a text input area and a reference media input area. The reference media input area can be used to receive reference media, which can be recommended media selected by the user within the first interactive interface or media uploaded by the user via local loading. The text input area can be used to enter prompt words. Similarly, the prompt words within the text input area can be target prompt words corresponding to the recommended media selected by the user within the first interactive interface. For example, as shown in the reference figure, after the user selects "Recommended Image M3," the terminal device directly enters prompt word T3 (target prompt word) corresponding to "Recommended Image M3" into the text input area. Of course, the prompt words within the text input area can also be manually entered by the user, which will not be described in detail. Furthermore, in response to a first trigger operation for the target recommended media, the user can then further modify, replace, add, or delete the target input prompt words or target reference media within the text input area and / or reference media input area, thereby achieving flexible adjustment of the generated information. Furthermore, as shown in the figure, the media generation control also includes an options component for configuring functional modes. The options component displays multiple alternative functional modes (e.g., Function Mode 01, Function Mode 02, etc.) via a drop-down menu, and in response to a user's selection instruction, configures one of them as the target functional mode. The text input area, reference media input area, and options component within the media generation control achieve the effect of configuring functional modes and receiving generated information.
[0053] Furthermore, in a possible implementation, before step S102, the following steps are further included:
[0054] Step S1001: Generate corresponding guidance text based on the target functional mode currently configured by the media generation control, where the guidance text represents prompt word input suggestions for the functional mode.
[0055] Step S1002: Display the guiding text in the text input area.
[0056] Exemplarily, the option component has a default value, that is, the first interactive interface has a default functional mode, such as "image generation", "video generation", etc. When the user does not make any adjustments, the default functional mode is the target functional mode of the current configuration. A guide text is displayed in the text input area within the media generation control, and the guide text is used to guide the user to apply a trigger operation that matches the target functional mode of the current configuration, that is, when the target functional mode of the current configuration is different, the displayed guide text is also different, thereby guiding the user to perform the correct configuration operation. Specifically, for example, when the target functional mode of the current configuration is "image generation", the guide text text_1 is displayed; and when the target functional mode of the current configuration is "video generation", the guide text text_2 is displayed.
[0057] In one possible implementation, the guidance text is displayed in gray text in the text input area. Gray text refers to text displayed in a different font than normal, such as a lower grayscale font or a certain degree of transparency. The guidance text is then removed after the user selects the text input area.
[0058] Furthermore, in a possible implementation, as Figure 5 As shown, the specific implementation of step S1001 includes:
[0059] Step S1001 - 1 : Acquire the media type and media quantity of the target reference media received in the reference media input area.
[0060] Step S1001 - 2 : Generate corresponding guidance text according to the target functional modality and the media type and media quantity of the target reference media.
[0061] Exemplarily, for the process of generating the corresponding guide text, after determining the target functional mode of the current configuration, the media type and media quantity of the target reference media received in the reference media input area can be further combined to refine the guide text, so that the guide text is more accurate and more operational, thereby improving the interaction efficiency. Specifically, the terminal device first obtains the media type and media quantity of the target reference media currently input in the reference media input area within the media generation control, wherein the media types include pictures, videos, and music. The media type and media quantity of the target reference media include, for example, one photo, two videos, and the like. Afterwards, based on the determination of the target functional mode, the corresponding guide text is determined according to the combination of the above media types and media quantities. The following is an example of a more specific embodiment. For example, if the target function modality corresponds to "image generation," when the reference media input area has not received (i.e., uploaded) any media, the generated guidance text is "Enter text to describe the image content and movement style you want to create, such as a 3D image of a little boy wearing a flight jacket, skateboarding in the park." For another example, if the target function modality corresponds to "image generation," when the reference media input area receives a single image, the generated guidance text is "Combining the images, describe the image and movement you want to create, such as waves lapping on the beach and a pink moon rising in the sky." For another example, if the target function modality corresponds to "video generation," when the reference media input area receives two images, the generated guidance text is "Try to keep the first and last frames with the same theme, and use text to describe the transition between the two images, such as "A cute little Antarctic penguin walking slowly on the ice."
[0062] Through the steps of the above embodiment, the generated guidance text can be made more accurate and more operable, thereby improving the efficiency of human-computer interaction and improving the quality of the output media finally generated.
[0063] In this embodiment, a first interactive interface is displayed, configured to simultaneously display a media generation control and at least two recommended media. The media generation control is used to configure a functional mode and receive generation information, wherein the functional mode is used to characterize the functional type of the media generation function, and the generation information is used to determine the content of the output media. In response to a first trigger operation on a target recommended media, generation information corresponding to the target recommended media is obtained and input into the media generation control. In response to a second trigger operation on the media generation control, output media is generated based on the target functional mode and generation information currently configured in the media generation control. By simultaneously displaying the media generation control and recommended media within the first interactive interface, a user can obtain generation information corresponding to the recommended media by selecting the recommended media and input it into the media generation control. Consequently, the corresponding output media is generated based on the target functional mode currently configured in the media generation control. This achieves rapid media generation based on the recommended media, improves media generation efficiency, and ensures the quality and effectiveness of the generated output media.
[0064] refer to Figure 6 , Figure 6 Schematic diagram of the interactive media generation method provided in the embodiment of the present disclosure Figure 2 In this embodiment Figure 2 Based on the embodiment shown, step S102 is further refined, and the interactive media generation method includes:
[0065] Step S201: Display a first interactive interface, which is configured to simultaneously display a media generation control and at least two recommended media, wherein the media generation control is used to configure a functional mode and receive generation information, the functional mode is used to characterize a functional type of a media generation function, and the generation information is used to determine the content of the output media.
[0066] Step S202: In response to the first sub-trigger operation for the target recommended media, jump to the target details page, which is configured to display target generation information for generating the target recommended media and a media generation control.
[0067] Step S203: In response to the second sub-trigger operation on the target details page, a target functional mode corresponding to the target recommended media is obtained.
[0068] Step S204: configuring the media generation control based on the target functional modality, and inputting the target generation information into the media generation control.
[0069] For example, in one possible implementation, the terminal device can jump to the details page corresponding to the target recommended media from the first interactive interface through the corresponding first sub-trigger operation. In the target details page, for example, conventional description information corresponding to the target recommended media, such as the author, release time, etc., can be displayed. In addition, the target generation information through the target recommended media, such as the prompt words and reference pictures used to generate the target recommended media, can also be displayed in the target details page. And the media generation control, that is, the media generation control is permanently displayed on the details page corresponding to the recommended media. Afterwards, in response to the second sub-trigger operation on the target details page, the trigger operation of the "Make the Same Style" button, the terminal device obtains the functional mode of the target recommended media, that is, the target functional mode, and configures the media generation control based on the target functional mode, and inputs the target generation information into the media generation control (and can further adjust the input target generation information as needed), thereby achieving the purpose of quickly and conveniently "imitating" the target recommended media to generate similar media works.
[0070] In a possible implementation, after step S202, the method further includes:
[0071] Step S202A: Display target recommended media in the target details page.
[0072] Step S202B: In response to the third trigger operation on the target recommended media displayed in the target detail page, target generation information of the target recommended media is displayed.
[0073] For example, in another possible implementation, the target generation information of the target recommended media is not displayed by default on the target details page, but the target recommended media (or its preview) is displayed. Then, the user inputs a third trigger operation, and the terminal device responds to the third trigger operation, and then displays the target generation information of the target recommended media. The third trigger operation can be, for example, a double-click operation, a mouse hover operation, etc. Through the above solution of this embodiment, the information display efficiency of the target details page can be further improved, and the display effect can be prevented from being affected by excessive information.
[0074] Furthermore, after completing the steps of the above-mentioned embodiment, the target generation information corresponding to the target recommended media is input into the media generation control through the details page, and the purpose of configuring the media generation control based on the target functional mode corresponding to the target recommended media is achieved. At the same time, since the media generation control is configured in a resident state (displayed in the details page), the target generation information input into the media generation control can be modified at any time as needed, and the corresponding output media can be generated at any time, thereby improving the interaction efficiency and information display efficiency in the entire media generation process.
[0075] Furthermore, optionally, an AI Agent control is also configured in the media generation control. After the AI Agent control is triggered, an AI Agent dialogue interface can be further displayed. Specifically, in this embodiment, the following is also included:
[0076] Step S205: Display the intelligent agent dialogue interface, which is configured to receive the demand statement input by the user, generate corresponding suggestion prompt words, and input the suggestion prompt words into the media generation control.
[0077] Exemplarily, the intelligent agent is constructed based on a large speech model, and is an intelligent program that can perceive the environment, understand it autonomously, make decisions, and execute actions. Through the intelligent agent dialogue interface, which is the interactive interface between the intelligent agent and the user, the intelligent agent dialogue interface can continuously receive the user's input demand statements, and after reasoning, generate generation information that matches their demand statements, including suggested prompt words. Afterwards, the suggested prompt words can be automatically or in response to user operations to input the above-mentioned media generation control as a supplement and improvement to the generation information corresponding to the target recommended media. Of course, the media generation control can also generate the corresponding output media based on the suggested prompt words alone. Through the above steps, another way of inputting generation information into the media generation control is provided, so that the user can complete the purpose of inputting generation information into the media generation control more quickly and efficiently, and improve the quality of the generation information input into the media generation control.
[0078] Step S206: In response to the second triggering operation on the media generation control, output media is generated based on the target functional mode and generation information currently configured by the media generation control.
[0079] Step S207: displaying a stream generation interface, where the stream generation interface is configured to display at least one output medium and to display a media generation control, wherein the output media in the stream generation interface are arranged based on generation time.
[0080] Further, exemplarily, after completing the step of inputting generation information into the media generation control, the functional logic corresponding to the media generation control can be triggered by responding to the second trigger operation on the media generation control, and the corresponding output media can be generated based on the above-mentioned generation information by calling the functional model corresponding to the target functional mode. This process has been introduced in the previous embodiment and will not be repeated here.
[0081] Afterwards, the terminal device jumps to the generation flow interface and displays the previously generated output media, as well as the historical output media generated earlier, based on the generation flow interface. Among them, the output media in the generation flow interface are arranged based on the generation time. That is, the generation flow interface displays different output media based on chronological order. Through the generation flow interface, the recording and display of previously generated output media can be achieved. At the same time, the generation flow interface is configured to display the media generation control, that is, the media generation control is also resident in the generation flow interface. Due to the presence of the media generation control, the purpose of generating new output media at any time can also be achieved in the generation flow interface.
[0082] According to specific needs, optionally, this embodiment also includes:
[0083] Step S207A: In response to the display switching operation on the generated stream interface, the output media in the generated stream interface is switched to be displayed, and at the same time, the media generation control is continuously displayed at a preset position of the generated stream interface.
[0084] For example, when a user is browsing the generation flow interface, the display switching operation is performed to switch the display of the content in the generation flow interface, that is, the output media displayed in the generation flow interface. During this process, the media generation control is continuously displayed at the preset position of the generation flow interface, that is, the media generation control is also resident in the generation flow interface, so that the user can input the generation information into the media generation control for editing and media generation at any time after "getting inspiration". At the same time, referring to the characteristics of the media generation control in the first interactive interface, before and after the start of switching to display the output media in the generated stream interface, the position and state of the media generation control will also change accordingly. For example, before the start of switching to display the output media in the generated stream interface, the media generation control in the first state is displayed at the first preset position in the generated stream interface; after the start of switching to display the output media in the generated stream interface, the media generation control in the second state is displayed at the second preset position in the generated stream interface; after the stop of switching to display the output media in the generated stream interface, the media generation control in the third state is displayed at the second preset position in the generated stream interface, wherein the control size of the media generation control in the first state or the third state is larger than the control size of the media generation control in the second state. The step of displaying the media generation control in the third state at the second preset position in the generated stream interface can be triggered manually or automatically. For details, please refer to the state switching scheme of the media generation control in the first interactive interface, which will not be repeated here.
[0085] Figure 7 A schematic diagram of a generation flow interface provided by an embodiment of the present disclosure, such as Figure 7As shown, exemplarily, within the generated stream interface, three output media items are displayed, including video video_1, picture group [p1, p2, p3, p4] (including pictures p1, p2, p3, and p4), and picture p5. Video video_1 is a video generated at 9:00 on January 1 by the media generation control of the configuration function module M1 (corresponding to the video generation function) based on the generation information info_1; picture group [p1, p2, p3, p4] is a multi-frame picture generated at 9:20 on January 2 by the media generation control of the configuration function module M2 (corresponding to the picture generation function) based on the generation information info_2; and picture p5 is a single-frame picture generated at 9:21 on January 2 by the media generation control of the configuration function module M3 (corresponding to the picture sharpening function) based on the generation information info_3. The above-mentioned video video_1, picture group [p1, p2, p3, p4], and picture p5 are each an output media item and are arranged based on the generation time. At the same time, a media generation control is displayed in a permanent manner at the bottom of the generation stream interface. The specific functions of the media generation control have been introduced in the previous embodiment and will not be repeated here.
[0086] The following continues to introduce the relevant interaction solutions based on the generated flow interface:
[0087] Step S208: In response to the fourth trigger operation on the first output media in the generation flow interface, first generation information corresponding to the first output media is obtained, and the first generation information is input into the media generation control.
[0088] Step S209: In response to the fifth trigger operation on the media generation control in the generation flow interface, the first generation information is modified into second generation information, and second output media is generated in the generation flow interface based on the second generation information.
[0089] Exemplarily, based on the previous introduction, at least one previously generated output media is displayed in the generation flow interface, for example, based on a chronological order. In response to the user's fourth trigger operation on the first output media in the generation flow interface, the terminal device can obtain the first generation information corresponding to the first output media. Thereafter, the first generation information is input into the resident media generation control in the generation flow interface, thereby achieving the purpose of quickly extracting generation information from the historically generated output media. Thereafter, the user can modify the first generation information into the second generation information through the fifth trigger operation (for example, manually deleting or adding the prompt word) based on the first generation information, and generate the second output media based on the second generation information, and display it in the survival flow interface. The above-mentioned interactive scheme provided by this embodiment further improves the efficiency and real-time performance of users in obtaining "inspiration" and converting "inspiration" into media works.
[0090] In a possible implementation, in this embodiment, after step S208, the following steps are further included:
[0091] Step S208A: Obtain a first functional mode corresponding to the first output media, and configure a media generation control based on the first functional mode.
[0092] Accordingly, after executing step S208A, the specific implementation of step S209 includes:
[0093] Step S2090: calling the image processing model corresponding to the first functional modality to process the second generated information, generate a second output medium, and display it in the generated flow interface.
[0094] Exemplarily, in the steps of this embodiment, in response to the fourth trigger operation for the first output media in the generation flow interface, while obtaining the first generation information corresponding to the first output media, the terminal device simultaneously obtains the first functional mode (corresponding to the image generation function) corresponding to the first output media, and configures the media generation control based on the first functional mode. If the media generation control is already configured as the first functional mode at this time, no adjustment is required; if the media generation control is not configured as the first functional mode at this time, for example, it is configured as the second functional mode (corresponding to the video generation function), the functional mode of the media generation control is adjusted to the first functional mode.
[0095] Furthermore, in another possible implementation, the output media includes at least two output preview media with different media contents. In this embodiment, the following is further included:
[0096] Step S210: Displaying each output preview media and the predicted difference parameters corresponding to each output preview media in the generated stream interface, wherein the predicted difference parameters are hidden prompt words that cause the media content differences between each output preview media, and the hidden prompt words are not included in the generated information of the generated output media.
[0097] After executing step S210, the specific implementation of step S208 includes:
[0098] Step S208B: In response to the fourth trigger operation for generating the first output preview media in the stream interface, obtain first generation information and predicted difference parameters corresponding to the first output preview media, and input the first generation information and predicted difference parameters into the media generation control.
[0099] For example, in one possible implementation, the media generation model called by the media generation control, when generating media such as pictures, usually outputs multiple results for the user to further select, that is, the output media includes multiple output preview media with different media content, such as multiple pictures and videos with different contents. Figure 7The image group [p1, p2, p3, p4] in the generated stream interface shown here represents multiple output preview media with different media content. However, due to the black-box nature of the media generation model, users cannot determine the cause of these image content differences (media content differences). Therefore, after modifying the generation information corresponding to a particular output preview media to generate the desired second generation information, the output media generated based on the second generation information may not reproduce the characteristics of the output preview media, resulting in inefficient content control.
[0100] In this embodiment, the media generation model invoked by the media generation control generates multiple output preview media simultaneously and also generates predicted difference parameters for each output preview media. These predicted difference parameters are hidden prompts that indicate differences in the media content between the output preview media, and these hidden prompts are not included in the generated information for the output media. Subsequently, as needed, the user can further select the predicted difference parameters corresponding to the first output preview media by applying a fourth trigger operation, input the first generated information and the predicted difference parameters into the media generation control, and further modify them as needed to achieve more refined control over the generated content.
[0101] Figure 8 This is another schematic diagram of a generated flow interface provided by an embodiment of the present disclosure, such as Figure 8As shown, in an exemplary embodiment, a picture group P (output media) is displayed within the generated stream interface. Picture group P includes pictures p1, p2, p3, and p4 (four output preview media). Picture group P is a multi-frame image generated by the media generation control of configuration function module M2 (corresponding to the picture generation function) based on generation information info_1. Generation information info_1 includes prompt words and reference images. The prompt words include: "A 3D image of a little boy wearing a flight jacket, skateboarding in a park." Pictures p1, p2, p3, and p4 in picture group P each have different media content. Each picture is displayed with a corresponding hidden prompt word that is not included in the prompt word corresponding to generation information info_1. For example, as shown in the figure, the hidden prompt words corresponding to picture p1 are: "Green shirt" and "Hands open," and the hidden prompt words corresponding to picture p2 are: "Clothes and skateboard are the same color" and "Hands in pockets," etc. Afterwards, for example, after the user clicks on picture p1, the hidden prompt words "green top" and "open hands" corresponding to picture p1, as well as the generation information info_1 used when generating picture group P, can be input into the media generation control together, thereby realizing differentiated hidden prompt word extraction for different output preview media, allowing users to extract personalized image elements in the picture more finely and improve the control accuracy of content generation.
[0102] In this embodiment, the implementation of step S201 and step S206 is the same as that of the present disclosure. Figure 2 The implementation methods of step S101 and step S103 in the illustrated embodiment are the same and will not be described in detail here.
[0103] Corresponding to the interactive media generation method of the above embodiment, Figure 9 This is a block diagram of the structure of an interactive media generation device provided in an embodiment of the present disclosure. The methods described in the above embodiments can be executed by this interactive media generation device, which can be implemented using software and / or hardware and integrated into an electronic device with certain data processing capabilities. These electronic devices may include, but are not limited to, mobile terminals with big data processing capabilities, as well as fixed terminals with big data processing capabilities, such as desktop computers and supercomputers.
[0104] For ease of explanation, only the parts related to the embodiments of the present disclosure are shown. Figure 9 , the interactive media generating device 3 includes:
[0105] A display module 31 is configured to display a first interactive interface, wherein the first interactive interface is configured to simultaneously display a media generation control and at least two recommended media. The media generation control is configured to configure a functional mode and receive generation information. The functional mode is used to represent the functional type of the media generation function. The generation information is used to determine the content of the output media.
[0106] A first interaction module 32 is configured to obtain generation information corresponding to the target recommended media in response to a first trigger operation on the target recommended media, and input the generation information into the media generation control;
[0107] The second interaction module 33 is configured to generate output media in response to a second triggering operation on the media generation control and based on the target functional mode and generation information currently configured for the media generation control.
[0108] According to one or more embodiments of the present disclosure, after displaying the first interactive interface, the display module 31 is also used to: in response to a display switching operation for the first interactive interface, switch to display the recommended media within the first interactive interface, and continuously display the media generation control at a preset position of the first interactive interface.
[0109] According to one or more embodiments of the present disclosure, before switching to display the recommended media in the first interactive interface in response to a display switching operation on the first interactive interface, the display module 31 is also used to: display a media generation control in a first state at a first preset position in the first interactive interface; when the display module 31 switches to display the recommended media in the first interactive interface and continuously displays the media generation control at the preset position of the first interactive interface, it is specifically used to: display the media generation control in the second state at a second preset position of the first interactive interface after starting to switch to display the recommended media in the first interactive interface; and display the media generation control in the third state at the second preset position of the first interactive interface after stopping switching to display the recommended media in the first interactive interface, wherein the control size of the media generation control in the first state or the third state is larger than the control size of the media generation control in the second state.
[0110] According to one or more embodiments of the present disclosure, when the display module 31 displays the media generation control in the third state at the second preset position of the first interactive interface, it is specifically used to: display the media generation control in the third state at the second preset position of the first interactive interface in response to a first trigger operation for recommended media.
[0111] According to one or more embodiments of the present disclosure, the generated information includes prompt words used to generate output media, and / or reference media, and the reference media includes at least one frame of picture and / or at least one video; the media generation control includes a text input area for receiving prompt words, and / or a reference media input area for receiving reference media; when the first interactive module 32 inputs the generated information into the media generation control, it is specifically used to: input the target prompt words used when generating target recommended media into the text input area, and / or input the target reference media used when generating target recommended media into the reference media input area, and display a thumbnail of the target reference media in the reference media input area.
[0112] According to one or more embodiments of the present disclosure, the first interaction module 32 is further used to: generate corresponding guidance text based on the target functional mode currently configured by the media generation control, the guidance text representing the prompt word input suggestion for the functional mode; and display the guidance text in the text input area.
[0113] According to one or more embodiments of the present disclosure, when the first interaction module 32 generates corresponding guidance text based on the target functional mode currently configured by the media generation control, it is specifically used to: obtain the media type and media quantity of the target reference media received by the reference media input area; and generate corresponding guidance text according to the target functional mode and the media type and media quantity of the target reference media.
[0114] According to one or more embodiments of the present disclosure, the first trigger operation includes a first sub-trigger operation and a second sub-trigger operation. When the first interactive module 32 obtains the generation information corresponding to the target recommended media in response to the first trigger operation for the target recommended media and inputs the generation information into the media generation control, it is specifically used to: in response to the first sub-trigger operation for the target recommended media, jump to the target details page, the target details page is configured to display the target generation information for generating the target recommended media, and the media generation control; in response to the second sub-trigger operation for the target details page, obtain the target function mode corresponding to the target recommended media; configure the media generation control based on the target function mode, and input the target generation information into the media generation control.
[0115] According to one or more embodiments of the present disclosure, the first interaction module 32 is further used to: display target recommended media in the target details page; and display target generation information of the target recommended media in response to a third trigger operation on the target recommended media displayed in the target details page.
[0116] According to one or more embodiments of the present disclosure, the first interaction module 32 is also used for at least one of the following: displaying a generation flow interface, the generation flow interface is configured to display at least one output media, wherein the output media in the generation flow interface are arranged based on the generation time; displaying an intelligent agent dialogue interface, the intelligent agent dialogue interface is configured to receive a demand statement input by a user and generate corresponding suggestion prompt words, and inputting the suggestion prompt words into the media generation control.
[0117] According to one or more embodiments of the present disclosure, the generation flow interface is configured to display a media generation control, and the first interaction module 32 is further used to: in response to a fourth trigger operation on the first output media in the generation flow interface, obtain first generation information corresponding to the first output media, and input the first generation information into the media generation control; in response to a fifth trigger operation on the media generation control in the generation flow interface, modify the first generation information into second generation information; the second interaction module 33 is specifically used to: generate second output media in the generation flow interface based on the second generation information.
[0118] According to one or more embodiments of the present disclosure, after obtaining the first generation information corresponding to the first output media, the first interaction module 32 is further used to: obtain the first functional mode corresponding to the first output media, and configure the media generation control based on the first functional mode; when the second interaction module 33 generates the second output media in the generation flow interface based on the second generation information, it is specifically used to: call the image processing model corresponding to the first functional mode to process the second generation information, generate the second output media, and display it in the generation flow interface.
[0119] According to one or more embodiments of the present disclosure, the output media includes at least two output preview media with media content differences. The first interactive module 32 is further used to: display each output preview media and the predicted difference parameters corresponding to each output preview media in the generated flow interface, wherein the predicted difference parameters are hidden prompt words that cause the media content differences between each output preview media, and the hidden prompt words are not included in the generated information of the generated output media; when the first interactive module 32 obtains the first generation information corresponding to the first output media in response to the fourth trigger operation for the first output media in the generated flow interface, and inputs the first generation information into the media generation control, it is specifically used to: obtain the first generation information and predicted difference parameters corresponding to the first output preview media in response to the fourth trigger operation for the first output preview media in the generated flow interface, and input the first generation information and predicted difference parameters into the media generation control.
[0120] The display module 31, the first interaction module 32 and the second interaction module 33 are connected in sequence. The interactive media generating device 3 provided in this embodiment can implement the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, which will not be repeated in this embodiment.
[0121] Figure 10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure is shown in FIG. Figure 10 As shown, the electronic device 4 includes:
[0122] A processor 41, and a memory 42 communicatively connected to the processor 41;
[0123] Memory 42 stores computer-executable instructions;
[0124] The processor 41 executes the computer execution instructions stored in the memory 42 to implement the following Figure 2-Figure 8 The interactive media generation method in the illustrated embodiment.
[0125] Optionally, the processor 41 and the memory 42 are connected via a bus 43 .
[0126] For related instructions, please refer to Figure 2-Figure 8 The relevant descriptions and effects corresponding to the steps in the corresponding embodiments can be understood, and no further details are given here.
[0127] The present invention provides a computer-readable storage medium that stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the computer-executable instructions are used to implement the present invention. Figure 2-Figure 8 The interactive media generation method provided in any one of the corresponding embodiments.
[0128] The present invention provides a computer program product, including a computer program, which implements the present invention when executed by a processor. Figure 2-Figure 8 The interactive media generation method provided in any one of the corresponding embodiments.
[0129] In order to implement the above embodiment, the embodiment of the present disclosure further provides an electronic device.
[0130] refer to Figure 11, which shows a schematic structural diagram of an electronic device 900 suitable for implementing an embodiment of the present disclosure. The electronic device 900 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 11 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0131] like Figure 11 As shown, the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0132] Typically, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 11 The electronic device 900 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0133] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0134] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0135] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0136] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0137] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0139] The units or modules involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit or module does not, in some cases, limit the unit itself.
[0140] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0141] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0142] In a first aspect, according to one or more embodiments of the present disclosure, there is provided an interactive media generation method, comprising:
[0143] A first interactive interface is displayed, wherein the first interactive interface is configured to simultaneously display a media generation control and at least two recommended media, wherein the media generation control is used to configure a functional mode and receive generation information, the functional mode is used to characterize a functional type of a media generation function, and the generation information is used to determine the content of the output media; in response to a first trigger operation for a target recommended media, generation information corresponding to the target recommended media is obtained, and the generation information is input into the media generation control; in response to a second trigger operation for the media generation control, the output media is generated based on the target functional mode currently configured for the media generation control and the generation information.
[0144] According to one or more embodiments of the present disclosure, after displaying the first interactive interface, the method further includes: in response to a display switching operation on the first interactive interface, switching to display the recommended media in the first interactive interface, and continuously displaying the media generation control at a preset position of the first interactive interface;
[0145] According to one or more embodiments of the present disclosure, before the display switching operation on the first interactive interface is performed to switch the display of the recommended media in the first interactive interface, it also includes: displaying a media generation control in a first state at a first preset position in the first interactive interface; the switching to display the recommended media in the first interactive interface while continuously displaying the media generation control at the preset position of the first interactive interface includes: after starting to switch to display the recommended media in the first interactive interface, displaying the media generation control in the second state at a second preset position of the first interactive interface; after stopping to switch to display the recommended media in the first interactive interface, displaying the media generation control in the third state at the second preset position of the first interactive interface, wherein the control size of the media generation control in the first state or the third state is larger than the control size of the media generation control in the second state.
[0146] According to one or more embodiments of the present disclosure, displaying the media generation control in the third state at the second preset position of the first interactive interface includes: displaying the media generation control in the third state at the second preset position of the first interactive interface in response to a first trigger operation for the recommended media.
[0147] According to one or more embodiments of the present disclosure, the generation information includes prompt words used to generate output media, and / or reference media, and the reference media includes at least one frame of picture and / or at least one video; the media generation control contains a text input area for receiving the prompt words, and / or a reference media input area for receiving the reference media; the inputting of the generation information into the media generation control includes: inputting the target prompt words used when generating the target recommended media into the text input area, and / or inputting the target reference media used when generating the target recommended media into the reference media input area, and displaying a thumbnail of the target reference media in the reference media input area.
[0148] According to one or more embodiments of the present disclosure, the method further includes: generating corresponding guide text based on the target functional mode currently configured by the media generation control, the guide text representing prompt word input suggestions for the functional mode; and displaying the guide text in the text input area.
[0149] According to one or more embodiments of the present disclosure, the target functional mode currently configured based on the media generation control is used to generate corresponding guidance text, including: obtaining the media type and media quantity of the target reference media received by the reference media input area; and generating corresponding guidance text according to the target functional mode and the media type and media quantity of the target reference media.
[0150] According to one or more embodiments of the present disclosure, the first trigger operation includes a first sub-trigger operation and a second sub-trigger operation, and the response to the first trigger operation for the target recommended media, obtaining the generation information corresponding to the target recommended media, and inputting the generation information into the media generation control, includes: responding to the first sub-trigger operation for the target recommended media, jumping to the target details page, the target details page is configured to display the target generation information for generating the target recommended media, and the media generation control; responding to the second sub-trigger operation for the target details page, obtaining the target function mode corresponding to the target recommended media; configuring the media generation control based on the target function mode, and inputting the target generation information into the media generation control.
[0151] According to one or more embodiments of the present disclosure, the method further includes: displaying the target recommended media in the target details page; and displaying target generation information of the target recommended media in response to a third trigger operation on the target recommended media displayed in the target details page.
[0152] According to one or more embodiments of the present disclosure, it also includes at least one of the following: displaying a generation flow interface, wherein the generation flow interface is configured to display at least one of the output media, wherein the output media in the generation flow interface are arranged based on the generation time; displaying an intelligent agent dialogue interface, wherein the intelligent agent dialogue interface is configured to receive a demand statement input by a user and generate corresponding suggestion prompt words, and inputting the suggestion prompt words into the media generation control.
[0153] According to one or more embodiments of the present disclosure, the generation flow interface is configured to display the media generation control, and the method further includes: in response to a fourth trigger operation on the first output media within the generation flow interface, obtaining first generation information corresponding to the first output media, and inputting the first generation information into the media generation control; in response to a fifth trigger operation on the media generation control within the generation flow interface, modifying the first generation information to second generation information, and generating second output media within the generation flow interface based on the second generation information.
[0154] According to one or more embodiments of the present disclosure, after obtaining the first generation information corresponding to the first output media, it also includes: obtaining the first functional mode corresponding to the first output media, and configuring the media generation control based on the first functional mode; generating the second output media in the generation flow interface based on the second generation information, including: calling the image processing model corresponding to the first functional mode to process the second generation information, generating the second output media, and displaying it in the generation flow interface.
[0155] According to one or more embodiments of the present disclosure, the output media includes at least two output preview media with media content differences, and the method further includes: displaying each of the output preview media and a predicted difference parameter corresponding to each of the output preview media in the generation flow interface, wherein the predicted difference parameter is a hidden prompt word that causes the media content difference between each of the output preview media, and the hidden prompt word is not included in the generation information for generating the output media; the obtaining of the first generation information corresponding to the first output media in response to the fourth trigger operation on the first output media in the generation flow interface, and inputting the first generation information into the media generation control, includes: obtaining the first generation information and predicted difference parameter corresponding to the first output preview media in response to the fourth trigger operation on the first output preview media in the generation flow interface, and inputting the first generation information and predicted difference parameter into the media generation control.
[0156] In a second aspect, according to one or more embodiments of the present disclosure, there is provided an interactive media generating apparatus, comprising:
[0157] a display module configured to display a first interactive interface configured to simultaneously display a media generation control and at least two recommended media, wherein the media generation control is configured to configure a functional mode and receive generation information, the functional mode being used to characterize a functional type of a media generation function, and the generation information being used to determine content of output media;
[0158] A first interaction module is configured to obtain generation information corresponding to a target recommended media in response to a first triggering operation on the target recommended media, and input the generation information into the media generation control;
[0159] The second interaction module is configured to generate the output media in response to a second trigger operation on the media generation control based on the target functional mode currently configured by the media generation control and the generation information.
[0160] According to one or more embodiments of the present disclosure, after displaying the first interactive interface, the display module is further used to: in response to a display switching operation for the first interactive interface, switch the display of the recommended media within the first interactive interface, and continuously display the media generation control at a preset position of the first interactive interface.
[0161] According to one or more embodiments of the present disclosure, before switching to display the recommended media in the first interactive interface in response to a display switching operation on the first interactive interface, the display module is further used to: display a media generation control in a first state at a first preset position in the first interactive interface; when the display module switches to display the recommended media in the first interactive interface and continuously displays the media generation control at the preset position of the first interactive interface, the display module is specifically used to: display the media generation control in the second state at a second preset position of the first interactive interface after starting to switch to display the recommended media in the first interactive interface; and display the media generation control in the third state at the second preset position of the first interactive interface after stopping switching to display the recommended media in the first interactive interface, wherein the control size of the media generation control in the first state or the third state is larger than the control size of the media generation control in the second state.
[0162] According to one or more embodiments of the present disclosure, when the display module displays the media generation control in the third state at the second preset position of the first interactive interface, it is specifically used to: display the media generation control in the third state at the second preset position of the first interactive interface in response to a first trigger operation for the recommended media.
[0163] According to one or more embodiments of the present disclosure, the generation information includes prompt words used to generate output media, and / or reference media, and the reference media includes at least one frame of picture and / or at least one video; the media generation control contains a text input area for receiving the prompt words, and / or a reference media input area for receiving the reference media; when the first interactive module inputs the generation information into the media generation control, it is specifically used to: input the target prompt words used when generating the target recommended media into the text input area, and / or input the target reference media used when generating the target recommended media into the reference media input area, and display a thumbnail of the target reference media in the reference media input area.
[0164] According to one or more embodiments of the present disclosure, the first interaction module is further used to: generate corresponding guidance text based on the target functional mode currently configured by the media generation control, the guidance text representing the prompt word input suggestion for the functional mode; and display the guidance text in the text input area.
[0165] According to one or more embodiments of the present disclosure, when the first interaction module generates corresponding guidance text based on the target functional mode currently configured by the media generation control, it is specifically used to: obtain the media type and media quantity of the target reference media received by the reference media input area; and generate corresponding guidance text according to the target functional mode and the media type and media quantity of the target reference media.
[0166] According to one or more embodiments of the present disclosure, the first trigger operation includes a first sub-trigger operation and a second sub-trigger operation. When the first interactive module obtains the generation information corresponding to the target recommended media in response to the first trigger operation for the target recommended media and inputs the generation information into the media generation control, it is specifically used to: in response to the first sub-trigger operation for the target recommended media, jump to the target details page, the target details page is configured to display the target generation information for generating the target recommended media, and the media generation control; in response to the second sub-trigger operation for the target details page, obtain the target function mode corresponding to the target recommended media; configure the media generation control based on the target function mode, and input the target generation information into the media generation control.
[0167] According to one or more embodiments of the present disclosure, the first interaction module is further used to: display the target recommended media in the target details page; and display target generation information of the target recommended media in response to a third trigger operation on the target recommended media displayed in the target details page.
[0168] According to one or more embodiments of the present disclosure, the first interaction module is also used for at least one of the following: displaying a generation flow interface, the generation flow interface is configured to display at least one of the output media, wherein the output media in the generation flow interface are arranged based on the generation time; displaying an intelligent agent dialogue interface, the intelligent agent dialogue interface is configured to receive a demand statement input by a user and generate corresponding suggestion prompt words, and inputting the suggestion prompt words into the media generation control.
[0169] According to one or more embodiments of the present disclosure, the generation flow interface is configured to display the media generation control, and the first interaction module is further used to: in response to a fourth trigger operation on the first output media in the generation flow interface, obtain first generation information corresponding to the first output media, and input the first generation information into the media generation control; in response to a fifth trigger operation on the media generation control in the generation flow interface, modify the first generation information to second generation information; the second interaction module is specifically used to: generate second output media in the generation flow interface based on the second generation information.
[0170] According to one or more embodiments of the present disclosure, after obtaining the first generation information corresponding to the first output media, the first interaction module is further used to: obtain the first functional mode corresponding to the first output media, and configure the media generation control based on the first functional mode; when the second interaction module generates the second output media in the generation flow interface based on the second generation information, it is specifically used to: call the image processing model corresponding to the first functional mode to process the second generation information, generate the second output media, and display it in the generation flow interface.
[0171] According to one or more embodiments of the present disclosure, the output media includes at least two output preview media with media content differences, and the first interactive module is further used to: display each of the output preview media and the predicted difference parameters corresponding to each of the output preview media in the generation flow interface, wherein the predicted difference parameters are hidden prompt words that cause the media content differences between the output preview media, and the hidden prompt words are not included in the generation information for generating the output media; when the first interactive module obtains the first generation information corresponding to the first output media in response to the fourth trigger operation on the first output media in the generation flow interface and inputs the first generation information into the media generation control, the first interactive module is specifically used to: obtain the first generation information and predicted difference parameters corresponding to the first output preview media in response to the fourth trigger operation on the first output preview media in the generation flow interface, and input the first generation information and predicted difference parameters into the media generation control.
[0172] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: at least one processor and a memory;
[0173] The memory stores computer-executable instructions;
[0174] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the interactive media generation method as described in the first aspect and various possible designs of the first aspect.
[0175] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the interactive media generation method described in the first aspect and various possible designs of the first aspect is implemented.
[0176] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the interactive media generation method as described in the first aspect and various possible designs of the first aspect.
[0177] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0178] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0179] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An interactive media generation method, characterized in that: include: Displaying a first interactive interface, wherein the first interactive interface is configured to simultaneously display a media generation control and at least two recommended media, wherein the media generation control is used to configure a functional mode and receive generation information, the functional mode is used to characterize a functional type of a media generation function, and the generation information is used to determine content of output media; In response to a first trigger operation on a target recommended media, obtaining generation information corresponding to the target recommended media, and inputting the generation information into the media generation control; In response to a second trigger operation on the media generation control, the output media is generated based on the target functional modality currently configured by the media generation control and the generation information.
2. The method according to claim 1, characterized in that After displaying the first interactive interface, the method further includes: In response to a display switching operation on the first interactive interface, the recommended media in the first interactive interface is switched to be displayed, and at the same time, the media generation control is continuously displayed at a preset position of the first interactive interface.
3. The method according to claim 2, characterized in that Before switching to display the recommended media in the first interactive interface in response to the display switching operation on the first interactive interface, the method further includes: Displaying a media generation control in a first state at a first preset position in the first interactive interface; The switching and displaying of the recommended media in the first interactive interface and continuously displaying the media generation control at a preset position in the first interactive interface includes: After starting to switch and display the recommended media in the first interactive interface, displaying a media generation control in a second state at a second preset position of the first interactive interface; After stopping switching to display the recommended media in the first interactive interface, a media generated control in a third state is displayed at a second preset position of the first interactive interface, wherein the control size of the media generated control in the first state or the third state is larger than the control size of the media generated control in the second state.
4. The method according to claim 3, characterized in that Displaying the media generation control in the third state at the second preset position of the first interactive interface includes: In response to a first triggering operation for the recommended media, the media generation control in the third state is displayed at a second preset position of the first interactive interface.
5. The method according to claim 1, characterized in that The generation information includes a prompt word used to generate the output media, and / or reference media, wherein the reference media includes at least one frame of an image and / or at least one video segment; the media generation control includes a text input area for receiving the prompt word, and / or a reference media input area for receiving the reference media; The inputting the generation information into the media generation control comprises: The target prompt word used when generating the target recommended media is input into the text input area, and / or the target reference media used when generating the target recommended media is input into the reference media input area, and a thumbnail of the target reference media is displayed in the reference media input area.
6. The method according to claim 5, characterized in that The method further comprises: Generate corresponding guidance text based on the target functional mode currently configured by the media generation control, wherein the guidance text represents a prompt word input suggestion for the functional mode; The guiding text is displayed in the text input area.
7. The method according to claim 6, characterized in that The generating corresponding guidance text based on the target function mode currently configured by the media generation control includes: Acquire the media type and media quantity of the target reference media received in the reference media input area; Generate corresponding guidance text according to the target functional modality and the media type and media quantity of the target reference media.
8. The method according to claim 1, characterized in that The first trigger operation includes a first sub-trigger operation and a second sub-trigger operation. In response to the first trigger operation for the target recommended media, obtaining generation information corresponding to the target recommended media and inputting the generation information into the media generation control includes: In response to a first sub-trigger operation for a target recommended media, jumping to a target details page, wherein the target details page is configured to display target generation information for generating the target recommended media and the media generation control; In response to a second sub-trigger operation on the target details page, obtaining a target functional mode corresponding to the target recommended media; The media generation control is configured based on the target functional modality, and the target generation information is input into the media generation control.
9. The method according to claim 8, characterized in that Also includes: Displaying the target recommended media in the target details page; In response to a third trigger operation on the target recommended media displayed in the target detail page, target generation information of the target recommended media is displayed.
10. The method according to claim 1, characterized in that Also includes: displaying a generation flow interface, wherein the generation flow interface is configured to display at least one of the output media and the media generation control, wherein the output media in the generation flow interface are arranged based on generation time; In response to a display switching operation on the generated stream interface, the output media in the generated stream interface is switched to be displayed, and at the same time, the media generation control is continuously displayed at a preset position of the generated stream interface.
11. The method according to claim 10, characterized in that The method further comprises: In response to a fourth trigger operation on a first output medium in the generation flow interface, obtaining first generation information corresponding to the first output medium, and inputting the first generation information into the media generation control; In response to a fifth trigger operation on a media generation control in the generation flow interface, the first generation information is modified into second generation information, and second output media is generated in the generation flow interface based on the second generation information.
12. The method according to claim 11, characterized in that After obtaining the first generation information corresponding to the first output medium, the method further includes: Acquire a first functional mode corresponding to the first output media, and configure the media generation control based on the first functional mode; Generating the second output media in the generation flow interface based on the second generation information includes: The image processing model corresponding to the first functional mode is called to process the second generation information, generate the second output media, and display it in the generation flow interface.
13. The method according to claim 11, characterized in that The output media includes at least two output preview media with different media contents, and the method further includes: Displaying, within the generated stream interface, each of the output preview media and a predicted difference parameter corresponding to each of the output preview media, wherein the predicted difference parameter is a hidden prompt word that causes the media content difference between the output preview media, and the hidden prompt word is not included in the generated information for generating the output media; The step of obtaining, in response to a fourth trigger operation on the first output media in the generation stream interface, first generation information corresponding to the first output media and inputting the first generation information into the media generation control includes: In response to a fourth trigger operation on the first output preview media in the generation stream interface, first generation information and predicted difference parameters corresponding to the first output preview media are obtained, and the first generation information and predicted difference parameters are input into the media generation control.
14. An interactive media generating device, characterized in that include: a display module configured to display a first interactive interface configured to simultaneously display a media generation control and at least two recommended media, wherein the media generation control is configured to configure a functional mode and receive generation information, the functional mode being used to characterize a functional type of a media generation function, and the generation information being used to determine content of output media; A first interaction module is configured to obtain generation information corresponding to a target recommended media in response to a first triggering operation on the target recommended media, and input the generation information into the media generation control; The second interaction module is configured to generate the output media in response to a second trigger operation on the media generation control based on the target functional mode currently configured by the media generation control and the generation information.
15. An electronic device, characterized in that: include: processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the interactive media generation method according to any one of claims 1 to 13.
16. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the interactive media generation method according to any one of claims 1 to 13 is implemented.
17. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the interactive media generating method according to any one of claims 1 to 13 is implemented.
Citation Information
Patent Citations
Video generation method, device and equipment, computer readable storage medium and product
CN118413717A
Media file generation method and device
CN118764686A
Data processing method and device, computer equipment and storage medium
CN118803372A
Multimedia data generation method and device, storage medium and electronic equipment
CN119941920A
Generation and presentation of interactive information cards for a video
US10620801B1
Cited By
Interaction system and method for generating Web end by AI content, and electronic equipment
CN121433485A