Method and apparatus for generating media content, and device and storage medium
By providing candidate media samples and attribute items, users can select and edit prompts, which solves the problem of insufficient accuracy of prompts in generative AI-generated media content and improves the quality of generated media content.
Patent Information
- Application Number
- PCT/CN2025/091825
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-29
- Filing Date
- 2025-04-28
- Publication Date
- 2025-11-06
AI Technical Summary
Existing technologies, when using generative artificial intelligence to generate media content, suffer from insufficient accuracy in the expression of prompts, which affects the quality of generated media content.
It provides a set of candidate media samples and multiple preset attribute items, allowing users to select and edit media samples and their attribute information, and generate target media content by refining the editing prompts.
This improved the accuracy of prompts and editing efficiency, thereby enhancing the quality of the generated media content.
Smart Images

Figure CN2025091825_06112025_PF_FP_ABST
Abstract
Description
Method, device, equipment and storage medium for generating media content
[0001] The present application claims priority to the Chinese patent application No. 202410534129.4, filed on April 29, 2024, entitled “Method, device, equipment and storage medium for generating media content”, the whole content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] The example embodiments of the present disclosure generally relate to the field of computer, and in particular, to a method, device, equipment and computer readable storage medium for generating media content. BACKGROUND
[0003] With the continuous development of computer technology, generative artificial intelligence technology is gradually applied in various fields. People can use generative artificial intelligence technology to create various types of media content, such as pictures, videos, etc. In such a creation process, how to better reference the generation process of media content to improve the generation quality of media content is the focus of people's attention. SUMMARY
[0004] In a first aspect of the present disclosure, a method for generating media content is provided. The method comprises: providing a set of candidate media samples in a target interface, the set of candidate media samples corresponding to a set of preset prompt information; presenting, based on a selection of a target media sample in the set of candidate media samples, first prompt information corresponding to the target media sample in the target interface; presenting, in the target interface, a set of attribute information of the target media sample with respect to a plurality of preset attribute items, the plurality of preset attribute items being used to describe different aspects of the media content to be generated; modifying, based on an editing of first attribute information in the set of attribute information, the first prompt information to second prompt information corresponding to the edited second attribute information; and generating, based on the second prompt information, the target media content.
[0005] In a second aspect of the disclosure, an apparatus for generating media content is provided. The apparatus includes: a sample providing module configured to provide a set of candidate media samples in a target interface, the set of candidate media samples corresponding to a set of preset prompt information; a first presenting module configured to present, in the target interface, first prompt information corresponding to a target media sample in the set of candidate media samples based on a selection of the target media sample; a second presenting module configured to present, in the target interface, a set of attribute information of the target media sample with respect to a plurality of preset attribute items, the plurality of preset attribute items being used to describe different aspects of the media content to be generated; an information processing module configured to modify the first prompt information to second prompt information corresponding to edited second attribute information based on an edit of first attribute information in the set of attribute information; and a content generating module configured to generate the target media content based on the second prompt information.
[0006] In a third aspect of the disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of the disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon a computer program, which is executable by a processor to implement the method of the first aspect.
[0008] It should be understood that the content described in this section is not intended to limit key features or important features of the embodiments of the disclosure, nor is it used to limit the scope of the disclosure. Other features of the disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS
[0009] The above and other features, advantages, and aspects of embodiments of the disclosure will become more apparent by describing in detail some embodiments thereof with reference to the annexed drawings in which:
[0010] FIG. 1 shows a schematic diagram of an example environment in which embodiments according to the disclosure can be implemented;
[0011] FIGS. 2A to 2D show example interfaces according to some embodiments of the disclosure;
[0012] FIG. 3 shows a flowchart of an example process of generating media content according to some embodiments of the disclosure;
[0013] FIG. 4 shows a schematic structural block diagram of an example apparatus for generating media content according to some embodiments of the disclosure; and
[0014] FIG. 5 illustrates a block diagram of an electronic device that can implement various embodiments of the disclosure. DETAILED DESCRIPTION
[0015] Embodiments of the disclosure will be described in more detail with reference to the accompanying drawings. While certain embodiments of the disclosure will be illustrated herein and described, it is to be understood that the disclosure is not limited to the embodiments set forth herein, but applies to any metal object, and thus generally to all compatible metal objects and methods embodied or of which equivalents exist that incorporate the principles of the disclosure. It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure, as claimed.
[0016] It should be noted that the headings provided herein are for convenience only and are not to be construed as limiting. Various embodiments are described herein, and any type of embodiment can be included under any heading. Further, embodiments described in any heading can be combined with any other embodiment described in the same heading and / or a different heading in any manner.
[0017] In the description of embodiments of the disclosure, the term "includes" and its derivatives, such as "including," should be understood in an open, inclusive sense, that is, "including, but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "an embodiment" should be understood as "at least one embodiment." The term "some embodiments" should be understood as "at least some embodiments." Other explicit or implicit definitions can also be included below. The terms "first," "second," etc. can refer to different or the same objects. Other explicit and implicit definitions can also be included below.
[0018] Data of users, acquisition and / or use of data, etc. can be involved in embodiments of the disclosure. These aspects all comply with corresponding laws and regulations and relevant provisions. In embodiments of the disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms. Accordingly, when implementing various embodiments of the disclosure, the type of data or information that can be involved, the scope of use, the use scenario, etc. should be notified to the user and authorized by the user in a proper manner according to relevant laws and regulations. The specific notification and / or authorization manner can vary according to the actual situation and application scenario, and the scope of the disclosure is not limited in this aspect.
[0019] In the specification and embodiments, if personal information processing is involved, it will be processed on the premise of legality (for example, obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and only within the prescribed or agreed range. The user refuses to process personal information other than the necessary information required for the basic function, which does not affect the user's use of the basic function.
[0020] As introduced above, when people use generative artificial intelligence technology to create media, accurate prompt items are a key factor affecting the generation quality of media content. Traditional solutions usually rely on user input text content as prompt items or simply specify preset templates. This limits the expression accuracy of prompt items and affects the generation quality of media content.
[0021] To this end, embodiments of the present disclosure propose a solution for generating media content. According to the solution, a set of candidate media samples corresponding to a set of preset prompt information can be provided in a target interface; based on selection of a target media sample in the set of candidate media samples, first prompt information corresponding to the target media sample is presented in the target interface; attribute information set of the target media sample with respect to a plurality of preset attribute items for describing different aspects of the media content to be generated is presented in the target interface; based on editing of first attribute information in the attribute information set, the first prompt information is modified to second prompt information corresponding to edited second attribute information; and based on the second prompt information, the target media content is generated.
[0022] In this way, embodiments of the present disclosure can support users to achieve refined editing of prompt information through preset media samples and corresponding attribute adjustments, thereby improving the accuracy of prompt information and the efficiency of editing prompt information, and further improving the quality of generated media content.
[0023] Various example implementations of the solution are described in further detail below in conjunction with the accompanying drawings.
[0024] Example Environment
[0025] FIG. 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG. 1, the example environment 100 can include an electronic device 110.
[0026] In this example environment 100, the electronic device 110 can run an application 120 that supports interface interaction. The application 120 can be any appropriate type of application for interface interaction, examples of which can include but are not limited to a video application, a clip application, or other appropriate application. A user 140 can interact with the application 120 via the electronic device 110 and / or its attached devices.
[0027] In the environment 100 of FIG. 1, if the application 120 is in an active state, the electronic device 110 can present an interface 150 for supporting interface interaction through the application 120.
[0028] In some embodiments, the electronic device 110 communicates with the server 130 to enable provisioning of services of the application 120. The electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a tablet computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a palmtop computer, a portable gaming terminal, a VR / AR device, a Personal Communication System (PCS) terminal, a personal navigation device, a Personal Digital Assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including an accessory or peripheral device for the foregoing, or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface to a user (such as “wearable” circuitry, etc.).
[0029] The server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and basic cloud computing services such as big data and artificial intelligence platforms. The server 130 may, for example, include a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like. The server 130 can provide background services for the application 120 in the electronic device 110 that supports a virtual scene.
[0030] A communication connection can be established between the server 130 and the electronic device 110. The communication connection can be established by wired or wireless means. The communication connection can include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, and the like, and embodiments of the present disclosure are not limited in this regard. In embodiments of the present disclosure, the server 130 and the electronic device 110 can implement signaling interaction through the communication connection therebetween.
[0031] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only, and do not imply any limitation on the scope of the present disclosure.
[0032] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.
[0033] Example Interaction
[0034] An example process of generating media content according to embodiments of the present disclosure will be described below with reference to the accompanying drawings. The example process of generating media content will be described below with video content as an example of media content. It should be understood that the process of generating according to the present disclosure can also be applied to other types of media content, e.g., images, audio, etc.
[0035] FIGS. 2A-2D illustrate example interfaces 200A-200D according to some embodiments of the present disclosure. The interfaces 200A-200D can be provided by the electronic device 110 shown in FIG. 1.
[0036] In some embodiments, the electronic device 110 presents a target interface based on a request of generating media content. FIG. 2A illustrates an example target interface according to some embodiments of the present disclosure. As shown in FIG. 2A, the target interface can include an input control 212 for inputting prompt information.
[0037] As shown in FIG. 2A, the user can input prompt information (e.g., a piece of description text) about the video content to be generated in the input control, i.e., the user’s description about the video content to be generated. The electronic device 110 can determine the video content to be generated based on the received prompt information.
[0038] In some embodiments, the electronic device 110 can also acquire at least one piece of material for generating the video content via the interface 200A. For example, the user can upload video material, image material, audio material, etc. based on the control 210. Alternatively, the user can also input script material through the control 212, e.g., as a subtitle of the script content to be generated, etc.
[0039] In some embodiments, the electronic device 110 can accept the user’s selection of the generation entry 216 and generate media content based on the prompt information and the material received in the interface 200A. As an example, the electronic device 110 can generate the corresponding media content, e.g., by utilizing generative artificial intelligence technology, without the present disclosure intending to limit the specific process of generating such media content.
[0040] In some embodiments, the electronic device 110 can also support the user to more efficiently edit the prompt information for generating media content. As shown in FIG. 2A, the electronic device 110 provides an entry 214 in association with the input control 212. Upon receiving a trigger for the entry 214, the electronic device 110 can present, e.g., the interface 200B as shown in FIG. 2B.
[0041] As shown in FIG. 2B, the electronic device 110 can provide a set of candidate media samples 222 in the interface 200B. In some embodiments, such candidate media samples can be associated with different preset prompt information. As an example, different candidate media samples can correspond to different media content styles.
[0042] In some embodiments, upon entering the interface 200B, the electronic device 110 can automatically recommend a target media sample in the set of candidate media samples 222, and can accordingly update the input control 220 based on the prompt information corresponding to the target candidate media sample.
[0043] For example, in a case where “media sample one” is determined as the target media sample, the electronic device 110 can display the preset prompt information “I want a song like music one, the theme is about XX, and the video duration is 3-5 seconds” corresponding to “media sample one” in the input control 220.
[0044] In some embodiments, the electronic device 110 can determine the target media sample to be recommended from the set of candidate media samples 222 based on the information obtained in the interface 200A.
[0045] As an example, the electronic device 110 can determine the target media sample to be recommended based on the analysis of at least one piece of material received in the interface 200A. Such material can include, for example, uploaded video material, picture material, music material, and text material to be used, etc.
[0046] In yet some embodiments, the electronic device 110 can also recommend the target media sample based on the text content that the user has input in the input control 212.
[0047] In some embodiments, the electronic device 110 can also determine the selected target media sample based on the user’s selection of the media sample, for example. In some examples, the electronic device 110 can not recommend any media sample, or recommend other media samples, when presenting the interface 200B, for example. Further, the electronic device 110 can receive the user’s selection of “media sample one”, and determine “media sample one” as the target media sample, for example.
[0048] In some embodiments, as shown in FIG. 2B, the electronic device 110 can also present a set of attribute information corresponding to a plurality of preset attribute items 224 (e.g., music, style, theme, video duration, etc.) in the interface 200B based on the target media sample (e.g., media sample one).
[0049] In some embodiments, taking video content as an example of the media content, a plurality of preset attribute items are used to describe different aspects of the video content to be generated. That is, the plurality of preset attribute items can correspond to a plurality of different video attributes of the video content to be generated, such as background music, video style, video theme, video duration, and the like.
[0050] In some embodiments, the electronic device 110 presents, in the target interface, a plurality of preset attribute values corresponding to one or more attribute items of the plurality of preset attribute items. As shown in FIG. 2B, taking the “music” attribute item as an example, the electronic device 110 can display, in the interface 200B, a plurality of preset attribute values corresponding to the “music” attribute item, such as “music one”, “music two”, and the like.
[0051] As an example, the music used by “media sample one” can be “music one”, for example. Accordingly, the attribute value “music one” can be displayed differently in the interface 200B. For example, the electronic device 110 can highlight the attribute value “music one” by bolding, highlighting, or the like, to represent that the music corresponding to “media sample one” is “music one”.
[0052] Similarly, the electronic device 110 can also provide a plurality of preset attribute values corresponding to other attribute items, and can display one or more attribute values of the plurality of preset attribute values corresponding to “media sample one” differently.
[0053] In some embodiments, the electronic device 110 can further receive an editing operation of the attribute information corresponding to the target media sample (e.g., media sample one) by the user. As an example, the electronic device 110 can receive an adjustment operation of the attribute value by the user, thereby completing the editing of the attribute information.
[0054] As shown in FIG. 2C, the electronic device 110 can modify the attribute information corresponding to the “music” attribute item from “music one” to “music two”, for example, by receiving a selection of the attribute value “music two” by the user.
[0055] Accordingly, the electronic device 110 can update the prompt information in the input control 220 based on the modified attribute information. As an example, the electronic device 110 can update the prompt information to “I want a song like music two, the theme is about XX, and the video duration is 3-5 seconds”.
[0056] In some embodiments, the editing of the attribute information can also include, for example, canceling the selection of one or more attribute values associated with the target media sample.
[0057] For example, the electronic device 110 can receive a user selection of the attribute value "music one" and deselect "music one". Accordingly, the electronic device 110 can delete the portion corresponding to "music one" from the preset prompt information corresponding to "media sample one". For example, the electronic device 110 can update the prompt information to "the theme is about XX, and the video duration is 3-5 seconds".
[0058] In some embodiments, the editing of the attribute information may, for example, further include selecting one or more attribute values that are not associated with the target media sample. For example, the target media sample can not include attribute information associated with the "style" attribute item. Further, the electronic device 110 may, for example, receive a user selection of "style one" in the "style" attribute item and can accordingly update the prompt information in the input control 220. For example, the electronic device 110 can update the prompt information to "I want a song like music one, the theme is about XX, the video duration is 3-5 seconds, and the video style is style one".
[0059] In this way, embodiments of the present disclosure can help users to more efficiently edit the prompt information by selecting media samples and modifying attributes, thereby improving the accuracy of the prompt information.
[0060] In some embodiments, the set of candidate media samples 222, the plurality of preset attribute items 224, and / or the plurality of preset attribute values 226 provided by the electronic device 110 can be determined based on a current creation scenario.
[0061] In some embodiments, the creation scenario can indicate a type or theme of the media content to be created. For example, the media samples, attribute items, attribute values provided for creating video content can be different from those provided for creating music content, and so on. For another example, the media samples, attribute items, attribute values provided for creating different types of video content (e.g., science popularization video content and advertising video content) can also be different.
[0062] In some embodiments, the creation scenario can indicate a creation link for generating the media content. For example, different creation links can correspond to different media samples, attribute items, attribute values, and so on. As an example, a script-based creation link can indicate matching corresponding materials based on script content to generate video content; a music effect-based creation link may, for example, indicate matching corresponding materials based on music effects (e.g., beats) to generate video content. In some examples, the script-based creation link and the music effect-based creation link can correspond to different media samples, attribute items, attribute values, and so on.
[0063] In some embodiments, to facilitate the user to better perceive the effect of each media sample, the electronic device 110 can further provide a preview interface corresponding to the candidate media sample. In some embodiments, the electronic device 110 can present the preview interface 200D as shown in FIG. 2D after receiving a preset operation (e.g., click, double-click, long press, etc.) on the candidate media sample 222.
[0064] As shown in FIG. 2D, the preview interface 200D can display the reference media content 230 and the reference prompt information 232 corresponding to the selected media sample (e.g., media sample one). Such reference media content 230 can be, for example, the media content generated by using the reference prompt information 232, to facilitate the perception of the generation result corresponding to the reference prompt information 232.
[0065] In some embodiments, the electronic device 110 can receive the user's selection of the use entry 234, and can return to the interface 200C and select "media sample one" as the target media sample, and accordingly update the input control 220 according to the reference prompt information 232. For example, the electronic device 110 can display the reference prompt information 232 in the input control 220.
[0066] In some embodiments, as shown in FIG. 2D, the electronic device 110 can receive the sliding operation 236 in the interface 200D, and can accordingly display the reference media content and the reference prompt information of another candidate media sample. In this way, the embodiments of the present disclosure can help the user to more effectively perceive the generation effect of each candidate media sample.
[0067] In some embodiments, after receiving the selection of the "confirm" button shown in FIG. 2B or FIG. 2C, the electronic device 110 can return to the target interface 200A as shown in FIG. 2A, and can accordingly display the prompt information determined based on FIG. 2B or FIG. 2C in the input control 212.
[0068] In some embodiments, the electronic device 110 can also support the user to edit the prompt information automatically generated by the electronic device 110 through the input control 212 or the input control 220.
[0069] Taking FIG. 2C as an example, the user can edit the prompt information generated by the electronic device 110 based on the selection of the attribute value "music two" in the input control 220, for example. Such editing can include but is not limited to: adding content, modifying content, deleting content, etc. For example, the electronic device 110 can receive the user's supplemented content "video style is style two".
[0070] As a further example, when the user returns to the interface 200A by, for example, clicking the confirmation button shown in FIG. 2C, the prompt information can be similarly edited by the input control 212 accordingly.
[0071] Further, the electronic device 110 can generate the target media content based on the prompt information (e.g., the description text about the media content to be generated) in the input control 212 based on the user’s triggering of the generation entry 218.
[0072] For example, taking the prompt information shown in the input control 220 of FIG. 2C as an example, the electronic device 110 may, for example, generate a video with “Music Two” as background music, related to the theme “XX”, and a duration of 3-5 seconds accordingly.
[0073] In this way, embodiments of the present disclosure support users to refine the editing of the prompt item through pre-set templates and corresponding attribute adjustments, thereby improving the accuracy of the prompt item and the efficiency of editing the prompt item, and further improving the quality of the generated media content.
[0074] Example process
[0075] FIG. 3 shows a flowchart of an example process 300 of generating media content, according to some embodiments of the present disclosure. The process 300 can be implemented at the electronic device 110. The process 300 is described below with reference to FIG. 1.
[0076] As shown, at block 310, the electronic device 110 provides a set of candidate media samples in a target interface, the set of candidate media samples corresponding to a set of pre-set prompt information.
[0077] At block 320, the electronic device 110 presents first prompt information corresponding to a target media sample in the target interface based on a selection of the target media sample in the set of candidate media samples.
[0078] At block 330, the electronic device 110 presents a set of attribute information of the target media sample about a plurality of pre-set attribute items in the target interface, the plurality of pre-set attribute items being used to describe different aspects of the media content to be generated.
[0079] At block 340, the electronic device 110 modifies the first prompt information to second prompt information corresponding to edited second attribute information based on an editing of first attribute information in the set of attribute information.
[0080] At block 350, the electronic device 110 generates target media content based on the second prompt information.
[0081] In some embodiments, presenting the set of attribute information in the target interface includes: presenting a plurality of preset attribute values corresponding to a target attribute item in the plurality of preset attribute items in the target interface; and differentially displaying a group of attribute values in the plurality of preset attribute values corresponding to the first attribute information.
[0082] In some embodiments, the editing of the first attribute information includes: canceling the selection of the first attribute value in the group of attribute values; or selecting a second attribute value in the plurality of preset attribute values different from the group of attribute values.
[0083] In some embodiments, modifying the first prompt information to the second prompt information corresponding to the second attribute information after the editing includes: deleting a first part in the first prompt information corresponding to the first attribute value; or adding a second part corresponding to the second attribute value to the first prompt information.
[0084] In some embodiments, generating the target media content based on the second prompt information includes: receiving an edit of the second prompt information to determine a third prompt information; and generating the target media content based on the third prompt information.
[0085] In some embodiments, the process 300 further includes: obtaining at least one media material for generating the target media content; and determining the target media sample from the group of candidate media samples based on the at least one media material.
[0086] In some embodiments, determining the target media sample from the group of candidate media samples based on the at least one media material includes: in response to obtaining fourth prompt information input by the user, determining the target media sample from the group of candidate media samples based on the at least one media material and the fourth prompt information.
[0087] In some embodiments, presenting the first prompt information corresponding to the target media sample in the group of candidate media samples in the target interface includes: based on the selection of the target media sample in the group of candidate targets, presenting the first prompt information corresponding to the target media sample in the target interface.
[0088] In some embodiments, the process 300 further includes: based on a first preset operation on the target media sample, presenting a preview interface, the preview interface displaying reference media content and reference prompt information corresponding to the target media sample; and based on the selection of a first entry in the preview interface, presenting the first prompt information corresponding to the target media sample in the target interface, the first prompt information corresponding to the reference prompt information.
[0089] In some embodiments, the process 300 further includes: based on a second preset operation received in the preview interface, presenting another reference media content and another reference prompt information corresponding to another media sample in the group of candidate media samples in the preview interface.
[0090] In some embodiments, providing the set of candidate media samples for controlling generation of the media content in the target interface includes: presenting the target interface based on the request for generating the media content, the target interface including an input control for inputting the prompt information; providing a second entry associated with the input control; and providing the set of candidate media samples in the target interface based on selection of the second entry.
[0091] In some embodiments, the set of candidate media samples and / or the plurality of preset attribute items are determined based on a creation scenario corresponding to the target interface.
[0092] In some embodiments, the first prompt information includes first text content corresponding to the first attribute information; or the second prompt information includes second text content corresponding to the second attribute information.
[0093] In some embodiments, the target media content includes video content, and the plurality of preset attribute items correspond to a plurality of different video attributes of the video content.
[0094] Example apparatuses and devices
[0095] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above method or process. FIG. 4 shows a schematic structural block diagram of an example media content generation apparatus 400 according to certain embodiments of the present disclosure. The apparatus 400 can be implemented as or included in the electronic device 110. Various modules / components in the apparatus 400 can be implemented by hardware, software, firmware, or any combination thereof.
[0096] As shown in FIG. 4, the apparatus 400 includes a sample providing module 410 configured to provide a set of candidate media samples in a target interface, the set of candidate media samples corresponding to a set of preset prompt information; a first presenting module 420 configured to present first prompt information corresponding to a target media sample in the set of candidate media samples in the target interface based on selection of the target media sample; a second presenting module 430 configured to present a set of attribute information of the target media sample with respect to a plurality of preset attribute items in the target interface, the plurality of preset attribute items for describing different aspects of the media content to be generated; an information processing module 440 configured to modify the first prompt information to second prompt information corresponding to edited second attribute information based on editing of first attribute information in the set of attribute information; and a content generation module 450 configured to generate the target media content based on the second prompt information.
[0097] In some embodiments, the second presenting module 430 is specifically configured to present a plurality of preset attribute values corresponding to a target attribute item in the plurality of preset attribute items in the target interface; and to distinctly display a set of attribute values in the plurality of preset attribute values corresponding to the first attribute information.
[0098] In some embodiments, the editing of the first attribute information comprises: canceling the selection of the first attribute value in the group of attribute values; or selecting a second attribute value in the plurality of preset attribute values, different from the group of attribute values.
[0099] In some embodiments, the information processing module 440 is specifically configured to delete a first part in the first prompt information corresponding to the first attribute value; or add a second part in the first prompt information corresponding to the second attribute value.
[0100] In some embodiments, the content generation module 450 is specifically configured to receive an edit of the second prompt information to determine a third prompt information; and generate the target media content based on the third prompt information.
[0101] In some embodiments, the apparatus 400 further comprises a material obtaining module configured to obtain at least one media material for generating the target media content; and determine the target media sample from the group of candidate media samples based on the at least one media material.
[0102] In some embodiments, the material obtaining module is specifically configured to, in response to obtaining a fourth prompt information input by the user, determine the target media sample from the group of candidate media samples based on the at least one media material and the fourth prompt information.
[0103] In some embodiments, the first presentation module 420 is specifically configured to present the first prompt information corresponding to the target media sample in the target interface based on the selection of the target media sample in the group of candidate targets.
[0104] In some embodiments, the apparatus 400 further comprises a third presentation module configured to present a preview interface based on the first preset operation on the target media sample, the preview interface displaying a reference media content and a reference prompt information corresponding to the target media sample; and present the first prompt information corresponding to the target media sample in the target interface based on the selection of a first entry in the preview interface, the first prompt information corresponding to the reference prompt information.
[0105] In some embodiments, the apparatus 400 further comprises a fourth presentation module configured to present another reference media content and another reference prompt information corresponding to another media sample in the group of candidate media samples in the preview interface based on a second preset operation received in the preview interface.
[0106] In some embodiments, the media sample providing module 410 is specifically configured to, based on the request of generating the media content, present a target interface, the target interface comprising an input control for inputting prompt information; provide a second entry associated with the input control; and based on selection of the second entry, provide a set of candidate media samples in the target interface.
[0107] In some embodiments, the set of candidate media samples and / or the plurality of preset attribute items are determined based on a creation scenario corresponding to the target interface.
[0108] In some embodiments, the first prompt information comprises first text content corresponding to the first attribute information; or the second prompt information comprises second text content corresponding to the second attribute information.
[0109] In some embodiments, the target media content comprises video content, and the plurality of preset attribute items correspond to a plurality of different video attributes of the video content.
[0110] The modules included in the apparatus 400 can be implemented utilizing a variety of means, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more of the units can be implemented using software and / or firmware, e.g., machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the modules in the apparatus 400 can be implemented at least partially by one or more hardware logic components. As an example and not by way of limitation, example types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.
[0111] FIG. 5 shows a block diagram of an electronic device 500 in which one or more embodiments of the disclosure can be implemented. It should be understood that the electronic device 500 illustrated in FIG. 5 is exemplary only and should not be taken as limiting the functionality and scope of the embodiments described herein. The electronic device 500 illustrated in FIG. 5 can be used to implement the electronic device 110 of FIG. 1.
[0112] As shown in FIG. 5, the electronic device 500 is in the form of a general electronic device. Components of the electronic device 500 can include, but are not limited to, one or more processors or processing units 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 can be a real or virtual processor and is capable of executing various processing according to programs stored in the memory 520. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the electronic device 500.
[0113] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 500.
[0114] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0115] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0116] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0117] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.
[0118] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0119] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0120] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0122] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive, nor is it limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for generating media content, comprising: providing a set of candidate media samples in a target interface, the set of candidate media samples corresponding to a set of preset prompt information; presenting, in the target interface, first prompt information corresponding to a target media sample in the set of candidate media samples based on a selection of the target media sample; presenting, in the target interface, a set of attribute information of the target media sample with respect to a plurality of preset attribute items for describing different aspects of the media content to be generated; modifying the first prompt information to second prompt information corresponding to edited second attribute information based on an edit of first attribute information in the set of attribute information; and generating target media content based on the second prompt information.
2. The method of claim 1, wherein presenting the set of attribute information in the target interface comprises: presenting a plurality of preset attribute values corresponding to a target attribute item in the plurality of preset attribute items in the target interface; and differentially displaying a set of attribute values in the plurality of preset attribute values corresponding to the first attribute information.
3. The method of claim 2, wherein the edit of the first attribute information comprises: de-selecting a first attribute value in the set of attribute values; or selecting a second attribute value in the plurality of preset attribute values different from the set of attribute values.
4. The method of claim 3, wherein modifying the first prompt information to second prompt information corresponding to edited second attribute information comprises: deleting a first portion in the first prompt information corresponding to the first attribute value; or adding a second portion to the first prompt information corresponding to the second attribute value.
5. The method of claim 1, wherein generating target media content based on the second prompt information comprises: receiving an edit of the second prompt information to determine third prompt information; and generating the target media content based on the third prompt information.
6. The method of claim 1, further comprising: acquiring at least one media material for generating the target media content; and determining the target media sample from the set of candidate media samples based on the at least one media material.
7. The method of claim 6, wherein determining the target media sample from the set of candidate media samples based on the at least one media material comprises: determining the target media sample from the set of candidate media samples based on the at least one media material and fourth prompt information responsive to acquiring the fourth prompt information input by a user.
8. The method of claim 1, wherein presenting, in the target interface, first prompt information corresponding to a target media sample in the set of candidate media samples comprises: presenting, in the target interface, first prompt information corresponding to the target media sample in the set of candidate media samples based on a selection of the target media sample.
9. The method of claim 8, further comprising: presenting, based on a first preset operation on the target media sample, a preview interface, the preview interface displaying reference media content and reference prompt information corresponding to the target media sample; and presenting, based on a selection of a first entry in the preview interface, the first prompt information corresponding to the target media sample in the target interface, the first prompt information corresponding to the reference prompt information.
10. The method of claim 9, further comprising: presenting, based on a second preset operation received in the preview interface, another reference media content and another reference prompt information corresponding to another media sample in the set of candidate media samples in the preview interface.
11. The method of claim 1, wherein providing, in a target interface, a set of candidate media samples for controlling generation of media content comprises: presenting, based on a request to generate media content, the target interface, the target interface including an input control for inputting prompt information; providing a second entry in association with the input control; and providing, based on a selection of the second entry, the set of candidate media samples in the target interface.
12. The method of claim 1, wherein the set of candidate media samples and / or the plurality of preset attribute items are determined based on a creation scenario corresponding to the target interface.
13. The method of claim 1, wherein: the first prompt information includes first textual content corresponding to the first attribute information; or the second prompt information includes second textual content corresponding to the second attribute information.
14. The method of claim 1, wherein the target media content includes video content, and the plurality of preset attribute items correspond to a plurality of different video attributes of video content.
15. An apparatus for generating media content, comprising: a sample providing module configured to provide, in a target interface, a set of candidate media samples corresponding to a set of preset prompt information; a first presenting module configured to present, based on a selection of a target media sample in the set of candidate media samples, first prompt information corresponding to the target media sample in the target interface; a second presenting module configured to present, in the target interface, a set of attribute information of the target media sample with respect to a plurality of preset attribute items for describing different aspects of media content to be generated; an information processing module configured to modify, based on an edit of first attribute information in the set of attribute information, the first prompt information to second prompt information corresponding to edited second attribute information; and a content generating module configured to generate target media content based on the second prompt information.
16. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method of any of claims 1-14. 17. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method of any one of claims 1 to 14.
Citation Information
Patent Citations
Image generation method and device, computer equipment and storage medium
CN116433825A
Story video generation method and device, storage medium and equipment
CN117332118A
Prompt text generation method and device, equipment and medium
CN117493013A
Multimedia data processing method and device, electronic equipment and storage medium
CN117726716A
Systems and Methods for Automated Generation of Video
US20200335132A1