Media content generation method and device, electronic equipment and storage medium

Through the two-stage generation model, combining the visual description information of the template media content and user-defined feature information, the media content generation process is optimized, and the problems of complex operation and poor effect in the existing technology are solved, and efficient and personalized media content creation is achieved.

CN120499162APending Publication Date: 2025-08-15BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510465038.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, the creation process of media content such as videos is complex, the creation efficiency is inefficient, and the generation effect is poor, especially when the uploaded media content is different from the layout of template media content.

Method used

The second media content is generated through the first preset generation model, combined with the preset visual description information of the target template media content and the user-defined media content feature information, and further optimized the generated media content using the second preset generation model to realize the construction of two-stage media content.

Benefits of technology

Without a large number of editing operations, ensure that the generated media content inherits the layout style of the template media content and incorporates user-defined personalized characteristics, improving creative efficiency and the presentation quality of media content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499162A_ABST
    Figure CN120499162A_ABST
Patent Text Reader

Abstract

The invention relates to a media content generation method and device, electronic equipment and a storage medium, and the method comprises the steps: in response to a first selection instruction of a target account for target template media content, displaying first media content and second media content customized by the target account on a preset page, the second media content is generated by the first preset generation model based on media content feature information corresponding to the first media content and preset visual description information corresponding to the target template media content; in response to the first content generation instruction, first target media content is displayed, and the first target media content is media content generated by the second preset generation model based on the second media content. According to the embodiment of the invention, the presentation effect and quality of the generated media content can be effectively ensured on the basis that a large amount of tedious editing processing operation is not needed, so that the operation convenience and creation efficiency in a media content creation process can be greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a method, device, electronic device, and storage medium for generating media content. Background Art

[0002] With the development of artificial intelligence (AI) technology, applications for creating media content such as videos that combine AI technology are becoming increasingly popular.

[0003] In the related art, when combining AI technology to create media content such as videos, users often first upload media content such as pictures or videos, and combine AI technology to generate media content that is the same as the template media content selected by the user; however, the related art often requires that the media content such as pictures or videos uploaded by users have a similar layout to the template media content, etc., which results in users often needing to edit and adjust the uploaded pictures or videos first, or, after generation, adjust the generated media content, resulting in complex operations and low creation efficiency in the media content creation process; if the uploaded media content and the template media content have significant differences in visual features such as layout, the final generated video and other media content often results in poor presentation effects. Therefore, in the related art, when combining artificial intelligence technology to create media content such as videos, there are often problems such as complex operations, low creation efficiency, and poor media content generation effects. Summary of the Invention

[0004] The present disclosure provides a method, device, electronic device, and storage medium for generating media content, aiming to at least address the technical issues in the related art of media content creation, such as complex operations, low creation efficiency, and poor presentation quality of the generated media content. The technical solutions of the present disclosure are as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, a method for generating media content is provided, including:

[0006] In response to a first selection instruction of a target account for target template media content, displaying first media content and second media content customized by the target account on a preset page, where the second media content is generated by a first preset generation model based on media content feature information corresponding to the first media content and preset visual description information corresponding to the target template media content;

[0007] In response to the first content generation instruction, first target media content is displayed, where the first target media content is media content generated by a second preset generation model based on the second media content.

[0008] In an optional embodiment, the method further includes:

[0009] In response to an update trigger instruction for the second media content, displaying the preset visual description information;

[0010] In response to a first update instruction for the preset visual description information, updating the preset visual description information to updated visual description information corresponding to the first update instruction;

[0011] In response to the second content generation instruction, display at least one updated media content corresponding to the second media content, where the at least one updated media content is generated by the first preset generation model based on the media content feature information and the updated visual description information;

[0012] The first target media content is replaced with the second target media content, and the displaying of the first target media content in response to the first content generation instruction includes:

[0013] In response to the first content generation instruction triggered by the target updated media content, presenting the second target media content;

[0014] The target updated media content is one or more updated media contents among the at least one updated media content, and the second target media content is media content generated by the second preset generation model based on the target updated media content.

[0015] In an optional embodiment, the first target media content is further replaced with third target media content, where the third target media content is media content generated by the second preset generation model based on the first media content and the target updated media content; and displaying the second target media content in response to the first content generation instruction triggered by the target updated media content includes:

[0016] In response to the first content generation instruction triggered based on the target updated media content, the third target media content is presented.

[0017] In an optional embodiment, in response to the update trigger instruction for the second media content, displaying the preset visual description information includes:

[0018] In response to the update trigger instruction, display at least one preset update control corresponding to the second media content, the at least one preset update control including a preset smart update control, the preset smart update control being used to trigger updating of the second media content generated by the first preset generation model by adjusting the preset visual description information;

[0019] In response to a second selection instruction triggered for the preset smart update control, the preset visual description information is displayed.

[0020] In an optional embodiment, the at least one preset update control further includes: at least one of a preset custom control and a preset recommendation control;

[0021] Among them, the preset custom control is used to trigger the update of the second media content through the media content customized by the target account; the preset recommendation control is used to trigger the update of the second media content through the media content recommended by the system.

[0022] In an optional embodiment, in response to the first content generation instruction triggered based on the target updated media content, presenting the second target media content includes:

[0023] In response to a third selection instruction for the target updated media content, updating the second media content displayed on the preset page to the target updated media content;

[0024] In response to the first content generation instruction triggered based on the preset page, the second target media content is displayed.

[0025] In an optional embodiment, the first target media content is replaced with fourth target media content, where the fourth target media content is media content generated by the second preset generation model based on the first media content and the second media content; and displaying the first target media content in response to the first content generation instruction includes:

[0026] In response to the first content generation instruction, the fourth target media content is presented.

[0027] In an optional embodiment, when the fourth target media content is a target video, the target template media content is a target template video, and the first media content includes a first frame image of the target video, the second media content is a last frame image of the target video, and the preset visual description information is image visual feature information of the last frame image in the target template video;

[0028] In a case where the fourth target media content is a target video, the target template media content is a target template video, and the first media content includes a last frame image of the target video, the second media content is a first frame image of the target video, and the preset visual description information is image visual feature information of the first frame image in the target template video;

[0029] The subject in the first media content and the subject in the second media content are both target objects.

[0030] In an optional embodiment, the first media content is based on the target object; the media content feature information is the object feature information of the target object; and the second media content is generated by the first preset generation model based on the object feature information and the preset visual description information, and is media content based on the target object.

[0031] In an optional embodiment, in response to the first selection instruction of the target account for the target template media content, displaying the first media content and the second media content customized by the target account on a preset page includes:

[0032] In response to the first selection instruction, displaying at least one preset content generation parameter corresponding to the first media content, the second media content, and the target template media content on the preset page;

[0033] The first target media content is replaced with the fifth target media content, and the displaying of the first target media content in response to the first content generation instruction includes:

[0034] In response to the first content generation instruction, the fifth target media content is displayed, where the fifth target media content is media content generated by the second preset generation model based on the second media content and the at least one preset content generation parameter.

[0035] In an optional embodiment, the method further includes:

[0036] In response to a second update instruction for the target content generation parameter, updating the target content generation parameter among the at least one preset content generation parameter displayed on the preset page to an updated content generation parameter corresponding to the second update instruction, the target content generation parameter being any one of the at least one preset content generation parameter;

[0037] The first target media content is replaced with the sixth target media content, and the displaying of the first target media content in response to the first content generation instruction includes:

[0038] In response to the first content generation instruction, display the sixth target media content, where the sixth target media content is media content generated by the second preset generation model based on the second media content and the at least one updated content generation parameter.

[0039] The at least one updated content generation parameter is a content generation parameter obtained by updating the target content generation parameter in the at least one preset content generation parameter based on the updated content generation parameter.

[0040] According to a second aspect of an embodiment of the present disclosure, there is provided a media content generating apparatus, including:

[0041] The first content display module is configured to execute, in response to a first selection instruction of a target account for target template media content, displaying first media content and second media content customized by the target account on a preset page, where the second media content is generated by a first preset generation model based on media content feature information corresponding to the first media content and preset visual description information corresponding to the target template media content;

[0042] The second content display module is configured to execute in response to the first content generation instruction and display the first target media content, where the first target media content is media content generated by a second preset generation model based on the second media content.

[0043] In an optional embodiment, the device further comprises:

[0044] a visual description information display module, configured to execute, in response to an update trigger instruction for the second media content, display the preset visual description information;

[0045] a visual description information updating module, configured to execute, in response to a first updating instruction for the preset visual description information, updating the preset visual description information to updated visual description information corresponding to the first updating instruction;

[0046] an updated media content display module, configured to execute, in response to a second content generation instruction, display at least one updated media content corresponding to the second media content, the at least one updated media content being generated by the first preset generation model based on the media content feature information and the updated visual description information;

[0047] The first target media content is replaced with the second target media content, and the second content display module includes:

[0048] a first content display unit configured to execute the first content generation instruction triggered in response to the target updated media content, and display the second target media content;

[0049] The target updated media content is one or more updated media contents among the at least one updated media content, and the second target media content is media content generated by the second preset generation model based on the target updated media content.

[0050] In an optional embodiment, the first target media content is further replaced with third target media content, where the third target media content is media content generated by the second preset generation model based on the first media content and the target updated media content; and the first content display unit includes:

[0051] The second content display unit is configured to execute the first content generation instruction in response to being triggered based on the target updated media content, and display the third target media content.

[0052] In an optional embodiment, the visual description information display module includes:

[0053] a preset update control display unit, configured to execute, in response to the update trigger instruction, display at least one preset update control corresponding to the second media content, the at least one preset update control including a preset smart update control, the preset smart update control being used to trigger an update of the second media content generated by the first preset generation model by adjusting the preset visual description information;

[0054] The visual description information display unit is configured to execute a second selection instruction in response to the second selection instruction triggered for the preset intelligent update control to display the preset visual description information.

[0055] In an optional embodiment, the at least one preset update control further includes: at least one of a preset custom control and a preset recommendation control;

[0056] Among them, the preset custom control is used to trigger the update of the second media content through the media content customized by the target account; the preset recommendation control is used to trigger the update of the second media content through the media content recommended by the system.

[0057] In an optional embodiment, the first content display unit includes:

[0058] a third content display unit configured to execute, in response to a third selection instruction for the target updated media content, updating the second media content displayed on the preset page to the target updated media content;

[0059] The fourth content display unit is configured to execute the first content generation instruction in response to being triggered based on the preset page, and display the second target media content.

[0060] In an optional embodiment, the first target media content is replaced with fourth target media content, where the fourth target media content is media content generated by the second preset generation model based on the first media content and the second media content; and the second content display module includes:

[0061] The fifth content display unit is configured to display the fourth target media content in response to the first content generation instruction.

[0062] In an optional embodiment, when the fourth target media content is a target video, the target template media content is a target template video, and the first media content includes a first frame image of the target video, the second media content is a last frame image of the target video, and the preset visual description information is image visual feature information of the last frame image in the target template video;

[0063] In a case where the fourth target media content is a target video, the target template media content is a target template video, and the first media content includes a last frame image of the target video, the second media content is a first frame image of the target video, and the preset visual description information is image visual feature information of the first frame image in the target template video;

[0064] The subject in the first media content and the subject in the second media content are both target objects.

[0065] In an optional embodiment, the first media content is based on the target object; the media content feature information is the object feature information of the target object; and the second media content is generated by the first preset generation model based on the object feature information and the preset visual description information, and is media content based on the target object.

[0066] In an optional embodiment, the first content display module includes:

[0067] a sixth content display unit, configured to display, in response to the first selection instruction, at least one preset content generation parameter corresponding to the first media content, the second media content, and the target template media content on the preset page;

[0068] The first target media content is replaced with fifth target media content, and the second content display module includes:

[0069] The seventh content display unit is configured to execute in response to the first content generation instruction and display the fifth target media content, where the fifth target media content is media content generated by the second preset generation model based on the second media content and the at least one preset content generation parameter.

[0070] In an optional embodiment, the device further comprises:

[0071] a content generation parameter updating module configured to execute, in response to a second update instruction for a target content generation parameter, an update of the target content generation parameter among the at least one preset content generation parameter displayed on the preset page to an updated content generation parameter corresponding to the second update instruction, wherein the target content generation parameter is any one of the at least one preset content generation parameter;

[0072] The first target media content is replaced with sixth target media content, and the second content display module includes:

[0073] an eighth content display unit, configured to execute, in response to the first content generation instruction, display the sixth target media content, the sixth target media content being media content generated by the second preset generation model based on the second media content and the at least one updated content generation parameter;

[0074] The at least one updated content generation parameter is a content generation parameter obtained by updating the target content generation parameter in the at least one preset content generation parameter based on the updated content generation parameter.

[0075] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement a method as described in any one of the above-mentioned media content generation methods.

[0076] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute any one of the media content generation methods of the embodiments of the present disclosure.

[0077] According to a fifth aspect of an embodiment of the present disclosure, there is provided a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute any one of the above-mentioned media content generation methods.

[0078] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects:

[0079] During the media content generation process, the first preset generation model and the second preset generation model can be combined to realize two-stage media content construction. In the first stage, in response to the first selection instruction of the target account for the target template media content, the first media content customized by the target account and the second media content generated by the first preset generation model based on the media content feature information corresponding to the first media content and the preset visual description information corresponding to the target template media content are displayed on the preset page. While inheriting the visual features such as the layout style in the target template media content, the user-defined personalized media content features can be integrated. In the second stage, in response to the first content generation instruction, the first target media content generated by the second preset generation model based on the second media content is displayed. The second media content that inherits the visual features such as the layout style in the target template media content and integrates the user-defined personalized media content features can be combined. Without the need for a large number of tedious editing and processing operations, the presentation effect and quality of the generated media content can be effectively guaranteed, thereby greatly improving the convenience of operation and the creation efficiency in the media content creation process.

[0080] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0082] Figure 1 is a schematic diagram showing an application environment according to an exemplary embodiment;

[0083] Figure 2 is a flow chart showing a method for generating media content according to an exemplary embodiment;

[0084] Figure 3 is a schematic diagram showing page changes in a process of triggering the generation of second media content according to an exemplary embodiment;

[0085] Figure 4 is a schematic diagram showing a preset page displaying multiple preset update controls according to an exemplary embodiment;

[0086] Figure 5 is a schematic diagram of page changes during a process of updating second media content by updating preset visual description information according to an exemplary embodiment;

[0087] Figure 6 is a block diagram of a media content generation device according to an exemplary embodiment;

[0088] Figure 7 It is a block diagram of an electronic device for generating media content according to an exemplary embodiment. DETAILED DESCRIPTION

[0089] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0090] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0091] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.

[0092] See also Figure 1 , Figure 1 FIG. 1 is a schematic diagram showing an application environment according to an exemplary embodiment. The application environment may include a terminal 100 and a server 200 .

[0093] In an optional embodiment, the terminal 100 can be used to provide services such as media content generation to any user. Specifically, the terminal 100 may include, but is not limited to, electronic devices such as smartphones, desktop computers, tablet computers, laptop computers, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. It may also be software running on these electronic devices, such as applications. Optionally, the operating system running on the electronic device may include, but is not limited to, Android, iOS, Linux, Windows, etc.

[0094] In an optional embodiment, the server 200 may provide background services for the terminal 100. Specifically, the server 200 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0095] In addition, it should be noted that Figure 1 What is shown is only one application environment provided by the present disclosure. In actual applications, other application environments may also be included, for example, more terminals may be included.

[0096] In the embodiments of this specification, the terminal 100 and the server 200 may be directly or indirectly connected via wired or wireless communication, which is not limited in this disclosure.

[0097] Figure 2 is a flow chart of a method for generating media content according to an exemplary embodiment. The method can be applied to a terminal, such as Figure 2 As shown, the method may include the following steps:

[0098] In step S201, in response to the first selection instruction of the target account for the target template media content, the first media content and the second media content customized by the target account are displayed on the preset page. The second media content is generated by the first preset generation model based on the media content feature information corresponding to the first media content and the preset visual description information corresponding to the target template media content.

[0099] In a specific embodiment, the target account can be the currently logged-in user account; the first selection instruction for the target template media content can be an instruction to select the target template media content. Specifically, the target template media content can be the template media content selected by the target account; specifically, the first media content can be the media content customized by the target account; specifically, the media content in the embodiment of the present application can include but is not limited to at least one modality of content such as images, text, graphics, and videos; illustratively, the first media content can include at least any one of an image customized by the target account, text information customized by the target account, an image generated based on the customized image generation description information (information describing the image to be generated), and a customized video.

[0100] In a specific embodiment, the first preset generation model may be an artificial intelligence model for generating media content. For example, when the second media content is an image, the first preset generation model may be an artificial intelligence model for generating images; for example, when the second media content is a video, the first preset generation model may be an artificial intelligence model for generating videos. Specifically, the preset visual description information may be information reflecting the visual features of the target template media content as a whole or in part (such as a frame image in a video); for example, the preset visual description information may include at least any one of style information and content description information. Specifically, the content description information may include information describing at least any one of the color, light and shadow, texture, layout, transition method, etc. of the corresponding media content.

[0101] In a specific embodiment, the media content feature information may be feature information representing the first media content. Specifically, the media content feature information corresponding to the first media content may be extracted in combination with a feature extraction model. In an optional embodiment, the first media content may be based on a target object. Accordingly, the media content feature information may be object feature information of the target object. Accordingly, the second media content may be media content based on the target object, generated by a first preset generation model based on object feature information and preset visual description information.

[0102] In a specific embodiment, the target object can be set by the user based on actual needs. For example, the target object can be a person, an animal, or other object. Accordingly, the feature extraction model can be a model for extracting object feature information corresponding to the target object. Specifically, the target object can be first identified from the first media content, and then feature extraction processing can be performed on the target object. Optionally, the object feature information corresponding to the target object can be feature information that characterizes the target object's position, posture, expression, and other characteristics of the object.

[0103] In the above embodiment, when the first media content is based on the target object as the main body, the second media content is generated by the first preset generation model based on the object feature information corresponding to the target object in the first media content and the preset visual description information corresponding to the target template media content. The media content with the target object as the main body can effectively ensure the consistency between the main body in the generated media content and the user-defined main body while ensuring that the generated second media content inherits the visual features in the template. Furthermore, in the same media content generation scenario, the template can be combined to effectively improve the presentation effect of the generated media content, while improving the generation efficiency and convenience, and improving the user's personalized needs.

[0104] In an optional embodiment, the target template media content may be selected first, and then the customized first media content may be input. Accordingly, in response to the target account's first selection instruction for the target template media content, displaying the target account's customized first media content and second media content on a preset page may include:

[0105] In response to a first selection instruction of the target account for the target template media content, an input control corresponding to the first media content is displayed on a preset page; in response to an input instruction of the first media content triggered based on the input control, the first media content and the second media content are displayed on the preset page.

[0106] In a specific embodiment, the input control can be used directly to trigger the input of the first media content, or the user can trigger the input of content generation description information (information describing the first media content to be generated, such as descriptive text corresponding to the image to be generated in the text-image scene) used to generate the first media content.

[0107] In another optional embodiment, the target account can first input the customized first media content and then select the target template media content; optionally, in the media content generation process, after inputting the customized first media content, the display of at least one preset template media content can be triggered first, and by selecting a preset template media content (target template media content), the first selection instruction for the target template media content is triggered, and a preset page displaying the first media content and the second media content is displayed.

[0108] In actual applications, after the first selection instruction is triggered, it often takes a certain amount of time for the first preset generation model to generate the second media content based on the media content feature information corresponding to the first media content and the preset visual description information corresponding to the target template media content; optionally, a preset page showing the corresponding generation prompt information of the first media content and the second media content (i.e., the prompt information that the second media content is being generated) can be displayed first. For example, Figure 3 As shown, Figure 3 FIG. 1 is a schematic diagram showing page changes during a process of triggering the generation of second media content according to an exemplary embodiment. Figure 3 The page shown in a displays description information of multiple preset template media contents (such as cover images and creative theme information). Furthermore, by clicking a touch operation such as a "generate the same" control corresponding to a preset template media content, a first selection instruction for the preset template media content (target template media content) can be triggered; further, as Figure 3 The preset page shown in b can display the generation prompt information 302 corresponding to the first media content 301 and the second media content; further, as shown in FIG. Figure 3In the preset page shown in c, when the second media content is generated, the generation prompt information 302 corresponding to the second media content in the preset page can be replaced with the second media content 303.

[0109] In step S203 , in response to the first content generation instruction, first target media content is displayed, where the first target media content is media content generated by the second preset generation model based on the second media content.

[0110] In a specific embodiment, the first content generation instruction may be an instruction for generating media content; optionally, the preset page may further display a media content generation control, for example, Figure 3 Correspondingly, the first content generation instruction can be triggered by a touch operation such as clicking or long pressing the media content generation control. Optionally, the first target media content can be displayed on a new page or on a preset page through a pop-up window or other means.

[0111] In a specific embodiment, the second preset generation model can be an artificial intelligence model for generating media content. Optionally, taking the case where a template video (target template media content) corresponding to the same video (first target media content) is generated by combining a user-defined object image (first media content), and the second media content is at least one key frame image in the same video as an example, the second media content can be an image whose main body is the main object in the first media content, but retains the image visual feature information corresponding to the key frame image in the template video; accordingly, the first target media content can be the same video as the template video whose main body is the main object in the first media content. Optionally, taking the case where a template video (target template media content) corresponding to the same video (first target media content) is generated by combining user-defined text information (first media content, for example, a character name), and the second media content is at least one key frame image in the same video as an example, the second media content can be an image whose main body is the character corresponding to the first media content, but retains the image visual feature information corresponding to the key frame image in the template video; accordingly, the first target media content can be the same video as the template video whose main body is the character corresponding to the first media content.

[0112] In addition, it should be noted that in actual applications, after triggering the first content generation instruction, it often takes a certain amount of time to generate the first target media content. Optionally, before displaying the first target media content, the corresponding generation prompt information (i.e., the prompt information that the first target media content is being generated) can be displayed first. Optionally, the corresponding generation progress information can also be displayed.

[0113] In an optional embodiment, the first target media content may be replaced with fourth target media content, where the fourth target media content may be media content generated by a second preset generation model based on the first media content and the second media content. Accordingly, in response to the first content generation instruction, displaying the first target media content includes:

[0114] In response to the first content generation instruction, fourth target media content is presented.

[0115] In a specific embodiment, taking the scenario of generating a video by combining the first and last frame images as an example, when the fourth target media content is the target video, and the target template media content is the target template video, and the first media content includes the first frame image of the target video, the above-mentioned second media content is the last frame image of the target video, and the preset visual description information is the image visual feature information of the last frame image in the target template video; optionally, when the fourth target media content is the target video, and the target template media content is the target template video, and the first media content includes the last frame image of the target video, the above-mentioned second media content can be the first frame image of the target video, and the above-mentioned preset visual description information can be the image visual feature information of the first frame image in the target template video.

[0116] In a specific embodiment, the subject in the first media content and the subject in the second media content are both target objects.

[0117] In a specific embodiment, suppose a special effects video is created by combining first and last frame images. The first frame image of the special effects template is "a dog running on the grass," and the last frame image is "a dog flying into the sky." If a user uploads an image of themselves at the beach as the first frame image, and the last frame image in the special effects template is directly used as the last frame image for the special effects video created by the user, the resulting special effects video will show an abrupt transition between the user at the beach (first frame) and the dog flying into the sky (last frame). Based on the technical solution provided in the embodiments of the present application, the user's personal feature information can be extracted from the above-mentioned image of the user at the beach, and combined with the last frame image in the template special effects, a last frame image (second media content) with similar visual feature information such as style to the last frame of the special effects template can be automatically generated, but with the main subject being "the user himself." Furthermore, the generated last frame image can be combined with the first frame image uploaded by the user to generate the same special effects video with the same theme.

[0118] In the above embodiment, in the scenario of generating a video by combining the first and last frame images, in the process of generating the second media content, the object feature information corresponding to the target object in the first media content and the preset visual description information corresponding to the last frame image in the target template media content are used to generate the second media content. While effectively ensuring the video generation effect by combining the preset visual description information corresponding to the last frame image in the template media content, the consistency of the subject in the first and last frame images can be effectively guaranteed, thereby effectively avoiding the abrupt video connection and frame skipping caused by the inconsistency of the subject in the first and last frame images, greatly improving the smoothness of the generated video and the user's subsequent video viewing experience.

[0119] In addition, in the scenario where the same video of the template video is generated in combination with the first and last frame images, the first media content can also include content generation description information corresponding to the second media content; accordingly, the above-mentioned media content feature information can be feature information of the content generation description information in the first media content, and accordingly, the above-mentioned second preset generation model can generate the fourth target media content based on the first frame image or the last frame image in the first media content, and the second media content.

[0120] In the above embodiment, the second preset generation model generates the final media content based on the user-defined first media content and the intelligently generated second media content, which can better ensure the personalized characteristics of the final generated media content and thus better meet the user's personalized creation needs.

[0121] In an optional embodiment, the above method may further include:

[0122] In response to an update trigger instruction for the second media content, displaying preset visual description information;

[0123] In response to a first update instruction for the preset visual description information, updating the preset visual description information to updated visual description information corresponding to the first update instruction;

[0124] In response to the second content generation instruction, presenting the second media content corresponding to at least one updated media content;

[0125] Accordingly, the first target media content may be replaced by the second target media content. The displaying of the first target media content in response to the first content generation instruction includes:

[0126] In response to a first content generation instruction triggered based on the target updated media content, presenting second target media content;

[0127] In an optional embodiment, in response to the update trigger instruction for the second media content, displaying the preset visual description information may include:

[0128] In response to the update trigger instruction, display at least one preset update control corresponding to the second media content, the at least one preset update control including a preset smart update control;

[0129] In response to a second selection instruction triggered for the preset smart update control, preset visual description information is displayed.

[0130] In a specific embodiment, the update trigger instruction may be an instruction for triggering the update of the second media content. Optionally, the update trigger instruction may be triggered by a touch operation such as clicking or long pressing a preset update trigger control corresponding to the second media content. For example, Figure 3 The control corresponding to 305 may be a preset update trigger control.

[0131] In a specific embodiment, at least one preset update control can be displayed on a preset page via a pop-up window or the like, or on a new page. Specifically, the preset intelligent update control can be used to trigger an update of the second media content generated by the first preset generation model by adjusting the preset visual description information.

[0132] In an optional embodiment, the second selection instruction may be an instruction for selecting a preset smart update control; optionally, the second selection instruction for the preset smart update control may be triggered by a touch operation such as clicking or long pressing the preset smart update control. Specifically, the preset visual description information may be displayed on a preset page through a pop-up window or the like, or may be displayed on a new page. Exemplarily, the preset visual description information may include style information and content description information; optionally, the style information may be presented in the form of an image and / or text of a corresponding style; specifically, the division of styles may be set in combination with actual applications, such as cartoon style, oil painting style, realistic style, etc. The content description information may be presented in the form of an image and / or text.

[0133] In the above embodiment, in response to an update trigger instruction of the second media content, at least one preset update control corresponding to the second media content is displayed, and when a second selection instruction for a preset intelligent update control in at least one preset update control is triggered, preset visual description information can be displayed, thereby facilitating the user to dynamically update the intelligently generated second media content by adjusting the preset visual description information, thereby better improving the matching between the visual presentation effect of the final generated video and other media content and user needs, and thereby better improving the user's personalized creation needs.

[0134] In an optional embodiment, the at least one preset update control may further include: at least one of a preset custom control and a preset recommendation control;

[0135] In a specific embodiment, the preset custom control can be used to trigger the target account to customize the media content to update the second media content. Specifically, the target account customized media content can be media content input by the target account, such as input through album upload or real-time shooting. The target account customized media content can also be media content generated based on the description information of the content input by the target account, such as input through text and image.

[0136] In a specific embodiment, the preset recommendation control can be used to trigger the updating of the second media content through the media content recommended by the system. Specifically, the media content recommended by the system can be the media content preset by the system and recommended to the user for selection.

[0137] In a specific embodiment, Figure 4 As shown, Figure 4 This is a page diagram of a preset page showing multiple preset update controls according to an exemplary embodiment; wherein the control corresponding to 401 may be a preset smart update control, the control corresponding to 402 may be a preset custom control, and the control corresponding to 403 may be a preset recommendation control.

[0138] In the above embodiment, on the basis of setting a preset intelligent update control, at least one of a preset custom control and a preset recommendation control is also set. In the process of dynamically updating the intelligently generated media content, a rich update method can be provided to the user, thereby greatly improving the user's flexibility in dynamic updating.

[0139] In a specific embodiment, the first update instruction may be an instruction to update the preset visual description information; optionally, the preset visual description information may be updated by selecting other visual description information, such as selecting other style information; or the preset visual description information may be updated by directly modifying a certain visual description information, such as re-editing the content description information.

[0140] In a specific embodiment, the second content generation instruction may be an instruction to generate at least one updated media content; optionally, while displaying the preset visual description information, the corresponding generation control may also be displayed. Accordingly, the second content generation instruction may be triggered by clicking, long pressing, or the like on the generation control. Specifically, the at least one updated media content may be displayed on a preset page via a pop-up window or the like, or may be displayed on a new page. Specifically, the at least one updated media content may be generated by the first preset generation model based on media content feature information and updated visual description information; optionally, the at least one updated media content may also be generated by the first preset generation model based on object feature information and updated visual description information corresponding to the target object in the first media content. The target updated media content may be one or more updated media contents in the at least one updated media content, and the second target media content may be media content generated by the second preset generation model based on the target updated media content.

[0141] In an optional embodiment, in response to the first content generation instruction triggered based on the target updated media content, presenting the second target media content may include:

[0142] In response to a third selection instruction for target updated media content, updating the second media content displayed on the preset page to the target updated media content;

[0143] In response to the first content generation instruction triggered based on the preset page, the second target media content is displayed.

[0144] In a specific embodiment, the third selection instruction can be an instruction for selecting the target updated media content, and accordingly, the second media content displayed on the preset page can be updated to the target updated media content; further, when the first content generation instruction is triggered, the second target media content generated by the second preset generation model based on the target updated media content can be displayed.

[0145] In the above embodiment, in response to the third selection instruction for the target updated media content, the second media content displayed on the preset page is updated to the target updated media content, and when the target updated media content is displayed on the preset page, the first content generation instruction is triggered, so that the final second target media content can be generated based on the target updated media content, thereby better improving the matching between the visual presentation effect of the final generated video and other media content and user needs.

[0146] In a specific embodiment, Figure 5 As shown, Figure 5 1 is a schematic diagram of page changes during the process of updating the second media content by updating the preset visual description information according to an exemplary embodiment. Specifically, Figure 5The page shown in a shows style information and content description information (preset visual description information); optionally, the user can update the style information and / or content description information according to actual needs. Further, the second content generation instruction can be triggered in combination with the "Generate Now" control. Accordingly, Figure 5 As shown in b, the second media content can be displayed corresponding to multiple updated media contents; further, the user can select a certain updated media content based on actual needs to update the second media content in the preset page to the selected updated media content, and then generate and create media content based on the selected updated media content.

[0147] In addition, it should be noted that in actual applications, after triggering the second content generation instruction, it often takes a certain amount of time to generate at least one updated media content. Optionally, before displaying at least one updated media content, the corresponding generation prompt information (i.e., the prompt information that the updated media content is being generated) can be displayed first. Optionally, the corresponding generation progress information can also be displayed.

[0148] In the above embodiment, in response to an update trigger instruction for the second media content, preset visual description information is displayed, and then the second media content corresponding to at least one updated media content can be generated by updating the preset visual description information. Then, the final target media content can be generated in combination with the target updated media content in the at least one updated media content. By dynamically adjusting the intelligently generated second media content, the quality of the media content can be effectively guaranteed while better meeting the user's personalized creation needs, thereby attracting more users to participate in the creation and improving the richness and diversity of the platform content.

[0149] In an optional embodiment, the first target media content may be replaced by a third target media content. Accordingly, in response to the first content generation instruction triggered by the target updated media content, displaying the second target media content may include:

[0150] In response to the first content generation instruction triggered based on the target updated media content, third target media content is presented.

[0151] In a specific embodiment, the third target media content is media content generated by the second preset generation model based on the first media content and the target updated media content.

[0152] In an optional embodiment, in a scenario where a video is generated by combining the first and last frame images, the second preset generation model can generate third target media content based on the first frame image or the last frame image in the first media content and the target updated media content.

[0153] In the above embodiment, after the intelligently generated second media content is updated, when the target updated media content is determined, the final media content can also be generated in combination with the user-defined first media content, which can better enhance the personalized characteristics of the final generated media content and thus better meet the user's personalized creation needs.

[0154] In an optional embodiment, in response to the target account's first selection instruction for the target template media content, displaying the target account's customized first media content and second media content on a preset page may include:

[0155] In response to the first selection instruction, displaying at least one preset content generation parameter corresponding to the first media content, the second media content, and the target template media content on a preset page;

[0156] Accordingly, the first target media content may be replaced by the fifth target media content. The displaying of the first target media content in response to the first content generation instruction includes:

[0157] In response to the first content generation instruction, fifth target media content is displayed, where the fifth target media content is media content generated by the second preset generation model based on the second media content and at least one preset content generation parameter.

[0158] In a specific embodiment, the at least one preset content generation parameter may be at least one parameter that can characterize the characteristics of the finally generated media content (fifth target media content); optionally, for example, at least one of video transition description information, camera movement mode, and the correlation between the fifth target media content and the target simulation media content. Exemplarily, the at least one preset content generation parameter may include Figure 3 The data corresponding to 306.

[0159] In addition, it should be noted that, during the generation process of the second target media content, the third target media content, and the fourth target media content, at least one preset content generation parameter can also be used for generation.

[0160] In the above embodiment, in the process of intelligently generating media content, combining at least one preset content generation parameter corresponding to the target template media content can better improve the correlation between the finally generated media content and the target template media content, thereby better improving the generation effect of the same media content.

[0161] In an optional embodiment, the above method may further include:

[0162] In response to a second update instruction for the target content generation parameter, updating the target content generation parameter in the at least one preset content generation parameter displayed on the preset page to the updated content generation parameter corresponding to the second update instruction, where the target content generation parameter is any one of the at least one preset content generation parameter;

[0163] Accordingly, the first target media content is replaced with the sixth target media content. In response to the first content generation instruction, displaying the first target media content includes:

[0164] In response to the first content generation instruction, sixth target media content is displayed, where the sixth target media content is media content generated by the second preset generation model based on the second media content and at least one updated content generation parameter.

[0165] In a specific embodiment, the second update instruction may be an instruction for updating target content generation parameters. Optionally, the target content generation parameters may be updated by selecting other content generation parameters or by directly modifying the target content generation parameters.

[0166] In a specific embodiment, the at least one updated content generation parameter may be a content generation parameter obtained by updating a target content generation parameter in at least one preset content generation parameter based on the updated content generation parameter.

[0167] In addition, it should be noted that, during the generation process of the second target media content, the third target media content, and the fourth target media content, at least one preset content generation parameter can also be used for generation.

[0168] In the above embodiment, the user can adjust and update at least one preset content generation data based on actual needs, thereby better meeting the user's personalized creation needs on the basis of improving the generation effect of the same media content.

[0169] In addition, it should be noted that in the embodiment of the present application, the instructions triggered based on touch operations can also be triggered by voice control operations, device movement operations, etc. in actual applications.

[0170] It can be seen from the technical solutions provided by the above embodiments of this specification that in the process of media content generation in this specification, the first preset generation model and the second preset generation model can be combined to realize two-stage media content construction, and in the first stage, in response to the first selection instruction of the target account for the target template media content, the first media content customized by the target account and the second media content generated by the first preset generation model based on the media content feature information corresponding to the first media content and the preset visual description information corresponding to the target template media content are displayed on the preset page. While inheriting the visual features such as the layout style in the target template media content, the user-defined personalized media content features can be integrated. In the second stage, in response to the first content generation instruction, the first target media content generated by the second preset generation model based on the second media content is displayed. The second media content can be combined with the visual features such as the layout style inherited from the target template media content and integrated with the user-defined personalized media content features. Without the need for a large number of tedious editing and processing operations, the presentation effect and quality of the generated media content are effectively guaranteed, thereby greatly improving the convenience of operation and the creation efficiency in the media content creation process.

[0171] Figure 6 FIG. 1 is a block diagram of a media content generation device according to an exemplary embodiment. Figure 6 , the device comprises:

[0172] The first content display module 610 is configured to execute, in response to a first selection instruction of the target account for the target template media content, displaying first media content and second media content customized by the target account on a preset page, where the second media content is generated by a first preset generation model based on media content feature information corresponding to the first media content and preset visual description information corresponding to the target template media content;

[0173] The second content display module 620 is configured to execute in response to the first content generation instruction and display the first target media content, where the first target media content is media content generated by the second preset generation model based on the second media content.

[0174] In an optional embodiment, the above device further includes:

[0175] a visual description information display module, configured to execute, in response to an update trigger instruction for the second media content, display preset visual description information;

[0176] a visual description information updating module configured to execute, in response to a first updating instruction for the preset visual description information, updating the preset visual description information to updated visual description information corresponding to the first updating instruction;

[0177] an updated media content display module, configured to execute, in response to the second content generation instruction, displaying at least one updated media content corresponding to the second media content, the at least one updated media content being generated by the first preset generation model based on the media content feature information and the updated visual description information;

[0178] The first target media content is replaced with the second target media content, and the second content display module 620 includes:

[0179] a first content display unit configured to execute a first content generation instruction in response to being triggered based on the target updated media content, and to display the second target media content;

[0180] The target updated media content is one or more updated media contents in the at least one updated media content, and the second target media content is media content generated by the second preset generation model based on the target updated media content.

[0181] In an optional embodiment, the first target media content is further replaced by third target media content, where the third target media content is media content generated by the second preset generation model based on the first media content and the target updated media content; and the first content display unit includes:

[0182] The second content display unit is configured to execute the first content generation instruction triggered in response to the target updated media content, and display the third target media content.

[0183] In an optional embodiment, the visual description information display module includes:

[0184] a preset update control display unit configured to execute, in response to an update trigger instruction, display at least one preset update control corresponding to the second media content, the at least one preset update control including a preset smart update control, the preset smart update control being used to trigger an update of the second media content generated by the first preset generation model by adjusting preset visual description information;

[0185] The visual description information display unit is configured to execute a second selection instruction in response to triggering a preset smart update control and display preset visual description information.

[0186] In an optional embodiment, the at least one preset update control further includes: at least one of a preset custom control and a preset recommendation control;

[0187] Among them, the preset custom control is used to trigger the update of the second media content through the media content customized by the target account; the preset recommendation control is used to trigger the update of the second media content through the media content recommended by the system.

[0188] In an optional embodiment, the first content display unit includes:

[0189] a third content display unit configured to execute, in response to a third selection instruction for target updated media content, updating the second media content displayed on the preset page to the target updated media content;

[0190] The fourth content display unit is configured to execute the first content generation instruction in response to the triggering based on the preset page, and display the second target media content.

[0191] In an optional embodiment, the first target media content is replaced with fourth target media content, where the fourth target media content is media content generated by the second preset generation model based on the first media content and the second media content; the second content display module 620 includes:

[0192] The fifth content display unit is configured to display fourth target media content in response to the first content generation instruction.

[0193] In an optional embodiment, when the fourth target media content is a target video, the target template media content is a target template video, and the first media content includes a first frame image of the target video, the second media content is a last frame image of the target video, and the preset visual description information is image visual feature information of the last frame image in the target template video;

[0194] When the fourth target media content is a target video, the target template media content is a target template video, and the first media content includes a last frame image of the target video, the second media content is a first frame image of the target video, and the preset visual description information is image visual feature information of the first frame image in the target template video;

[0195] The subject in the first media content and the subject in the second media content are both target objects.

[0196] In an optional embodiment, the first media content is based on the target object; the media content feature information is the object feature information of the target object; and the second media content is generated by a first preset generation model based on the object feature information and preset visual description information, and is media content based on the target object.

[0197] In an optional embodiment, the first content display module 610 includes:

[0198] a sixth content display unit, configured to display, in response to the first selection instruction, at least one preset content generation parameter corresponding to the first media content, the second media content, and the target template media content on a preset page;

[0199] The first target media content is replaced with the fifth target media content, and the second content display module 620 includes:

[0200] The seventh content display unit is configured to display fifth target media content in response to the first content generation instruction, where the fifth target media content is media content generated by the second preset generation model based on the second media content and at least one preset content generation parameter.

[0201] In an optional embodiment, the above device further includes:

[0202] a content generation parameter updating module configured to execute, in response to a second update instruction for a target content generation parameter, an update instruction for updating a target content generation parameter among at least one preset content generation parameter displayed on a preset page to an updated content generation parameter corresponding to the second update instruction, wherein the target content generation parameter is any one of the at least one preset content generation parameter;

[0203] The first target media content is replaced with the sixth target media content, and the second content display module 620 includes:

[0204] an eighth content display unit configured to execute, in response to the first content generation instruction, display sixth target media content, where the sixth target media content is media content generated by the second preset generation model based on the second media content and at least one updated content generation parameter;

[0205] The at least one updated content generation parameter is a content generation parameter obtained by updating a target content generation parameter in at least one preset content generation parameter based on the updated content generation parameter.

[0206] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0207] Figure 7 This is a block diagram of an electronic device for generating media content according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 7 The electronic device may include an RF (Radio Frequency) circuit 710, a memory 720 including one or more computer-readable storage media, an input unit 730, a display unit 740, a sensor 750, an audio circuit 760, a WiFi (wireless fidelity) module 770, a processor 780 including one or more processing cores, and a power supply 790. Those skilled in the art will understand that Figure 7 The terminal structure shown in the figure does not constitute a limitation to the terminal, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0208] The RF circuit 710 can be used to receive and transmit signals during information transmission or calls. Specifically, it receives downlink information from the base station and transmits it to one or more processors 780 for processing. Furthermore, it transmits uplink data to the base station. Typically, the RF circuit 710 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a subscriber identity module (SIM) card, a transceiver, a coupler, an LNA (low noise amplifier), a duplexer, and the like. Furthermore, the RF circuit 710 can communicate with the network and other terminals via wireless communication. Wireless communication can utilize any communication standard or protocol, including but not limited to GSM (Global System of Mobile Communications), GPRS (General Packet Radio Service), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), LTE (Long Term Evolution), email, and SMS (Short Messaging Service).

[0209] The memory 720 can be used to store software programs and modules. The processor 780 executes various functional applications and data processing by running the software programs and modules stored in the memory 720. The memory 720 may mainly include a program storage area and a data storage area. The program storage area can store the operating system, application programs required for functions, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, the memory 720 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 720 may also include a memory controller to provide access to the memory 720 by the processor 780 and the input unit 730.

[0210] The input unit 730 can be used to receive digital or character input and generate keyboard, mouse, joystick, optical, or trackball signal input related to user settings and function control. Specifically, the input unit 730 may include a touch-sensitive surface 731 and other input devices 732. The touch-sensitive surface 731, also known as a touch display or touchpad, can detect user touch operations on or near it (for example, operations performed by a user using a finger, stylus, or any other suitable object or accessory on or near the touch-sensitive surface 731) and drive corresponding connected devices according to a pre-set program. Optionally, the touch-sensitive surface 731 may include a touch detection device and a touch controller. The touch detection device detects the user's touch position and detects signals generated by the touch operation, transmitting the signals to the touch controller. The touch controller receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 780. It can also receive and execute commands from the processor 780. In addition, the touch-sensitive surface 731 can be implemented using various types of touch devices, including resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 731, the input unit 730 may further include other input devices 732. Specifically, the other input devices 732 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power keys, etc.), a trackball, a mouse, and a joystick.

[0211] The display unit 740 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the terminal. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 740 may include a display panel 741. Optionally, the display panel 741 can be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), or the like. Furthermore, the touch-sensitive surface 731 can cover the display panel 741. When the touch-sensitive surface 731 detects a touch operation on or near it, it transmits the information to the processor 780 to determine the type of touch event. The processor 780 then provides a corresponding visual output on the display panel 741 based on the type of touch event. The touch-sensitive surface 731 and the display panel 741 can be two independent components to implement input and output functions. However, in some embodiments, the touch-sensitive surface 731 and the display panel 741 can also be integrated to implement input and output functions.

[0212] The terminal may also include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor may adjust the brightness of the display panel 741 according to the brightness of the ambient light, and the proximity sensor may turn off the display panel 741 and / or the backlight when the terminal is moved to the ear. As a type of motion sensor, the gravity acceleration sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that identify the terminal posture (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that can be configured in the terminal, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be described here.

[0213] Audio circuit 760, speaker 761, and microphone 762 provide an audio interface between the user and the terminal. Audio circuit 760 converts received audio data into electrical signals and transmits them to speaker 761, which then converts them into sound signals for output. Microphone 762, on the other hand, converts collected sound signals into electrical signals, which are then received by audio circuit 760 and converted into audio data. The audio data is then processed by processor 780 and transmitted via RF circuit 710 to, for example, another terminal. Alternatively, the audio data may be output to memory 720 for further processing. Audio circuit 760 may also include an earphone jack to allow communication between an external headset and the terminal.

[0214] WiFi is a short-range wireless transmission technology. The terminal can help users send and receive emails, browse web pages and access streaming media through the WiFi module 770. It provides users with wireless broadband Internet access. Figure 7 A WiFi module 770 is shown, but it is understandable that it is not an essential component of the terminal and can be omitted as needed without changing the essence of the invention.

[0215] Processor 780 is the terminal's control center, connecting all components of the terminal using various interfaces and circuits. By running or executing software programs and / or modules stored in memory 720 and accessing data stored in memory 720, it performs various terminal functions and processes data, thereby providing overall terminal monitoring. Optionally, processor 780 may include one or more processing cores; preferably, processor 780 may integrate an application processor and a modem processor, with the application processor primarily handling the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 780.

[0216] The terminal also includes a power supply 790 (e.g., a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the processor 780 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 790 can also include any of one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other components.

[0217] Although not shown, the terminal may further include a camera, a Bluetooth module, etc., which will not be described in detail herein. Specifically, in this embodiment, the display unit of the terminal is a touch screen display, and the terminal further includes a memory and one or more programs, wherein the one or more programs are stored in the memory and configured to execute instructions in the method embodiment of the present invention by one or more processors.

[0218] In an exemplary embodiment, an electronic device is further provided, including: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the media content generation method in the embodiment of the present disclosure.

[0219] In an exemplary embodiment, a computer-readable storage medium is further provided. When instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the media content generating method in the embodiment of the present disclosure.

[0220] In an exemplary embodiment, a computer program product containing instructions is also provided. When the computer program product is run on a computer, the computer is caused to execute the media content generating method in the embodiment of the present disclosure.

[0221] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, which can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0222] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0223] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for generating media content, characterized in that: include: In response to a first selection instruction of a target account for target template media content, displaying first media content and second media content customized by the target account on a preset page, where the second media content is generated by a first preset generation model based on media content feature information corresponding to the first media content and preset visual description information corresponding to the target template media content; In response to the first content generation instruction, first target media content is displayed, where the first target media content is media content generated by a second preset generation model based on the second media content.

2. The media content generation method according to claim 1, wherein: The method further comprises: In response to an update trigger instruction for the second media content, displaying the preset visual description information; In response to a first update instruction for the preset visual description information, updating the preset visual description information to updated visual description information corresponding to the first update instruction; In response to the second content generation instruction, display at least one updated media content corresponding to the second media content, where the at least one updated media content is generated by the first preset generation model based on the media content feature information and the updated visual description information; The first target media content is replaced with the second target media content, and the displaying of the first target media content in response to the first content generation instruction includes: In response to the first content generation instruction triggered by the target updated media content, presenting the second target media content; The target updated media content is one or more updated media contents among the at least one updated media content, and the second target media content is media content generated by the second preset generation model based on the target updated media content.

3. The media content generation method according to claim 2, characterized in that: The first target media content is further replaced by third target media content, where the third target media content is media content generated by the second preset generation model based on the first media content and the target updated media content; In response to the first content generation instruction triggered by the target updated media content, presenting the second target media content includes: In response to the first content generation instruction triggered based on the target updated media content, the third target media content is presented.

4. The method for generating media content according to claim 2, wherein: The displaying of the preset visual description information in response to the update trigger instruction for the second media content includes: In response to the update trigger instruction, display at least one preset update control corresponding to the second media content, the at least one preset update control including a preset smart update control, the preset smart update control being used to trigger updating of the second media content generated by the first preset generation model by adjusting the preset visual description information; In response to a second selection instruction triggered for the preset smart update control, the preset visual description information is displayed.

5. The media content generation method according to claim 4, characterized in that: The at least one preset update control further includes: at least one of a preset custom control and a preset recommendation control; Among them, the preset custom control is used to trigger the update of the second media content through the media content customized by the target account; the preset recommendation control is used to trigger the update of the second media content through the media content recommended by the system.

6. The media content generation method according to claim 2, characterized in that: In response to the first content generation instruction triggered by the target updated media content, presenting the second target media content includes: In response to a third selection instruction for the target updated media content, updating the second media content displayed on the preset page to the target updated media content; In response to the first content generation instruction triggered based on the preset page, the second target media content is displayed.

7. The media content generation method according to claim 1, characterized in that: The first target media content is replaced with fourth target media content, where the fourth target media content is media content generated by the second preset generation model based on the first media content and the second media content; The displaying of the first target media content in response to the first content generation instruction includes: In response to the first content generation instruction, the fourth target media content is presented.

8. The media content generation method according to claim 7, characterized in that: In a case where the fourth target media content is a target video, the target template media content is a target template video, and the first media content includes a first frame image of the target video, the second media content is a last frame image of the target video, and the preset visual description information is image visual feature information of the last frame image in the target template video; In a case where the fourth target media content is a target video, the target template media content is a target template video, and the first media content includes a last frame image of the target video, the second media content is a first frame image of the target video, and the preset visual description information is image visual feature information of the first frame image in the target template video; The subject in the first media content and the subject in the second media content are both target objects.

9. The method for generating media content according to any one of claims 1 to 8, wherein: The first media content is based on a target object; the media content feature information is object feature information of the target object; The second media content is generated by the first preset generation model based on the object feature information and the preset visual description information, and is media content with the target object as the main body.

10. The media content generation method according to any one of claims 1 to 8, characterized in that: In response to the first selection instruction of the target account for the target template media content, displaying the first media content and the second media content customized by the target account on the preset page includes: In response to the first selection instruction, displaying at least one preset content generation parameter corresponding to the first media content, the second media content, and the target template media content on the preset page; The first target media content is replaced with the fifth target media content, and the displaying of the first target media content in response to the first content generation instruction includes: In response to the first content generation instruction, the fifth target media content is displayed, where the fifth target media content is media content generated by the second preset generation model based on the second media content and the at least one preset content generation parameter.

11. The media content generation method according to claim 10, characterized in that: The method further comprises: In response to a second update instruction for the target content generation parameter, updating the target content generation parameter among the at least one preset content generation parameter displayed on the preset page to an updated content generation parameter corresponding to the second update instruction, the target content generation parameter being any one of the at least one preset content generation parameter; The first target media content is replaced with the sixth target media content, and the displaying of the first target media content in response to the first content generation instruction includes: In response to the first content generation instruction, display the sixth target media content, where the sixth target media content is media content generated by the second preset generation model based on the second media content and the at least one updated content generation parameter. The at least one updated content generation parameter is a content generation parameter obtained by updating the target content generation parameter in the at least one preset content generation parameter based on the updated content generation parameter.

12. A media content generating device, characterized in that: include: The first content display module is configured to execute, in response to a first selection instruction of a target account for target template media content, displaying first media content and second media content customized by the target account on a preset page, where the second media content is generated by a first preset generation model based on media content feature information corresponding to the first media content and preset visual description information corresponding to the target template media content; The second content display module is configured to execute in response to the first content generation instruction and display the first target media content, where the first target media content is media content generated by a second preset generation model based on the second media content.

13. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the media content generation method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the media content generation method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Video generation platform based on AI

    CN118748742A

  • Multimedia content generation method and device, electronic equipment and storage medium

    CN118778860A

  • Media content generation method and device, equipment, medium and program product

    CN119668456A

  • Media data generation method and device, equipment, medium and product

    CN119718121A

  • Video generating method and apparatus, and terminal device and storage medium

    US20240118787A1