A method, apparatus, device and medium for generating media content
By acquiring and analyzing trending information, and generating creative prompts for special effects based on models, the problem of low efficiency in generating special effects packages is solved, and efficient and high-quality inspiration-related media content generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2026-06-18
- Publication Date
- 2026-08-04
AI Technical Summary
In existing special effects platforms, creators rely on discovering trending topics on their own when making special effects packages, which is slow and inefficient. When creative ideas run dry, the quality of inspiration-related media content is also poor.
By acquiring initial information, the type of special effects is determined based on model analysis, and special effects creation prompts are obtained. Media content used to prompt the generation of special effects packages is generated, and inspiration-related media content is automatically filtered and generated using a multi-dimensional model.
It enables the rapid filtering of potential information suitable for generating special effects packages from massive amounts of trending information, improving the efficiency and quality of generating inspiration-related media content and reducing the cost of understanding and trial and error.
Smart Images

Figure CN122513636A_ABST
Abstract
Description
Technical Field
[0001] This article relates to the field of computer technology, and in particular to a method, apparatus, device and medium for generating media content. Background Technology
[0002] Special effects are widely used on content platforms and social media platforms, and are very popular with users. For example, users can create and publish videos based on the special effects provided by the platform, and other users can then create similar videos using the same effects.
[0003] However, the creation of special effects themselves requires special effects creators to create based on sources of inspiration or creative ideas. Summary of the Invention
[0004] To address the aforementioned technical issues, this paper provides a method, apparatus, device, and medium for generating media content.
[0005] This article provides a method for generating media content, the method comprising: Obtain first information; Obtain at least one special effect type corresponding to the first information, and obtain at least one special effect creation prompt content corresponding to each special effect type; Based on the first information and the at least one special effects creation prompt, at least one first media content is generated, wherein the first media content is used to prompt the generation of a special effects package.
[0006] This article also provides a media content generation apparatus, the apparatus comprising: The acquisition module is used to acquire the initial information. The type module is used to obtain at least one special effect type corresponding to the first information, and to obtain at least one special effect creation prompt content corresponding to each special effect type; The generation module is used to generate at least one first media content based on the first information and the at least one special effects creation prompt content, wherein the first media content is used to prompt the generation of a special effects package.
[0007] This document also provides an electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the executable instructions to implement a method for generating media content as provided herein.
[0008] This document also provides a computer-readable storage medium storing a computer program for performing a method for generating media content as provided herein.
[0009] The technical solution provided in this paper has the following advantages compared with existing technologies: The media content generation scheme provided in this paper obtains first information; obtains at least one special effect type corresponding to the first information, and obtains at least one special effect creation prompt content corresponding to each special effect type; based on the first information and at least one special effect creation prompt content, at least one first media content is generated, wherein the first media content is used to prompt the generation of a special effect package. By adopting the above scheme, the corresponding special effect type is determined for information with high attention, and the special effect creation prompt content for the special effect type is obtained. Based on this information and the special effect creation prompt content, media content used to prompt the generation of a special effect package is generated. This achieves rapid generation of inspiration-related media content based on the special effect creation prompt content corresponding to the special effect type of hot information, not only improving the generation efficiency of inspiration-related media content, but also resulting in higher generation quality. Attached Figure Description
[0010] The above and other features, advantages, and aspects of this document will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0011] Figure 1 This is a schematic diagram of an interactive scenario provided in this article; Figure 2 A flowchart illustrating a method for generating media content provided in this paper; Figure 3 A flowchart illustrating another method for generating media content provided in this article; Figure 4 A diagram illustrating the primary media content provided in this article; Figure 5 This is a schematic diagram illustrating the media content generation process provided for this article; Figure 6 A schematic diagram illustrating the image generation process provided in this article; Figure 7 A schematic diagram of the structure of a media content generation device provided in this article; Figure 8 This is a schematic diagram of the structure of an electronic device provided in this article. Detailed Implementation
[0012] The present invention will now be described in more detail with reference to the accompanying drawings. While some aspects of the present invention are illustrated in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the situations set forth herein; rather, these situations are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and content are for illustrative purposes only and are not intended to limit the scope of this invention.
[0013] It is understood that before using the technical solutions described in this article, relevant users should be informed of the type, scope of use, and usage scenarios of the information involved in this article through appropriate means in accordance with relevant laws and regulations, and their authorization should be obtained. Relevant users may include any type of rights holder, such as individuals, enterprises, or groups.
[0014] For example, in response to receiving an active request from a user, a prompt message is sent to the relevant user to clearly inform the user that the requested operation will require obtaining and using the user's information. This allows the relevant user to choose whether to provide information to the software or hardware such as electronic devices, applications, servers, or storage media that perform the operation of the technical solution described herein, based on the prompt message.
[0015] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user. This could be a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" with the information provided to the electronic device.
[0016] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation methods described in this article. Other methods that comply with relevant laws and regulations may also be applied to the implementation methods described in this article.
[0017] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0018] It should be understood that the steps described in the method embodiments herein may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this document is not limited in this respect.
[0019] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one situation" means "at least one situation"; the term "another situation" means "at least one additional situation"; the term "some situations" means "at least some situations". Definitions of other terms will be given in the following description.
[0020] It should be noted that the concepts of "first" and "second" mentioned in this article are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.
[0021] It should be noted that the terms "one" and "more" used in this document are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0022] The names of messages or information exchanged between multiple devices in the embodiments herein are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0023] Currently, when creating special effects packages on special effects platforms, creators usually rely on discovering trending topics on their own, resulting in slow response and low efficiency. Even if creators are aware of trending topics, they still need to conceive creative ideas to generate media content as a source of inspiration. When creative ideas run dry, efficiency is low, and the media content related to the inspiration is not rich enough and of poor quality.
[0024] To address the aforementioned issues, this paper presents a method for generating media content, which will be described below with specific examples.
[0025] For example, Figure 1 This article provides a schematic diagram of an interactive scenario, such as... Figure 1 As shown, the media content generation method is applied to a media content generation platform, which is set in the electronic device 101 shown in the figure. In response to the user's operation to generate media content, the electronic device 101 can obtain first information, obtain at least one special effect type corresponding to the first information, and obtain at least one special effect creation prompt content corresponding to each special effect type. Based on the first information and at least one special effect creation prompt content, at least one first media content for prompting the generation of special effect packages is generated. Then, the first media content can be sent to the client device 102 for display. The user of the electronic device 101 can be a technician who generates media content related to the backend special effect packages, and the user of the client device 102 can be a special effect personnel who generates special effect packages or a user who uses special effect packages.
[0026] Figure 2 This document presents a flowchart illustrating a method for generating media content. This method can be executed by a media content generation device, which can be implemented using software and / or hardware, and is typically integrated into an electronic device. Figure 2 As shown, the method includes: Step 201: Obtain the first information.
[0027] The first piece of information can be information that is predicted to become a hot topic in the future, information that has the potential to be used for special effects packaging or is more suitable for special effects packaging, and information can include at least one of text, video, and image.
[0028] In some cases, obtaining the first information may include: obtaining multiple pieces of second information, wherein the second information is information whose attention meets the first condition; filtering the multiple pieces of second information to obtain the first information, wherein the number of first information is at least one.
[0029] The second piece of information can be information that has been widely discussed and disseminated over a period of time; it can be understood as trending information with high attention. The second piece of information can be determined based on its level of attention, which can be determined by at least one of the following: page views, search volume, interaction volume, and dissemination volume. It can also be predicted based on experience. Interaction volume can include the amount of data from interactive actions such as likes, comments, reposts, and plays. The second piece of information that meets the first condition can be predicted based on experience, for example, through manual or model prediction. For instance, if a person predicts that a particular holiday will become a trending topic, then the second piece of information is information related to that holiday. A model can be trained on given data to complete the prediction task. The first condition can be a condition set for attention, specifically based on the actual situation. For example, when attention is determined by page views, the first condition can be met if the page views exceed a corresponding threshold. Alternatively, the first condition can be either manually set or output by a model; that is, anything manually set or output by a model meets the first condition.
[0030] After receiving a user's generation operation, the media content generation device can acquire multiple sources of second information. These sources can include preset programs, offline data, prediction platforms, and manual input. Preset programs can be video programs, live streaming programs, etc., and can acquire the hot topic ranking results of these programs at regular intervals, extracting second information whose attention meets the first condition from the ranking results. The interval can be, for example, one hour. Offline data can be acquired from the hot topic data module of a data warehouse, extracting second information whose attention meets the first condition from the offline data. The prediction platform can call a prediction model to output future second information based on the current time, existing second information, and historical preset time predictions. For example, it can predict future hot topics such as festivals, solar terms, trends, and sports. The historical preset time can be, for example, one week or one day. Manually input second information can be obtained in response to the input controls of the media content generation platform.
[0031] The media content generation device can then filter multiple pieces of second information to obtain at least one piece of first information. Specifically, it can acquire third information, determine the compatibility between the multiple pieces of second information and the third information, and filter to obtain the first information based on the compatibility. The third information includes feature information of a preset special effects package. This enables the rapid filtering of hot information with the potential for creating special effects packages or suitable for special effects package creation from a large amount of hot information, thereby increasing the likelihood of corresponding special effects packages becoming hot topics in subsequent inspiration media.
[0032] The compatibility between the second and third information can be understood as the compatibility between the second information and the special effects package. It can be measured by the degree of special effects generation, assessing the matching and fit between the content characteristics of the second information and the special effects package in terms of visual, auditory, and interactive presentation. It represents a quantitative value indicating the probability or likelihood that the second information is suitable for creating a special effects package. A special effects package can be a collection of resources such as props and code that process media content to produce corresponding special effects. For example, a user can change the background of an uploaded image to a sunset background using a sunset background effect package.
[0033] The third information can be the constraint information used by the filtering model to filter the second information. The preset effects package can be an effects package whose historical attention meets the second condition. The meaning of attention is as described above. The second condition can be the same as or different from the first condition. The feature information of the preset effects package can include generation rules, keywords, etc. The generation rules can be the formula or rule of the preset effects package during creation. For example, the generation rule can be a combination of text and enlargement. Keywords can be representative words in the associated text information of the special effects media content corresponding to the preset effects package. The special effects media content can be the media content generated by using the preset effects package. The associated text information of the special effects media content can include the title, description, and comments of the special effects media content.
[0034] When the media content generation device obtains the first information by filtering multiple second pieces of information, it can first acquire the third information, fill the prompt word template of the filtering model with the multiple second and third pieces of information to generate prompt words, input the prompt words into the filtering model to judge the fit between each second piece of information and the third piece of information, and extract the second information with a fit greater than the corresponding threshold as the first information output. The filtering model can be a model for performing special effects generation probability analysis and information filtering, and can be a pre-trained large model, multimodal model, deep model or machine learning model, etc.; then, at least one piece of information can be sorted in order of fit from high to low to obtain a sorting result, and at least one piece of information can be displayed according to the sorting result.
[0035] The above solution uses a model to predict future hot topics for special effects production from multiple fields and dimensions, enabling automated screening of potential information suitable for generating special effects packages from massive amounts of hot topics, thereby improving the efficiency and accuracy of information screening in the special effects generation process.
[0036] Step 202: Obtain at least one special effect type corresponding to the first information, and obtain at least one special effect creation prompt content corresponding to each special effect type.
[0037] Special effects types can be categorized according to the visual, auditory, and other effects they present. For example, special effects types can include game-related, performance-related, and makeup-related types, with no specific limitations. Special effects creation tips can be summaries of content that can be referenced when creating a special effects package of a certain type. This can be a form of experience-based content. Optionally, special effects creation tips for a particular special effects type can include at least one of the following: special effects type definition information, content structure information, requirement information, special effects generation logic information, and example information. Special effects type definition information can be an understanding of the definition of the special effects type. Content structure information can be the compositional structure of the media content corresponding to the special effects type, such as the compositional structure of a video. Requirement information can be the basic requirements of the special effects package corresponding to the special effects type, such as formatting. Special effects generation logic information can be a description of the creative rules or formulas for the special effects package corresponding to the special effects type. Example information can be an example of generating first media content by combining the special effects type with first information, that is, an example of generating media content as a source of inspiration by combining trending information. Example information can include example descriptions and example images.
[0038] After acquiring the first information, the media content generation device can analyze the first information through a model and map it to a special effect type, and extract the corresponding special effect creation prompts based on the special effect type.
[0039] In one scenario, obtaining at least one effect type corresponding to the first information may include: inputting the first information into a first model, obtaining content understanding information output by the first model, wherein the content understanding information includes text content understanding information or video content understanding information; and determining at least one effect type of the first information based on the content understanding information. Optionally, determining at least one effect type of the first information based on the content understanding information may include: inputting the content understanding information into a second model, obtaining the information type corresponding to the first information output by the second model; and mapping the information type to the corresponding at least one effect type using a preset rule.
[0040] Content understanding information can be information generated from different dimensions of information related to the first information, reflecting its inherent logic and unique characteristics. It is obtained through in-depth analysis of the first information using a first model. This content understanding information allows for a more comprehensive and accurate understanding of the first information. The first model can be a model that defines the content understanding information, and the specific model type is not limited. When the first information is video-type information, the content understanding information is video content understanding information; when the first information is text-type information, the content understanding information is text content understanding information. Information type can be a type formed by systematically classifying the first information based on specific criteria, which can be understood as a hotspot classification. It can be determined through analysis using a second model, which can classify information into preset hotspot classifications; the specific model type is not limited. Preset rules can be rules that establish a mapping relationship between information types and special effects types. One information type can correspond to one or more special effects types, and preset rules can be set in a special effects knowledge base.
[0041] When the media content generation device acquires at least one special effect type corresponding to the first information, it can input the first information into a first model. If the first information is text-based, the first model searches for causal information, related tasks, etc., of the first information to extract, deduplicate, and summarize information to generate and output the corresponding text content understanding information. If the first information is video-based, the first model performs multimodal decomposition and analysis on the video, extracting visual elements, auditory elements, and atmospheric elements to generate and output corresponding video content understanding information. Visual elements may include, for example, camera movement, makeup, and clothing; auditory elements may include background music, dialogue, and rhythm; and atmospheric elements may include the overall atmosphere. The content understanding information can then be input into a second model, which classifies and outputs the corresponding information type according to a preset classification system. The device then retrieves preset rules and searches for at least one special effect type mapped to the information type within those rules, ultimately determining that at least one special effect type as the at least one special effect type of the first information. In the above scheme, based on the model's interpretation and classification capabilities, unstructured hot topic information can be mapped to a standardized hot topic classification system, and then mapped to appropriate special effects types based on preset rules, ensuring the accuracy of special effects type mapping.
[0042] In another scenario, obtaining at least one effect type corresponding to the first information may include: in response to the first information being video type information and generated based on an effect package, obtaining at least one effect type corresponding to the effect package.
[0043] When the media content generation device acquires at least one special effect type corresponding to the first information, if the first information is video type information (i.e., the first information is a video), it can determine whether the video was generated using a special effect package. If so, it can acquire at least one special effect type corresponding to the special effect package and determine that at least one special effect type corresponding to the first information. By inheriting the special effect types of popular videos that use special effect packages, the accuracy of special effect type determination is improved. If the video does not use a special effect package, the information type of the first information can be determined based on its video content understanding information in the manner described above and mapped to the corresponding special effect type through preset rules. Alternatively, the video can be deleted; that is, for video type information, the unused portions can be deleted to improve the accuracy of special effect type mapping.
[0044] In one scenario, obtaining at least one special effect creation prompt for each special effect type includes at least one of the following: obtaining at least one special effect creation prompt for each special effect type from a special effect knowledge base; wherein the special effect knowledge base stores multiple special effect types and special effect creation prompts for each special effect type; inputting each special effect type into a third model, and obtaining at least one special effect creation prompt for each special effect type output by the third model; wherein the third model is obtained by training based on the special effect knowledge base.
[0045] The special effects knowledge base can be a pre-created knowledge base for the creation of special effects packages, including the underlying creation logic and creation-related knowledge of various special effects packages. It can be a vector knowledge base, which includes multiple special effects types and corresponding creation prompts for each special effects type. The third model can be a model used to determine the creation prompts for special effects types. The specific model type is not limited. The third model can be trained based on the special effects knowledge base to learn the relationship between special effects types and creation prompts.
[0046] When the media content generation device acquires special effects creation prompts for each special effects type, it can either query a special effects knowledge base to determine at least one corresponding special effects creation prompt based on the special effects type, or input the special effects type into a third model, which will analyze and determine at least one corresponding special effects creation prompt and output it. By constructing a special effects knowledge base or determining the special effects creation prompts through a model, the accuracy of the determination of special effects creation prompts is ensured.
[0047] Step 203: Based on the first information and at least one special effects creation prompt, generate at least one first media content, wherein the first media content is used to prompt the generation of a special effects package.
[0048] The first media content can be inspiration or creative material for creating special effects packages. Based on this first media content, special effects artists or users can refer to it to create their own special effects packages. The first media content can include images and text. Special effects packages can be used to apply effects to second media content. The second media content can be the media content that needs to be generated, or it can be user-uploaded media content. The second media content can be images or videos.
[0049] In some cases, primary media content is provided to special effects creators through special effects creation platforms to help them create popular and viral special effects, thereby enriching the media content on the content platform and providing people with a richer spiritual life.
[0050] After acquiring the special effects creation prompts corresponding to the first information's special effects type, the media content generation device can input the first information and the special effects creation prompts into a content generation model. The content generation model, combined with basic content generation knowledge and generation rules, outputs the first media content corresponding to the special effects type. Each first media content corresponds to at least one special effects type, with the specific number set according to actual conditions. For example, each piece of first information can correspond to 5 special effects types, and each special effects type can generate 10 pieces of first media content, ultimately generating 50 pieces of first media content. This batch generates inspiration information for special effects packages corresponding to real-time hot topics. The aforementioned basic content generation knowledge can be underlying fundamental knowledge specific to inspiration or creative generation, obtained through input by professionals. The generation rules can be constraints imposed during the generation of first media content, such as anomaly detection rules.
[0051] For example, Figure 3 A flowchart illustrating another method for generating media content provided in this article, such as... Figure 3 As shown, in one scenario, generating at least one first media content based on the first information and at least one special effects creation prompt may include: Step 301: Obtain the content understanding information of the first information.
[0052] The first information, or content understanding information, can be generated from different dimensions of information related to the content understanding information, reflecting its inherent logic and unique characteristics. It is obtained by in-depth analysis of the first information using the first model. Through this content understanding information, the second information can be understood more comprehensively and accurately.
[0053] Step 302: Input the content understanding information of the first information and the special effects creation prompts into the fourth model, and obtain the first description output by the fourth model.
[0054] The fourth model can be a model that generates the text portion of the first media content, or it can be a text sub-model of the aforementioned content generation model. The first description can be text content in the first media content that describes the effects and characteristics of the special effects package to be generated, or it can be text-based inspirational information.
[0055] The media content generation device can input the content understanding information of the first information and the creation prompts for each special effect into a fourth model. The fourth model, combined with generation rules, outputs a corresponding first description. There is at least one first description, and each first description corresponds to one special effect creation prompt. Optionally, the first description may also include first information, used to indicate that the current first media content is generated based on this first information, distinguishing it from first media content generated by other methods.
[0056] Step 303: Generate a first image based on the special effect type corresponding to the first information and the first description.
[0057] The first image may be a schematic diagram of the effect of converting the first description into an image in the first media content on the special effects package to be generated. The first media content includes the first description and the first image.
[0058] In some cases, generating a first image based on the special effect type corresponding to the first information and the first description may include: inputting the first information into a fifth model to obtain first constraint information output by the fifth model, wherein the first constraint information is used to describe the style corresponding to the first information; extracting example information from the corresponding special effect creation prompt content based on the special effect type corresponding to the first information, and inputting the example information into a sixth model to obtain second constraint information output by the sixth model, wherein the example information includes an example description and an example image; obtaining third constraint information, inputting the first constraint information, the second constraint information, the third constraint information, and the first description into a seventh model to obtain the first image output by the seventh model, wherein the third constraint information includes at least one attribute of the image.
[0059] The first constraint information can be a descriptive description reflecting the style and imagery of the first information. The fifth model can be a model that generates the first constraint information, and the specific model type is not limited. Example information can be an example of generating the first media content based on the first information, that is, an example of generating media content as inspiration based on trending information. Example information can include example descriptions and example images. The example description can be an example of text describing the effect and characteristics of the upcoming special effects package, and the example image can be a schematic diagram corresponding to the example description. The example description and example image correspond to each other. The second constraint information can be key conditions with immutable properties during model deployment and inference. These conditions can maintain the style effect during the generation of the first image. The second constraint information can include model parameters, static knowledge, etc. The second constraint information can be extracted from the example information of the special effects type to which the first information belongs through the sixth model. The sixth model can be a model that generates the second constraint information, and the specific model type is not limited. The third constraint information can be basic constraint information on the first image. The third constraint information can include at least one attribute of the image, such as image size, image content restrictions, etc. Image content restrictions can include, for example, abnormal words not allowed in the image. The seventh model can be a model for image generation, such as a pre-trained large model, a multimodal model, etc., and is not limited in specifics.
[0060] When generating the first image of the first media content, the media content generation device can first extract the first constraint information of the first information. Specifically, the first information, or the first information and its content understanding information, can be input into the fifth model for analysis to extract style information as the first constraint information. Next, example information is extracted from the special effect creation prompts corresponding to the special effect type of the first information. This example information is input into the sixth model to output the second constraint information. Third constraint information is then obtained. The first, second, and third constraint information, along with the first description, are input into the seventh model, which generates a first image that conforms to the constraints and corresponds to the first description. In this scheme, the first description is input when generating the first image of the first media content, and multi-dimensional constraint information such as style, immutability, and image attributes is added for different special effect types, achieving the effect of generating corresponding inspiration diagrams with differentiated generation for different special effect types.
[0061] Optionally, the seventh model may include a text-based image model and an image-based image model. Based on the effect type, one can be selected for generating the first image. Each effect type has a pre-configured corresponding model. For example, if the effect type is game-related, the first image can be generated using the image-based image model in the seventh model. When generating the first image using the image-based image model, multiple reference images corresponding to the effect type can also be obtained from the style library. These multiple reference images, along with the aforementioned first constraint information, second constraint information, third constraint information, and first description, are input into the image-based image model to generate the first image, further improving the accuracy and quality of the inspired image generation. The style library may include multiple effect types and multiple reference images corresponding to each effect type. The reference images can reflect the content style, layout style, and element style corresponding to the effect type, serving as constraint information input to the model. The style library can be generated based on the first information of the video type corresponding to different effect types. Specifically, keyframes can be extracted from the first information. Keyframes can be video frames without blank spaces, black screens, or transition effects. Abnormal information in the keyframes is removed and placed into the style library.
[0062] Step 304: Based on the first description and the first image, generate first media content, wherein each piece of first media content corresponds to a special effect type and a special effect creation prompt content corresponding to the first information.
[0063] The media content generation device can generate a corresponding first description and a first image for a special effect creation prompt content corresponding to a special effect type of the first information, and combine the first description and the first image to determine a first media content, thereby obtaining at least one first media content.
[0064] For example, Figure 4 An illustration of the primary media content provided in this article, such as Figure 4 As shown in the figure, the first media content 400 is displayed. The first media content 400 includes a first image 401 and a first description 402. In addition, the first media content 400 may also include second information 403. For example, the second information 403 in the figure is "Stage Online", which means that the current first media content 400 is generated based on the theme of "Stage Online". The title of the first description 402 is "Your Stage". By browsing the first media content 400, a corresponding stage-related special effects package can be generated.
[0065] In the above solution, by acquiring content understanding information of hot topics, and based on this content understanding information and the corresponding special effects creation prompts, descriptive text as inspiration can be generated. Based on multi-dimensional constraint information such as style, immutable conditions, and image attributes, as well as the descriptive text, illustrative images as inspiration can be generated. The descriptive text and illustrative images are combined to form media content of inspiration, realizing the automated batch generation of media content for special effects metapack production.
[0066] The multiple models mentioned in this article may be the same model or different models, depending on the actual situation.
[0067] The media content generation scheme presented in this paper involves: obtaining first information; obtaining at least one special effect type corresponding to the first information, and obtaining at least one special effect creation prompt for each special effect type; and generating at least one first media content based on the first information and at least one special effect creation prompt, wherein the first media content is used to prompt the generation of a special effect package. By adopting the above scheme, the corresponding special effect type is determined for information with high public interest, and the special effect creation prompt content for that special effect type is obtained. Based on this information and the special effect creation prompt content, media content used to prompt the generation of a special effect package is generated. This achieves rapid generation of inspiration-related media content based on the special effect creation prompt content corresponding to the special effect type of hot information, not only improving the generation efficiency of inspiration-related media content but also resulting in higher generation quality.
[0068] The following example further illustrates the media content generation scheme. For instance, Figure 5 This article provides a schematic diagram of the media content generation process, such as... Figure 5 As shown in the diagram, the media content generation process can include three parts: information acquisition, information classification, and content generation. The workflow in the diagram can represent a model processing process; different workflows can correspond to the same model or different models. Specifically, in the information acquisition process, second information can be obtained from preset programs, offline data, prediction platforms, and manual input. The prediction workflow is used to predict second information based on hot topics from preset programs. Multiple second inputs are filtered by a workflow that selects first information based on the compatibility of each second information with the special effects package. In the information classification process, in response to the first information being text-type information, the text understanding workflow can generate corresponding text content understanding information. Since the first piece of information is video-type information, corresponding video content understanding information can be generated through the video understanding workflow. Then, either text content understanding information or video content understanding information is input into the classification workflow to determine the information type of the first piece of information. Using preset rules, the information type is mapped to an effect type to obtain the effect type of the first piece of information. During content generation, the content understanding information of the first piece of information, basic knowledge of content generation, and effect creation prompts for the corresponding effect type are input into the text workflow to generate the first description. The content understanding information of the first piece of information and the effect creation prompts for the corresponding effect type are input into the raw image workflow to generate the first image. For details on the first image generation process, please refer to [link to documentation]. Figure 6 Then, anomaly detection is performed on the first image. If the detection result is abnormal, the image is discarded and regenerated. If the detection result is normal, the first image and the second description are output as the first media content.
[0069] For example, Figure 6 This is a schematic diagram of the image generation process provided in this article, such as... Figure 6 As shown, the image generation process can include two parts: preprocessing and raw image generation. The preprocessing part can extract first constraint information based on the first information or its content understanding, including style extraction, information features, and filtering of anomalous information. It can also extract second constraint information based on example information corresponding to the effect type, using invariable conditions. During this extraction, invariable conditions can be extracted from the example image and description in the example information. The raw image generation part inputs the first, second, and third constraint information, along with the first description, into the raw image model corresponding to the effect type and outputs the first image. The raw image model includes a text-based raw image model and a graph-based raw image model; different effect types correspond to different models. When processing with the graph-based raw image model, reference images corresponding to the effect type can be obtained from the style library and combined to generate the first image. The first image can then be compressed before output. The style library is constructed based on the first information of the video type corresponding to different effect types, specifically including keyframe extraction, anomalous information filtering, and library storage.
[0070] This paper addresses the issues of slow human inspiration response, limited production capacity, and high costs of understanding and trial and error in the generation chain of media content corresponding to special effects packages. It proposes a model-based solution that integrates hot information from multiple sources, filters information that is more suitable for generating special effects packages, and compresses the chain into a mode of automatically generating inspirational media content and then creating special effects packages based on that inspirational media content. Firstly, at the data level, the system integrates real-time trending information from pre-set applications and pre-constructs a special effects knowledge base. This knowledge base provides corresponding creation prompts for different special effects types and includes pre-defined rules mapping between effect types and information types. Secondly, at the understanding level, the system can filter trending information with potential for creating special effects packages from multiple sources. It automatically selects suitable trending information from a massive pool of information and deeply analyzes the context, atmosphere, memes, and gameplay of trending topics through real-time search. The special effects knowledge base helps adapt trending information to different effect types, providing a foundation for generating inspirational media content. Finally, based on the aforementioned understanding of effect types and trending topics, and combined with the special effects creation prompts from the knowledge base, the system automatically generates inspirational descriptions that can be directly used for effect creation. It also supports generating highly relevant effect illustrations for each inspirational description using strategies such as text-to-image and image-to-image, serving as a direct reference for creating special effects packages.
[0071] This paper achieves the following beneficial effects: it realizes the upgrade from passively chasing hot topics to actively predicting and adapting special effects to hot topics; special effects packages generated based on special effects adapted to hot topics are more likely to gain user attention and use; it automatically outputs high-quality inspiration-related descriptions and illustrations to special effects personnel, eliminating the time cost of manually finding inspiration and coming up with creative ideas; it reduces the cost of understanding and trial and error of inspiration; it shortens the time from obtaining inspiration to generating special effects; and it effectively improves the efficiency of generating inspiration-related content.
[0072] Figure 7 This is a schematic diagram of a media content generation device provided in this paper. This device can be implemented by software and / or hardware and is generally integrated into an electronic device. Figure 7 As shown, the device includes: Module 701 is used to acquire the first information; Type module 702 is used to obtain at least one special effect type corresponding to the first information, and to obtain at least one special effect creation prompt content corresponding to each special effect type; The generation module 703 is used to generate at least one first media content based on the first information and the at least one special effects creation prompt content, wherein the first media content is used to prompt the generation of a special effects package.
[0073] Optionally, the acquisition module 701 is used for: Obtain multiple pieces of second information, wherein the second information is information whose attention meets the first condition; The first information is obtained by filtering the plurality of second information, and the number of first information is at least one.
[0074] Optionally, type module 702 includes a first unit, the first unit being used for: The first information is input into the first model to obtain the content understanding information output by the first model, wherein the content understanding information includes text content understanding information or video content understanding information; Based on the content understanding information, at least one special effect type of the first information is determined.
[0075] Optionally, the first unit is specifically used for: Input the content understanding information into the second model, and obtain the information type corresponding to the first information output by the second model; The information type is mapped to at least one corresponding special effect type using preset rules.
[0076] Optionally, the first unit is further configured to: In response to the first information being video type information and generated based on an effects package, at least one effects type corresponding to the effects package is obtained.
[0077] Optionally, type module 702 includes a second unit, the second unit including at least one of the following: The first subunit is used to obtain at least one special effect creation prompt from the special effects knowledge base for each special effect type; wherein, the special effects knowledge base stores multiple special effect types and special effect creation prompts corresponding to each special effect type; The second subunit is used to input each of the special effects types into the third model and obtain at least one special effects creation prompt content corresponding to each of the special effects types output by the third model; wherein, the third model is obtained by training based on the special effects knowledge base.
[0078] Optionally, the generation module 703 includes: Content unit, used to acquire content understanding information of the first information; The text unit is used to input the content understanding information of the first information and the special effects creation prompts into the fourth model to obtain the first description output by the fourth model; An image unit is configured to generate a first image based on the special effect type corresponding to the first information and the first description; The generation unit is configured to generate the first media content based on the first description and the first image, wherein each piece of the first media content corresponds to a special effect type and a special effect creation prompt content corresponding to the first information.
[0079] Optionally, the image unit is used for: The first information is input into the fifth model to obtain the first constraint information output by the fifth model, wherein the first constraint information is used to describe the style corresponding to the first information; Based on the special effect type corresponding to the first information, example information is extracted from the corresponding special effect creation prompt content, and the example information is input into the sixth model to obtain the second constraint information output by the sixth model. The example information includes an example description and an example image. Obtain third constraint information, input the first constraint information, the second constraint information, the third constraint information, and the first description into the seventh model, and obtain the first image output by the seventh model, wherein the third constraint information includes at least one attribute of the image.
[0080] The media content generation apparatus provided herein can execute the media content generation method provided herein, and has the corresponding functional modules and beneficial effects of executing the method.
[0081] This document also provides a computer program product, including a computer program / instruction that, when executed by a processor, implements the media content generation method provided herein and has the same beneficial effects as the execution method.
[0082] Figure 8 This document provides a schematic diagram of the structure of an electronic device. The document also provides an electronic device comprising: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the executable instructions to implement the media content generation method provided herein, achieving the same beneficial effects as the execution method.
[0083] The following is a detailed reference. Figure 8 The diagram illustrates a suitable structural schematic for implementing the electronic device 800 described herein. The electronic device 800 described herein may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), tablet computers (PADs), portable media players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 8 The electronic device shown is merely an example and should not impose any limitations on the functionality and scope of this article.
[0084] like Figure 8 As shown, the electronic device 800 may include a processing unit 801 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 802 or a program loaded from storage device 808 into random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0085] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0086] In particular, according to this document, the processes described in the above-referenced flowchart can be implemented as computer software programs. For example, this document includes a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such an example, the computer program can be downloaded and installed from a network via communication device 809, or installed from storage device 808, or installed from ROM 802. When the computer program is executed by processing device 801, it performs the functions defined in the media content generation method of this document.
[0087] It should be noted that the computer-readable medium mentioned above can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an electrically erasable programmable read-only memory (EPROM), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this document, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, radio frequency (RF), etc., or any suitable combination thereof.
[0088] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as Hypertext Transfer Protocol (HTTP), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0089] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0090] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the aforementioned one or more programs, the electronic device causes the electronic device to: acquire first information; acquire at least one special effect type corresponding to the first information, and acquire at least one special effect creation prompt content corresponding to each special effect type; and generate at least one first media content based on the first information and the at least one special effect creation prompt content, wherein the first media content is used to prompt the generation of a special effect package. The computer-readable storage medium implementing the media content generation method provided herein has the same beneficial effects as the execution method.
[0091] Computer program code for performing the operations described herein may be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer may be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0092] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to the various examples herein. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the figures. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0093] The units described herein can be implemented in software or hardware. The names of the units are not, in some cases, limiting to the unit itself.
[0094] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field-Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0095] In the context of this document, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, flash memory, optical fiber, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0096] It is understandable that before using the technical solutions disclosed herein, users should be informed of the type, scope of use, and usage scenarios of the information involved in this article in an appropriate manner in accordance with relevant laws and regulations, and their authorization should be obtained.
[0097] The above description is merely a preferred example and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure herein is not limited to technical solutions formed by specific combinations of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed herein that have similar functions.
[0098] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. Multitasking and parallel processing may be advantageous in certain contexts. Similarly, while some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of this paper. Certain features described in the context of a single example can also be implemented in combination in a single example. Conversely, various features described in the context of a single example can also be implemented individually or in any suitable sub-combination in multiple examples.
[0099] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method for generating media content, comprising: Obtain first information; Obtain at least one special effect type corresponding to the first information, and obtain at least one special effect creation prompt content corresponding to each special effect type; Based on the first information and the at least one special effects creation prompt, at least one first media content is generated, wherein the first media content is used to prompt the generation of a special effects package.
2. The method according to claim 1, wherein obtaining the first information includes: Obtain multiple pieces of second information, wherein the second information is information whose attention meets the first condition; The first information is obtained by filtering the plurality of second information, and the number of first information is at least one.
3. The method according to claim 1, wherein obtaining at least one special effect type corresponding to the first information includes: The first information is input into the first model to obtain the content understanding information output by the first model, wherein the content understanding information includes text content understanding information or video content understanding information; Based on the content understanding information, at least one special effect type of the first information is determined.
4. The method according to claim 3, wherein determining at least one effect type of the first information based on the content understanding information includes: Input the content understanding information into the second model, and obtain the information type corresponding to the first information output by the second model; The information type is mapped to at least one corresponding special effect type using preset rules.
5. The method according to claim 1, wherein obtaining at least one special effect type corresponding to the first information includes: In response to the first information being video type information and generated based on an effects package, at least one effects type corresponding to the effects package is obtained.
6. The method according to claim 1, wherein obtaining at least one special effect creation prompt for each special effect type includes at least one of the following: Retrieve at least one special effect creation prompt for each of the aforementioned special effect types from the special effects knowledge base; wherein... The special effects knowledge base stores multiple special effects types and corresponding special effects creation tips for each type. Each of the aforementioned special effects types is input into a third model, and at least one special effects creation prompt corresponding to each of the aforementioned special effects types is obtained from the output of the third model; wherein, the third model is obtained by training based on the special effects knowledge base.
7. The method according to claim 1, wherein based on the first information and the at least one special effects creation prompt, at least one first media content is generated, comprising: Obtain the content understanding information of the first information; Input the content understanding information of the first information and the special effects creation prompts into the fourth model, and obtain the first description output by the fourth model; Based on the special effect type corresponding to the first information and the first description, a first image is generated; Based on the first description and the first image, the first media content is generated, wherein each piece of the first media content corresponds to a special effect type and a special effect creation prompt content corresponding to the first information.
8. The method according to claim 7, wherein generating a first image based on the special effect type corresponding to the first information and the first description, comprises: The first information is input into the fifth model to obtain the first constraint information output by the fifth model, wherein the first constraint information is used to describe the style corresponding to the first information; Based on the special effect type corresponding to the first information, example information is extracted from the corresponding special effect creation prompt content, and the example information is input into the sixth model to obtain the second constraint information output by the sixth model. The example information includes an example description and an example image. Obtain third constraint information, input the first constraint information, the second constraint information, the third constraint information, and the first description into the seventh model, and obtain the first image output by the seventh model, wherein the third constraint information includes at least one attribute of the image.
9. A media content generation apparatus, comprising: The acquisition module is used to acquire the initial information. The type module is used to obtain at least one special effect type corresponding to the first information, and to obtain at least one special effect creation prompt content corresponding to each special effect type; The generation module is used to generate at least one first media content based on the first information and the at least one special effects creation prompt content, wherein the first media content is used to prompt the generation of a special effects package.
10. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the media content generation method according to any one of claims 1-8.
11. A computer-readable storage medium storing a computer program for performing the method for generating media content according to any one of claims 1-8.