Multimedia content generation method, apparatus, device, and storage medium
By analyzing creative needs and generating multimedia content through large-scale artificial intelligence models, it solves the creative bottleneck for novice creators, provides detailed multimedia content generation solutions, and improves creation efficiency and content relevance.
Patent Information
- Application Number
- CN202310692511.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-09
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2043-06-09
AI Technical Summary
Multimedia content creators, especially beginners, often face creative bottlenecks, lacking reference materials and not knowing how to create content.
By calling a pre-configured large-scale artificial intelligence model, the system analyzes the user's creative needs, extracts the tag information of the specified type, and selects the matching target template from the preset prompt format template set to generate multimedia content that meets the user's needs.
It provides detailed multimedia content generation solutions for users to reference, improving creation efficiency and generating content that better meets user needs.
Smart Images

Figure CN116701669B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of large-scale artificial intelligence model technology, and more specifically, to a multimedia content generation method, apparatus, device, and storage medium. Background Technology
[0002] With the development of internet technology, the quantity and types of multimedia content have experienced explosive growth. Creators of multimedia content can collect information of interest or current online trends to create content such as articles, videos, and posters, which they can then publish online for other users to browse.
[0003] However, multimedia content creators also encounter creative bottlenecks, especially beginners who often face problems such as not knowing how to create content or lacking reference materials. Summary of the Invention
[0004] In view of the above problems, this application is made to provide a multimedia content generation method, apparatus, device, and storage medium to automatically generate multimedia content that meets the user's creative needs, for user reference and use. The specific solution is as follows:
[0005] Firstly, a method for generating multimedia content is provided, including:
[0006] Obtain users' creative needs;
[0007] The pre-configured large-scale artificial intelligence model is invoked to parse the creation requirements and extract the tag information of the specified type contained in the creation requirements;
[0008] Based on the extracted tag information of each type, a matching target prompt format template is selected from a preset prompt format template set. The prompt format template set contains prompt format templates that match different types of tags, and the prompt format templates contain knowledge information related to the matched type tags.
[0009] Based on the target prompt format template and the user's creative needs, a prompt instruction is generated and input into a large-scale artificial intelligence model to obtain multimedia content that meets the user's creative needs.
[0010] Preferably, the user's creative needs include the creation needs of any one of the following content types: articles, images, and video scripts. Correspondingly, the multimedia content generated by the large-scale artificial intelligence model includes any one of the following content types: articles, images, and video scripts.
[0011] Preferably, when the multimedia content generated by the large-scale artificial intelligence model that meets the user's creative needs is a video script, the method further includes:
[0012] Render the video script and output it.
[0013] Preferably, a pre-configured large-scale artificial intelligence model is invoked to parse the creative requirements and extract the tag information of the specified type contained in the creative requirements, including:
[0014] Obtain a pre-configured first prompt format template, which includes a user creation requirement slot. The first prompt format template is used to instruct a large artificial intelligence model to understand the user creation requirements within the user creation requirement slot and output the labels of various set types contained therein.
[0015] The acquired user creation requirements are filled into the user creation requirement slot to obtain the edited first prompt instruction, which is then sent to a large-scale artificial intelligence model to obtain the tag information of the set type included in the creation requirements output by the model.
[0016] Preferably, the preset prompt format template set includes a subset of prompt format templates corresponding to different content types. The prompt format templates in the subset are used to instruct the large artificial intelligence model to generate multimedia content of the content type corresponding to the subset.
[0017] Based on the extracted tag information of various types, a matching target prompt format template is selected from the preset prompt format template set, including:
[0018] The extracted tags of various types are combined to obtain tag groups;
[0019] From the subset of prompt format templates corresponding to the type of content to be created according to the user's creative needs, a target prompt format template that matches the tag group is selected. The subset of prompt format templates contains prompt format templates that match different tag groups.
[0020] Preferably, generating a prompt instruction based on the target prompt format template and the user's creative needs includes:
[0021] The target prompt format template and the user's creative requirements are combined into a prompt instruction.
[0022] Preferably, before generating the prompt instruction based on the target prompt format template and the user's creative needs, the method further includes:
[0023] For the specified type of tags missing in the user's creative needs, the specified type of tags are predicted by combining the user's account information and behavioral data on the multimedia platform;
[0024] The process of generating a prompt instruction based on the target prompt format template and the user's creative requirements includes:
[0025] The target prompt format template, the user's creative requirements, and the predicted tags of the specified type are concatenated into a prompt instruction.
[0026] Preferably, before generating the prompt instruction based on the target prompt format template and the user's creative needs, the method further includes:
[0027] Acquire scenario knowledge information related to users' creative needs;
[0028] The process of generating a prompt instruction based on the target prompt format template and the user's creative requirements includes:
[0029] The target prompt format template, the user's creative requirements, and the scene knowledge information are combined into a prompt instruction.
[0030] Preferably, the user's creative requirements include at least one of the following: the platform to which the content belongs, the content type, the theme, and keywords that the user is interested in.
[0031] Secondly, a multimedia content generation device is provided, comprising:
[0032] The creative needs acquisition unit is used to acquire users' creative needs;
[0033] The creation requirement parsing unit is used to call a pre-configured large-scale artificial intelligence model to parse the creation requirement and extract the tag information of the set type contained in the creation requirement;
[0034] The template matching unit is used to select a target prompt format template from a preset prompt format template set based on the extracted tag information of each type. The prompt format template set contains prompt format templates that match different types of tags, and the prompt format templates contain knowledge information related to the matched type tags.
[0035] The model call generation unit is used to generate a prompt instruction based on the target prompt format template and the user's creative needs, and input it into a large-scale artificial intelligence model to obtain multimedia content generated by the model that meets the user's creative needs.
[0036] Thirdly, a multimedia content generation device is provided, including: a memory and a processor;
[0037] The memory is used to store programs;
[0038] The processor is used to execute the program to implement the various steps of the multimedia content generation method as described above.
[0039] Fourthly, a storage medium is provided on which a computer program is stored, which, when executed by a processor, implements the various steps of the multimedia content generation method as described above.
[0040] Using the above technical solution, this application generates multimedia content corresponding to user creation needs by calling a large-scale artificial intelligence model. It supports users in submitting their creation requests, and then calls a pre-configured large-scale artificial intelligence model to parse these requests, extracting tag information of specific types, such as the creation platform being Xiaohongshu or WeChat, and the content type being product introduction or brand promotion. Considering that different tag information represents different user creation needs, and that the large-scale artificial intelligence model may need to rely on different knowledge prompts to better understand and generate multimedia content that meets these needs, this application pre-configures a set of prompt format templates. These templates match various tag types and contain knowledge information related to the matched tag types. Based on this, the application selects a matching target prompt format template from the set according to the tag information extracted from the user's creation needs. A prompt instruction is generated based on this target prompt format template and the user's creation needs and sent to the large-scale artificial intelligence model to obtain the generated multimedia content. Since the target prompt template contains knowledge information related to the tags extracted from the user's creative needs, the generated prompt also contains this knowledge information. After being fed into a large-scale artificial intelligence model, the model can better understand the user's creative needs based on this knowledge information, and thus generate multimedia content that better meets the user's creative needs for reference. Attached Figure Description
[0041] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0042] Figure 1 This is a schematic flowchart of a multimedia content generation method provided in an embodiment of this application;
[0043] Figure 2 This example illustrates a content creation interface for creating articles.
[0044] Figure 3 An example diagram of a content creation interactive interface with a text editing box is provided.
[0045] Figure 4 An example of a content creation interface for creating images is shown in the diagram.
[0046] Figure 5 This example illustrates a content creation interface for creating video scripts.
[0047] Figure 6 This is a schematic diagram of a multimedia content generation device provided in an embodiment of this application;
[0048] Figure 7 This is a schematic diagram of the structure of a multimedia content generation device provided in an embodiment of this application. Detailed Implementation
[0049] Before introducing the proposed solution, let's first explain the English terms used in this document:
[0050] Prompt: Instructions. When conversing with AI (such as large artificial intelligence models), you need to send instructions to the AI. These can be a text description, such as "Please recommend a popular song for me" when you talk to the AI, or a parameter description in a certain format, such as describing the relevant drawing parameters to ask the AI to draw a picture in a certain format.
[0051] Large-scale artificial intelligence models, also known as large-scale deep learning models, are artificial intelligence models based on deep learning technology. They consist of hundreds of millions of parameters and can perform complex tasks such as natural language processing, image recognition, and speech recognition through learning and training on massive amounts of data. Large-scale artificial intelligence models can include large models and large language models. Both large models and large language models refer to machine learning models with a very large number of parameters, but their application scenarios and focuses are slightly different. The definitions of the two are explained below.
[0052] Large-Scale Models (LSMs) typically refer to large-scale neural network models, such as deep neural networks and convolutional neural networks, which can have millions or even billions of parameters. These models are often used to process large-scale datasets, such as in image recognition and natural language processing. The advantage of large-scale models is that they can learn more complex features and patterns, thereby improving the model's accuracy and generalization ability.
[0053] Large language models (LLMs) are language models built using techniques such as probabilistic graphical models or recurrent neural networks, and they typically have a very large number of parameters. These models are commonly used for natural language processing tasks, such as machine translation, text generation, and question-answering systems. The advantage of large language models lies in their ability to understand and generate natural language, thereby enabling more intelligent interaction methods. Common large language models include GT4 and other large language models developed by various companies.
[0054] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0055] This application provides a multimedia content generation solution that can be applied to generate various types of multimedia content, such as article types, image types, and video script types. The generated multimedia content can be used by users for reference, such as secondary processing and editing based on the generated multimedia content, direct publication, or use as reference material.
[0056] The proposed solution can be implemented based on a terminal with data processing capabilities, such as a mobile phone, computer, learning machine, or intelligent robot.
[0057] Next, combined Figure 1 The multimedia content generation method of this application may include the following steps:
[0058] Step S100: Obtain the user's creative needs.
[0059] Specifically, this application supports users in creating various types of multimedia content, such as article types, image types, and video script types.
[0060] This embodiment can provide users with a content creation interactive interface, including an input box for creation requirements on the interactive page, for reference. Figure 2 , Figure 3As shown, an input box is provided at the bottom of the example's interactive interface, where users can enter their creative needs.
[0061] Creation requirements can include the platform the content belongs to, the content type, the theme, keywords that users are interested in, etc. For example... Figure 2 In the example, on the interactive interface, the user's creative requirements are: "I would like you to help me create copy for the Xiaohongshu platform in a lively style. My theme is cute pets, my target audience is college students and women, and the value I bring to users is emotional comfort."
[0062] Understandably, the more detailed the user's creative needs are, the more detailed the subsequent model can be based on this information to generate multimedia content that truly meets the user's needs. Considering that some users may not be able to describe their creative needs well, or their descriptions may not be comprehensive enough, this embodiment can also provide a question template space on the interactive interface, such as... Figure 2 The "Use Question Template" control at the bottom allows users to trigger one or more pre-set question templates to pop up on the screen, guiding them to describe their creative needs more comprehensively and accurately.
[0063] Step S110: Call the pre-configured large-scale artificial intelligence model to parse the creation requirements and extract the tag information of the set type contained in the creation requirements.
[0064] Specifically, users can describe their creative needs from different dimensions, and these descriptions can be categorized into several different types of tags. For example, tags under the "creation platform" category can include various multimedia platforms such as Xiaohongshu, WeChat, and Weibo; tags under the "content type" category can include product promotion, knowledge dissemination, product recommendations, etc. Of course, there can be other types of tags, and each type can have multiple tags.
[0065] Different tag information represents different user creative needs. Large-scale artificial intelligence models may need to rely on different knowledge prompts to better understand these needs and generate multimedia content that meets them. Therefore, this embodiment pre-designs a set of prompt format templates matching different tag types. To select the optimal prompt format template, it is necessary to extract the tag information of the specified types from the user's creative needs. Here, the specified types can be one or more, for example, the specified types may include the creation platform and content type.
[0066] This step involves calling a large-scale artificial intelligence model to analyze the user's creative needs and extract the tag information for the specified type.
[0067] This embodiment provides an optional model invocation method, specifically:
[0068] A pre-configured first prompt template can be obtained. This first prompt template includes user creation requirement slots. The first prompt template is used to instruct a large-scale artificial intelligence model to understand the user creation requirements within these slots and output the tags of various defined types contained therein. For example, the first prompt template may include:
[0069] "Please understand the user's creative needs and extract the following tags for each setting type: <tag1, tag2, tag3...>."
[0070] The user's creative requirements are: [User Creative Requirements Slot].
[0071] Furthermore, the acquired user creative requirements are filled into the user creative requirement slot to obtain the edited first prompt instruction, which is then sent to a large-scale artificial intelligence model to obtain the tag information of the set type contained in the creative requirements output by the model.
[0072] In this step, the user's creative needs can be in text form, and the large-scale artificial intelligence model invoked can be a large language model with natural language processing capabilities.
[0073] Step S120: Based on the extracted tag information of each type, select the matching target prompt format template from the preset prompt format template set.
[0074] The prompt format template set contains prompt format templates that match different type tags, and the prompt format templates contain knowledge information related to the matched type tags.
[0075] It should be noted that the prompt format templates matching different types of tags can specifically be: prompt format templates matching different tags under each type, or prompt format templates matching combinations of multiple tags under multiple types, as shown in Table 1 below:
[0076] Table 1
[0077]
[0078] Table 1 illustrates the matching of two tag types with prompt format templates, namely tag type 1 and tag type 2. Each type contains several different tags; for example, type 1 includes tag X1, tag X2, etc. Different tag combinations are matched with different prompt format templates.
[0079] When setting a prompt template, you can set the knowledge information within the template according to the matching tag information. Taking the prompt template corresponding to the tag "Xiaohongshu" under the platform type as an example, the knowledge information contained therein can be a description of the characteristics of the multimedia content of the "Xiaohongshu" platform.
[0080] Further optionally, considering that the proposed solution supports generating multimedia content of different content types, such as generating articles, images, and video scripts, separate subsets of prompt format templates can be established for each content type. These prompt format templates within the subsets instruct the large-scale artificial intelligence model to generate multimedia content of the content type corresponding to the subset. For example, a subset of prompt format templates can be established for the article type, instructing the large-scale artificial intelligence model to generate articles; a subset can be established for the image type, instructing the large-scale artificial intelligence model to generate images; and a subset can be established for the video script type, instructing the large-scale artificial intelligence model to generate video scripts.
[0081] Based on this, the specific implementation methods of this step may include:
[0082] The extracted tags of various types are combined to obtain tag groups. From the subset of prompt format templates corresponding to the type of content the user needs to create, a target prompt format template matching the tag group is selected. This subset of prompt format templates contains prompt format templates matching different tag groups.
[0083] Step S130: Generate a prompt instruction based on the target prompt format template and the user's creative needs, and input it into a large-scale artificial intelligence model to obtain multimedia content generated by the model that meets the user's creative needs.
[0084] In one alternative implementation, the target prompt format template and the user's creative requirements can be directly combined into a prompt instruction, which is then fed into a large-scale artificial intelligence model to obtain multimedia content output by the model.
[0085] The multimedia content generation method provided in this application generates multimedia content corresponding to user creation needs by calling a large-scale artificial intelligence model. It supports users submitting their own creation needs, and then calls a pre-configured large-scale artificial intelligence model to parse these needs, extracting tag information of specific types, such as the creation platform being Xiaohongshu or WeChat, and the content type being product introduction or brand promotion. Considering that different tag information represents different user creation needs, and that the large-scale artificial intelligence model may need to rely on different knowledge prompts to better understand and generate multimedia content that meets these needs, this application pre-configures a set of prompt format templates. These templates contain prompt format templates matching different types of tags, and include knowledge information related to the matched type tags. Based on this, a matching target prompt format template is selected from the set according to the tag information extracted from the user's creation needs. A prompt instruction is generated based on this target prompt format template and the user's creation needs and sent to the large-scale artificial intelligence model to obtain the generated multimedia content. Since the target prompt template contains knowledge information related to the tags extracted from the user's creative needs, the generated prompt also contains this knowledge information. After being fed into a large-scale artificial intelligence model, the model can better understand the user's creative needs based on this knowledge information, and thus generate multimedia content that better meets the user's creative needs for reference.
[0086] As previously mentioned, user creation requests can include any type of content creation such as articles, images, and video scripts. When a user requests to create an article, the large-scale artificial intelligence model invoked in step S130 can be a large language model with natural language processing capabilities, and the multimedia content generated by the model can be an article.
[0087] When a user requests to create an image, the large-scale artificial intelligence model called in step S130 can be a large model with the ability to generate images based on text, such as the Stable Diffusion model, and the multimedia content generated by the model can be an image.
[0088] When a user requests to create a video script, the large-scale artificial intelligence model called in step S130 can be a large language model with natural language processing capabilities, and the multimedia content generated by the model can be a video script.
[0089] Furthermore, when a user requests to create a video script, after generating the video script in step S130, a rendering process for the video script can be added, and the rendered video script can be output.
[0090] Reference Figure 2 When a user's creative request is, "I want you to help me create a lively copy for the Xiaohongshu platform. My theme is cute pets, my target audience tags are college students and women, and the value I bring to users is emotional comfort," the article generated by the large AI model displayed on the interactive interface will be: "Title: Cute Pets Heal the Heart, Heartwarming Daily Life on Xiaohongshu. I. Introduction: For many college students and women, pets are indispensable companions in life. They not only bring us joy but also silently accompany us when we are unhappy, becoming our emotional comfort. II. Cute Pets Heal the Heart: There are many cute pet daily life posts on Xiaohongshu, they always..." So cute and so well-behaved, they make you want to laugh out loud. These pets also bring their owners a lot of joy and warmth, enriching your life. Third, User Value: In our fast-paced lives, people often neglect their inner needs, but the daily lives of these adorable pets allow us to temporarily set aside our worries and appreciate the beauty of life. Therefore, paying attention to the daily lives of these cute pets is a wonderful way to heal. Fourth, Concluding Resonance: Finally, I want to say that pets are an important part of our lives. I hope everyone can find their own adorable pet to spend wonderful times together. #XiaohongshuAdorablePetsDailyLife #CollegeStudents #Women #SoulCare
[0091] Considering that users may have temporary editing needs after browsing the article generated by the model, this embodiment can also set a text editing window on the interactive interface. The text editing window can be collapsed, and when the user triggers the expand control (e.g., Figure 2 The gray control in the middle can display a text editing window on the right for users to edit text.
[0092] Furthermore, to facilitate user editing of the articles generated by the model, a text editing control can be added below the dialog box containing the generated article displayed on the interactive interface, such as... Figure 2 The example shows two controls: "Copy" and "Output to Right". Users can use the "Copy" control to copy part or all of the text in the dialog box, and use the "Output to Right" control to generate the text in the dialog box into the text editing window on the right with one click.
[0093] Reference Figure 4When a user's creative request is "Help me generate an image of a child holding an umbrella in the rain, with a building in front of him," the interactive interface displays an image generated by a large artificial intelligence model.
[0094] Reference Figure 5 When a user's creative request is "Help me write a video script for *notebook", the interactive interface displays the rendered result of a video script generated by a large artificial intelligence model. This includes a title and script content. The title is: "Easy on the go, excellent sound recording - *notebook experience sharing". The script content is rendered in table format, as shown in Table 2 below:
[0095] Table 2
[0096]
[0097] In some embodiments of this application, another method for generating multimedia content is further provided. Considering that the user's input of creative requirements may not be comprehensive, lacking certain types of tag information, such as missing style tags or audience tags for the multimedia content to be created, this embodiment can supplement the missing tags of the specified types in the creative requirements before calling a large-scale artificial intelligence model to generate the created multimedia content in the aforementioned step S130.
[0098] Specifically, in this embodiment, the user's account information and behavioral data on a multimedia platform can be combined to predict the specified type of tag. The user's account information on the multimedia platform includes, but is not limited to, account name, profile, and published content. The user's behavioral data on the multimedia platform includes, but is not limited to, the platforms the user frequently logs into, the types of multimedia content the user asks, searches, and publishes, and the user's actions on content published by others, such as liking, commenting, and saving. It is understood that the aforementioned user account information and behavioral data on the multimedia platform contain some personal characteristics of the user, such as style and preferences. Based on this information, tags of some types missing from the current creative requirement can be predicted. For example, the user's current creative requirement does not indicate the style of the article content to be created. However, by analyzing the style of the user's historically published articles and the styles of articles the user follows, it can be known that the user likes "humorous" articles; therefore, it can be predicted that the missing content style in the user's current creative requirement is "humorous."
[0099] Based on this, the process of generating the prompt instruction in step S130 may include:
[0100] The target prompt format template, the user's creative requirements, and the predicted tags of the specified type are concatenated into a prompt instruction.
[0101] In other words, when concatenating a prompt, the target prompt format template, the user's creative requirements, and the prediction results for the specified types of tags missing in the creative requirements can be concatenated together to obtain the prompt. This makes the information contained in the prompt more comprehensive and in line with the user's actual creative requirements. This allows for better guidance to large-scale artificial intelligence models to generate multimedia content that meets the user's creative needs.
[0102] Furthermore, considering the various scenarios in which users create multimedia content, different creative scenarios may require reliance on external contextual knowledge. For example, when creating an article on a Weibo platform, a user might need to incorporate trending topics to generate the article. Similarly, when creating a video script introducing a product, a user might need to combine it with contextual knowledge such as the product manual.
[0103] Therefore, in this embodiment, in addition to obtaining the user's creative needs, scene knowledge information related to the user's creative needs can be further obtained. Based on this, the implementation process of generating the prompt instruction in step S130 above can include:
[0104] The target prompt format template, the user's creative requirements, and the scene knowledge information are combined into a prompt instruction.
[0105] In other words, when splicing prompt instructions, scene knowledge information related to creative needs can be further incorporated into the prompt instructions so that the large-scale artificial intelligence model can generate multimedia content that matches the user's creative needs based on the scene knowledge information.
[0106] It is understood that the two steps S130 described in the above embodiments, the process of generating the prompt instruction, can be combined. That is, the target prompt format template, the user's creative needs, the predicted tags of the specified type, and the scene knowledge information can be concatenated into the prompt instruction.
[0107] The multimedia content generation apparatus provided in the embodiments of this application is described below. The multimedia content generation apparatus described below and the multimedia content generation method described above can be referred to in correspondence.
[0108] See Figure 6 , Figure 6 This is a schematic diagram of the structure of a multimedia content generation device disclosed in an embodiment of this application.
[0109] like Figure 6 As shown, the device may include:
[0110] The creative needs acquisition unit 11 is used to acquire users' creative needs;
[0111] The creation requirement parsing unit 12 is used to call a pre-configured large-scale artificial intelligence model to parse the creation requirement and extract the tag information of the set type contained in the creation requirement;
[0112] The template matching unit 13 is used to select a matching target prompt format template from a preset prompt format template set based on the extracted tag information of each type. The prompt format template set contains prompt format templates that match different types of tags, and the prompt format template contains knowledge information related to the matched type tags.
[0113] The model call generation unit 14 is used to generate a prompt instruction based on the target prompt format template and the user's creative needs, and input it into a large-scale artificial intelligence model to obtain multimedia content generated by the model that meets the user's creative needs.
[0114] Optionally, the user's creation requirements obtained by the aforementioned creation requirement acquisition unit include creation requirements for any type of content among articles, images, and video scripts. Correspondingly, the multimedia content generated by the large-scale artificial intelligence model obtained by the model call generation unit includes any type of content among articles, images, and video scripts.
[0115] Optionally, the apparatus of this application may further include:
[0116] The video script rendering unit is used to render and output the video script when the multimedia content generated by the large artificial intelligence model meets the user's creative needs.
[0117] Optionally, the process by which the aforementioned creative requirement parsing unit calls a pre-configured large-scale artificial intelligence model to parse the creative requirement and extract the tag information of a set type contained in the creative requirement may include:
[0118] Obtain a pre-configured first prompt format template, which includes a user creation requirement slot. The first prompt format template is used to instruct a large artificial intelligence model to understand the user creation requirements within the user creation requirement slot and output the labels of various set types contained therein.
[0119] The acquired user creation requirements are filled into the user creation requirement slot to obtain the edited first prompt instruction, which is then sent to a large-scale artificial intelligence model to obtain the tag information of the set type included in the creation requirements output by the model.
[0120] Optionally, the preset prompt format template set includes subsets of prompt format templates corresponding to different content types. These subsets of prompt format templates are used to instruct the large-scale artificial intelligence model to generate multimedia content corresponding to the content type of the subset. Based on this, the process by which the template matching unit selects a matching target prompt format template from the preset prompt format template set according to the extracted tag information for each type may include:
[0121] The extracted tags of various types are combined to obtain tag groups;
[0122] From the subset of prompt format templates corresponding to the type of content to be created according to the user's creative needs, a target prompt format template that matches the tag group is selected. The subset of prompt format templates contains prompt format templates that match different tag groups.
[0123] Optionally, the process by which the model call generation unit generates a prompt instruction based on the target prompt format template and the user's creative requirements may include:
[0124] The target prompt format template and the user's creative requirements are combined into a prompt instruction.
[0125] Optionally, the apparatus of this application may further include:
[0126] The missing tag prediction unit is used to predict the missing tags of a specified type in the user's creative requirements by combining the user's account information and behavioral data on the multimedia platform. Based on this, the process by which the model call generation unit generates a prompt instruction based on the target prompt format template and the user's creative requirements may include:
[0127] The target prompt format template, the user's creative requirements, and the predicted tags of the specified type are concatenated into a prompt instruction.
[0128] Optionally, the apparatus of this application may further include:
[0129] The scene knowledge information acquisition unit is used to acquire scene knowledge information related to the user's creative needs. Based on this, the process by which the aforementioned model call generation unit generates a prompt instruction based on the target prompt format template and the user's creative needs may include:
[0130] The target prompt format template, the user's creative requirements, and the scene knowledge information are combined into a prompt instruction.
[0131] Optionally, the user's creation needs obtained by the aforementioned creation needs acquisition unit include at least one of the following: the platform to which the creation content belongs, the content type, the theme, and keywords that the user is interested in.
[0132] The multimedia content generation apparatus provided in this application embodiment can be applied to multimedia content generation devices, such as mobile phones, computers, learning machines, and intelligent robots. Optionally, Figure 7 A hardware block diagram of a multimedia content generation device is shown, with reference to... Figure 7 The hardware structure of a multimedia content generation device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;
[0133] In this embodiment of the application, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4;
[0134] Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.
[0135] Memory 3 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;
[0136] The memory stores a program, which the processor can call. The program is used for:
[0137] Obtain users' creative needs;
[0138] The pre-configured large-scale artificial intelligence model is invoked to parse the creation requirements and extract the tag information of the specified type contained in the creation requirements;
[0139] Based on the extracted tag information of each type, a matching target prompt format template is selected from a preset prompt format template set. The prompt format template set contains prompt format templates that match different types of tags, and the prompt format templates contain knowledge information related to the matched type tags.
[0140] Based on the target prompt format template and the user's creative needs, a prompt instruction is generated and input into a large-scale artificial intelligence model to obtain multimedia content that meets the user's creative needs.
[0141] Optionally, the refined and extended functions of the program can be found in the description above.
[0142] This application embodiment also provides a storage medium that can store a program suitable for execution by a processor, the program being used for:
[0143] Obtain users' creative needs;
[0144] The pre-configured large-scale artificial intelligence model is invoked to parse the creation requirements and extract the tag information of the specified type contained in the creation requirements;
[0145] Based on the extracted tag information of each type, a matching target prompt format template is selected from a preset prompt format template set. The prompt format template set contains prompt format templates that match different types of tags, and the prompt format templates contain knowledge information related to the matched type tags.
[0146] Based on the target prompt format template and the user's creative needs, a prompt instruction is generated and input into a large-scale artificial intelligence model to obtain multimedia content that meets the user's creative needs.
[0147] Optionally, the refined and extended functions of the program can be found in the description above.
[0148] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0149] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.
[0150] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating multimedia content, characterized in that, include: Obtain users' creative needs; The pre-configured large-scale artificial intelligence model is invoked to parse the creation requirements and extract the tag information of the specified type contained in the creation requirements; Based on the extracted tag information of each type, a matching target prompt format template is selected from a preset prompt format template set. The prompt format template set contains prompt format templates that match different types of tags, and the prompt format templates contain knowledge information related to the matched type tags. Based on the target prompt format template and the user's creative needs, a prompt instruction is generated and input into a large-scale artificial intelligence model to obtain multimedia content generated by the model that meets the user's creative needs. The preset set of prompt format templates includes subsets of prompt format templates corresponding to different content types. The prompt format templates in the subsets are used to instruct large artificial intelligence models to generate multimedia content of the content type corresponding to the subset. Based on the extracted tag information of various types, a matching target prompt format template is selected from the preset prompt format template set, including: The extracted tags of various types are combined to obtain tag groups; From the subset of prompt format templates corresponding to the type of content to be created according to the user's creative needs, a target prompt format template that matches the tag group is selected. The subset of prompt format templates contains prompt format templates that match different tag groups.
2. The method according to claim 1, characterized in that, Users' creative needs include the creation of any type of content among articles, images, and video scripts. Correspondingly, the multimedia content generated by large-scale artificial intelligence models includes any type of content among articles, images, and video scripts.
3. The method according to claim 2, characterized in that, When the multimedia content generated by a large-scale artificial intelligence model that meets the user's creative needs is a video script, the method also includes: Render the video script and output it.
4. The method according to claim 1, characterized in that, The pre-configured large-scale artificial intelligence model is invoked to parse the creation requirements and extract the tag information of the specified type contained in the creation requirements, including: Obtain a pre-configured first prompt format template, which includes a user creation requirement slot. The first prompt format template is used to instruct a large artificial intelligence model to understand the user creation requirements within the user creation requirement slot and output the labels of various set types contained therein. The acquired user creation requirements are filled into the user creation requirement slot to obtain the edited first prompt instruction, which is then sent to a large-scale artificial intelligence model to obtain the tag information of the set type included in the creation requirements output by the model.
5. The method according to claim 1, characterized in that, Based on the target prompt format template and the user's creative requirements, a prompt instruction is generated, including: The target prompt format template and the user's creative requirements are combined into a prompt instruction.
6. The method according to any one of claims 1-4, characterized in that, Before generating the prompt instruction based on the target prompt format template and the user's creative requirements, the method further includes: For the specified type of tags missing in the user's creative needs, the specified type of tags are predicted by combining the user's account information and behavioral data on the multimedia platform; The process of generating a prompt instruction based on the target prompt format template and the user's creative requirements includes: The target prompt format template, the user's creative requirements, and the predicted tags of the specified type are concatenated into a prompt instruction.
7. The method according to any one of claims 1-4, characterized in that, Before generating the prompt instruction based on the target prompt format template and the user's creative requirements, the method further includes: Acquire scenario knowledge information related to users' creative needs; The process of generating a prompt instruction based on the target prompt format template and the user's creative requirements includes: The target prompt format template, the user's creative requirements, and the scene knowledge information are combined into a prompt instruction.
8. The method according to any one of claims 1-4, characterized in that, The user's creative requirements include at least one of the following: the platform to which the content belongs, the content type, the theme, and keywords that the user is interested in.
9. A multimedia content generation device, characterized in that, include: The creative needs acquisition unit is used to acquire users' creative needs; The creation requirement parsing unit is used to call a pre-configured large-scale artificial intelligence model to parse the creation requirement and extract the tag information of the set type contained in the creation requirement; The template matching unit is used to select a target prompt format template from a preset prompt format template set based on the extracted tag information of each type. The prompt format template set contains prompt format templates that match different types of tags, and the prompt format templates contain knowledge information related to the matched type tags. The model call generation unit is used to generate a prompt instruction based on the target prompt format template and the user's creative needs, and input it into a large-scale artificial intelligence model to obtain multimedia content generated by the model that meets the user's creative needs. The preset set of prompt format templates includes subsets of prompt format templates corresponding to different content types. The prompt format templates in the subsets are used to instruct large artificial intelligence models to generate multimedia content of the content type corresponding to the subset. The template matching unit selects a target prompt format template from a preset set of prompt format templates based on the extracted tag information of various types, including: The extracted tags of various types are combined to obtain tag groups; From the subset of prompt format templates corresponding to the type of content to be created according to the user's creative needs, a target prompt format template that matches the tag group is selected. The subset of prompt format templates contains prompt format templates that match different tag groups.
10. A multimedia content generation device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the multimedia content generation method as described in any one of claims 1 to 8.
11. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the multimedia content generation method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Multimedia file generation method, device and equipment and computer readable storage medium
CN110826080A
Generation method and device, processing method and device, electronic equipment and medium
CN113095056A