Multimedia data generation method and device
By acquiring project topics and marketing information, analyzing multimedia data to generate themes, and using large language models to automatically generate multimedia content, the problem of low efficiency and poor consistency in multimedia content production in existing technologies has been solved, achieving efficient and accurate multimedia content generation.
Patent Information
- Application Number
- CN202511674182.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-02-17
AI Technical Summary
The current production of multimedia content relies on manual production, which leads to cumbersome processes and long cycles, making it difficult to meet the demand for massive amounts of content. Furthermore, the lack of unified standards and collaborative mechanisms results in inconsistent visual styles and deviations in the transmission of core information, affecting the consistency and professionalism of the dissemination effect.
By acquiring the project topics, marketing information, and multimedia data of the target project, determining the multimedia generation theme, analyzing multimedia reference information, generating target multimedia data, and using a large language model and multimedia generation device to achieve automated and intelligent multimedia content generation.
It enables intelligent and batch generation of multimedia content, improves the project relevance and content appeal of the generated content, shortens the content production cycle, reduces reliance on manual design, and improves dissemination efficiency.
Smart Images

Figure CN121547656A_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of computer technology, and in particular to a multimedia data generation method and a multimedia data generation apparatus. Background Technology
[0002] With the rapid development of information technology and digital media, multimedia content has become an indispensable core support for the promotion and dissemination of various projects. Whether it is product promotion, event promotion, or brand building, rich media forms such as images and videos, with their intuitiveness, appeal, and high transmissibility, have become key carriers of information transmission.
[0003] However, the production of multimedia content still relies heavily on manual labor. While this method can meet basic needs, it is generally plagued by cumbersome processes and lengthy cycles. Especially when faced with massive content demands or frequent revisions, the traditional model often suffers from slow response times and struggles to guarantee delivery efficiency. Furthermore, the lack of unified standards and collaborative mechanisms easily leads to inconsistent visual styles and deviations in the delivery of core information, affecting not only the consistency and professionalism of the dissemination effect but also resulting in inefficient consumption of human and time resources.
[0004] Therefore, there is an urgent need for a more efficient method for generating multimedia data to improve overall dissemination effectiveness and resource utilization efficiency. Summary of the Invention
[0005] In view of the above, embodiments of this specification provide a multimedia data generation method. One or more embodiments of this specification also relate to a multimedia data generation apparatus, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0006] According to a first aspect of the embodiments of this specification, a multimedia data generation method is provided, comprising:
[0007] Acquire the target project's project topics, marketing information, and corresponding multimedia data.
[0008] Based on the project topic and marketing information, determine the multimedia generation theme, which constrains the direction of multimedia data generation for the target project.
[0009] Analyze multimedia data to obtain multimedia reference information;
[0010] Based on the multimedia generation theme and multimedia reference information, target multimedia data corresponding to the target project is generated.
[0011] According to a second aspect of the embodiments of this specification, a multimedia data generation apparatus is provided, comprising:
[0012] The first acquisition module is configured to acquire the project topic, project marketing information and multimedia data corresponding to the project topic of the target project.
[0013] The determination module is configured to determine a multimedia generation theme based on the project topic and the project marketing information, wherein the multimedia generation theme is used to constrain the direction of multimedia data generation for the target project;
[0014] The parsing module is configured to parse the multimedia data to obtain multimedia reference information;
[0015] The generation module is configured to generate target multimedia data corresponding to the target project based on the multimedia generation theme and the multimedia reference information.
[0016] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising:
[0017] Memory and processor;
[0018] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the multimedia data generation method described above.
[0019] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the multimedia data generation method described above.
[0020] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the multimedia data generation method described above.
[0021] The video generation method provided in one or more embodiments of this specification acquires the project topic, project marketing information, and multimedia data corresponding to the project topic of a target project; determines the multimedia generation theme based on the project topic and project marketing information, wherein the multimedia generation theme is used to constrain the direction of multimedia data generation for the target project; parses the multimedia data to obtain multimedia reference information; and generates target multimedia data corresponding to the target project based on the multimedia generation theme and multimedia reference information. By accurately extracting a guiding multimedia generation theme, the direction and tone of subsequent content generation are effectively constrained, ensuring a high degree of consistency between the target multimedia data and the project marketing information; through in-depth analysis of the original multimedia data, multimedia reference information is extracted, constraining the generation direction of the generation process; and under the dual constraints of the semantic guidance of the multimedia generation theme and the multimedia reference information, the target multimedia data is automatically generated. This significantly improves the project relevance and content attractiveness of the generated content, while greatly reducing reliance on manual design, shortening the content production cycle, and improving response efficiency. It realizes the intelligent and batch generation of multimedia content and can be widely applied to practical business scenarios such as advertising, social media operation, event promotion, and brand digital construction, demonstrating good practical value and promotion prospects. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating a multimedia data generation method provided in one embodiment of this specification;
[0023] Figure 2 This is an interface diagram of a topic retrieval and classification result provided in one embodiment of this specification;
[0024] Figure 3 This is an interface diagram of another topic retrieval and classification result provided in one embodiment of this specification;
[0025] Figure 4 This is an interface diagram of a multimedia data reference provided in one embodiment of this specification;
[0026] Figure 5 This is an interface diagram of a multimedia theme generation method provided in one embodiment of this specification;
[0027] Figure 6 This is a schematic diagram of a multimedia reference information provided in one embodiment of this specification;
[0028] Figure 7 This is a flowchart illustrating a target audio generation process according to one embodiment of this specification;
[0029] Figure 8 This is a diagram showing a target audio interface provided in one embodiment of this specification;
[0030] Figure 9 This is a diagram illustrating a display interface of a reference digital human, provided in one embodiment of this specification.
[0031] Figure 10 This is a flowchart illustrating a target video generation process provided in one embodiment of this specification;
[0032] Figure 11 This is a flowchart illustrating an audio text generation process provided in one embodiment of this specification;
[0033] Figure 12 This is a schematic diagram of an audio text provided in one embodiment of this specification;
[0034] Figure 13 This is a diagram showing a target video display interface provided in one embodiment of this specification;
[0035] Figure 14 This is a flowchart illustrating the processing steps of a multimedia data generation method provided in one embodiment of this specification.
[0036] Figure 15 This is a flowchart illustrating a video script creation process provided in one embodiment of this specification;
[0037] Figure 16 This is a flowchart illustrating a target video creation process according to one embodiment of this specification;
[0038] Figure 17 This is a schematic diagram of the structure of a multimedia data generation device provided in one embodiment of this specification;
[0039] Figure 18 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0040] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0041] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items. The term “at least one” in one or more embodiments of this application means “one or more,” and “a plurality of” means “two or more.” The term “comprising” is an open-ended description and should be understood as “including but not limiting,” and may include other content in addition to what has been described.
[0042] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0043] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0044] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0045] n-gram: A statistical language model that predicts the probability of a word appearing based on the preceding n-1 words.
[0046] A Recurrent Neural Network (RNN) is a neural network model that uses a recurrent structure to process sequential data. At each time step, it receives the current input and the hidden state from the previous time step, outputs the current result, and updates the state. The hidden state can theoretically carry all previous information.
[0047] SubRip subtitle file (.srt): .srt is a commonly used subtitle file format used to display subtitles during video playback. It contains the subtitle text and its timecode information, specifying when the subtitles appear and disappear. .srt files are stored in plain text format and can be opened and edited using any text editor.
[0048] FFmpeg technology: FFmpeg is an open-source, powerful multimedia processing toolset capable of handling various media formats such as audio, video, and subtitles. It is widely used in scenarios such as audio and video transcoding, editing, streaming media transmission, recording, filtering, and format conversion, and is one of the core tools in the modern digital media technology stack.
[0049] Publishing content to the Douyin platform is one of the important means for internet products to expand traffic and promote sales growth, both now and in the future. However, currently, it requires manual content creation by business personnel, collection of a large number of excellent cases for reference, and manual booking of anchors and shooting of videos. The work is tedious and cannot be accumulated. This requires a lot of time investment and is not efficient. Therefore, it is necessary to create high-quality digital human videos efficiently to attract traffic and expand sales revenue.
[0050] Because manual content creation for the product requires scheduling livestreamers to film and edit the videos, and this process is limited by factors such as individual expertise, work schedule, and available time. For example, producing dozens of product-related videos per week requires significant time to develop a thesis on topics relevant to the product, and other urgent matters can disrupt the flow of thought. Therefore, as the amount of content to be released increases in the future, with multiple products or hundreds of videos per week, the manual input becomes extremely costly, inefficient, and time-consuming.
[0051] To address the shortcomings of existing content creation methods, this specification's embodiments identify suitable marketing hot topics by searching for and categorizing them. These tagged hot topics are then used to find viral content, which is then broken down and analyzed to form an overall copywriting framework and provide writing suggestions. Combined with a creative outline outlining the writing style for the introduction, body, and conclusion, the copy is created. A large-scale model is then used to evaluate and automatically modify the copy based on product characteristics and prohibited words across various platforms, resulting in high-quality copy. Speech synthesis technology is used to create an audio file from the copy, which is then combined with extracted and recorded digital human speaking videos. Speech recognition technology is used to obtain SRT subtitle files from the audio files, ultimately synthesizing a marketing video with subtitles.
[0052] See Figure 1 , Figure 1This is a flowchart of a multimedia data generation method provided in one embodiment of this specification, which specifically includes the following steps:
[0053] Step 102: Obtain the target project's topic, marketing information, and multimedia data corresponding to the topic.
[0054] It's important to clarify that the target project refers to the specific content that needs to be promoted. The target project can be a product, activity, service, application, etc. The target project determines the purpose and direction of content generation; all generated multimedia content is customized around the needs of the target project, thereby achieving efficient and precise brand exposure and market promotion.
[0055] Project topics refer to trending events, popular trends, or issues of public concern that are currently receiving widespread attention and discussion online (especially on multimedia publishing platforms) and are related to the target project. Project topics can be trending hashtags on social media, news events or hot social topics, popular challenges or memes on platforms, cyclical or seasonal themes related to the project category, etc., such as: #AutumnOutfits, #AITechnology, a new technology release, a certain video shooting style, the journey home for Chinese New Year, etc. Project topics can be obtained from various multimedia publishing platforms, such as online platforms, applications, browsers, and other channels.
[0056] Multimedia data corresponding to a project topic refers to popular multimedia content that has been published on the internet, is highly relevant to the selected project topic, and has been proven successful by the market. Multimedia data can take various forms, including video, audio, images, and text. Multimedia data corresponding to a project topic serves as core reference material and a source of inspiration for generating high-quality new content. Its main function is to extract the successful elements that constitute a "viral hit" after analysis.
[0057] By analyzing the narrative structure, rhythm, and opening style of multimedia data corresponding to project topics, we can learn from the popular phrases, expressions, and emotional tone used, extract the language style, understand the core viewpoints conveyed, and understand why they resonate with users. Combining these key elements extracted from multimedia data with the marketing information of the target project can guide the generation of target multimedia data. This will generate target multimedia data that not only conforms to hot trends but also promotes the product itself, thereby greatly improving the quality and popularity of new content.
[0058] Project marketing information refers to the core information and promotional appeals that the target project wants to communicate to the outside world, such as the core selling points and functions of the product, the brand's slogan or philosophy, the specific rules and offers of marketing activities, and the unique advantages or usage scenarios of the service.
[0059] In practical applications, there are multiple ways to obtain the project topics, marketing information, and multimedia data corresponding to the project topics of a target project. One possibility is to use automated scraping tools to obtain the project topics, marketing information, and multimedia data corresponding to the project topics from a multimedia publishing platform. Another possibility is that analysts can obtain the project topics, marketing information, and multimedia data corresponding to the project topics of a target project from a database based on the target project. The specific method depends on the actual situation, and this specification does not limit this method in the embodiments.
[0060] This step provides a precise direction for content generation by simultaneously acquiring the target project's marketing needs, relevant trending topics, and popular content references under those topics, effectively ensuring the quality and commercial conversion value of the content.
[0061] In one optional embodiment of this specification, obtaining the project topic, project marketing information, and multimedia data corresponding to the project topic of the target project may include the following steps:
[0062] Obtain project-related topics and marketing information for the target project;
[0063] Multimedia data corresponding to the project topic can be retrieved from the multimedia publishing platform.
[0064] It should be noted that a multimedia publishing platform refers to an online internet platform or social media platform that allows users to upload, share, and consume multimedia content. Multimedia publishing platforms can take the form of mainstream short video platforms, comprehensive social media platforms, lifestyle sharing communities, news aggregation platforms, or information flow platforms.
[0065] Retrieval refers to the act of searching, filtering, and retrieving specific information from massive data sources. Retrieval is typically an automated process based on keywords, tags, categories, or algorithmic recommendations, but it can also be done manually, depending on the specific circumstances. Retrieval can be performed by the system based on user needs or instructions, or it can be proactively performed according to pre-defined procedures or functions.
[0066] See examples Figure 2 , Figure 2This is an interface diagram of a topic retrieval and classification result provided in one embodiment of this specification. After obtaining the "project topic," the system further analyzes popular videos related to that topic. For example, under the hot topic "Four days on, five days off next week," the system identifies a popular video titled "Latest Holiday Notice: Five-Day Holiday! #MayDayHoliday #LaborDay #FiveDayHoliday #Xuzhou" as a reference, which has received 229 likes, showing certain dissemination potential. The system provides two operation options: "View Original Video" and "View AI Imitation." The "View AI Imitation" function allows users to preview how the system can automatically generate a new promotional video script or copy with a similar style but different content based on the structure, language, and emotion of this popular video, combined with the marketing information of the target project (such as a tourism product). This not only intuitively demonstrates the system's ability to learn and imitate popular content, but also provides a clear reference and expected results for subsequent content creation.
[0067] By applying the solutions described in this specification, the system can efficiently combine external traffic trends with internal business needs by proactively retrieving and filtering popular topics and their corresponding viral content that match the target project from multimedia publishing platforms. This ensures that content creation closely aligns with user interests, effectively guaranteeing the quality, appeal, and dissemination potential of the produced content.
[0068] An optional embodiment of this specification, which obtains project topics and project marketing information of a target project, may include the following steps: obtaining project marketing information of the target project; reading multiple candidate topics from a multimedia publishing platform and identifying the topic categories of the multiple candidate topics; and filtering out project topics that match the target project category from the multiple candidate topics based on the topic categories.
[0069] It should be noted that candidate topics refer to the set of issues initially retrieved by the system from multimedia publishing platforms. Candidate topics are a series of potential trending topics that are currently popular or widely discussed. Candidate topics can be selected from various types of content from multimedia publishing platforms, such as trending search lists on social media platforms, popular challenge tags on short video platforms, and highly interactive content themes appearing in the platform's recommended information feed.
[0070] Topic categories refer to a general classification of the content areas or thematic ranges involved in candidate topics or project topics. Topic categories are tags used to describe the attributes of a topic, helping the system or users understand which field a topic belongs to. Common topic categories include entertainment, technology, sports, finance, education, automobiles, beauty, travel, food, and social news.
[0071] For example, the system can use automated content scraping tools (such as RPA) to read multiple topics from a multimedia publishing platform and categorize and sort these topics by popularity. As shown in the figure, the system may obtain a list of hot topics containing "hourly rankings" or "daily rankings," where each entry corresponds to information about a candidate topic. Each entry includes the topic's serial number, specific topic content (e.g., "How fun is the Water Splashing Festival in Y location?"), current popularity score (e.g., 10,236,149), and a link to inspire creative content. The system will identify the "topic category" to which these candidate topics belong (e.g., insurance marketing, health science popularization, etc.) and categorize each candidate topic into the corresponding set according to the topic category, providing a reference for the system or users. See details below. Figure 3 , Figure 3 This is an interface diagram of another topic search and classification result provided in one embodiment of this specification.
[0072] The system will then categorize multiple candidate topics based on their topic categories and automatically filter the corresponding candidate topics under the "Insurance Marketing" topic category tag. For example... Figure 3 As shown, the system has successfully categorized the original massive number of candidate topics, focusing on candidate topics strongly related to "insurance marketing", such as "four days on, five days off next week" and "the tenth National Security Education Day".
[0073] There are several ways to obtain target projects and their marketing information. One possibility is that users can upload target projects and their marketing information through a preset interface. Another possibility is that the system can automatically obtain the target projects to be published and their marketing information.
[0074] There are several ways to read multiple candidate topics from a multimedia publishing platform and identify their topic categories. One possible approach is to read multiple candidate topics from the platform and identify their categories using a large language model. Another possible approach is to read multiple candidate topics from the platform, perform optical character recognition (OCR), and classify them based on the OCR results to obtain their topic categories.
[0075] Users can select from categorized candidate topics, such as... Figure 4 , Figure 4This is an interface diagram of a multimedia data reference provided in one embodiment of this specification. The diagram displays the video title, such as "Latest Holiday Notice: Five-Day Holiday! #MayDayHoliday #LaborDay #FiveDayHoliday #Xuzhou", indicating that the video revolves around the hot topic of the "May Day holiday". Likes: Displays the video's interaction data, here "229", reflecting its popularity and dissemination potential on the platform. Operations: Provides two interactive options: "View Original Video" and "View AI Imitation". The former allows users to directly watch the original viral video for inspiration; the latter demonstrates how the system automatically generates a new promotional copy or script with a similar style but different content based on the video's structure, language style, and other elements.
[0076] By applying the solutions described in the embodiments of this specification, the system eliminates interference from irrelevant topics through this precise screening process, ensuring that the selected project topics are highly aligned with the target project in terms of theme and audience. This provides solid data support and creative references for subsequently generating promotional content that both conforms to market trends and accurately conveys the value of insurance products.
[0077] Step 104: Based on the project topic and project marketing information, determine the multimedia generation theme, which is used to constrain the direction of multimedia data generation for the target project.
[0078] A multimedia generated theme refers to a core idea or central theme determined during the creation of multimedia content, based on the project's objectives, marketing message, and the characteristics of the selected topic. It can also serve as a source of creative inspiration. A multimedia generated theme ensures that the final content effectively conveys the intended message and resonates with the target audience.
[0079] In practical applications, there are several ways to determine the multimedia generation theme based on the project topic and project marketing information. One possible scenario is that the system uses a rule extraction tool to determine the multimedia generation theme based on the project topic and project marketing information.
[0080] Another possibility is to use prompt words to guide a pre-trained large language model to determine the multimedia generation topic based on the project topic and project marketing information.
[0081] For example, if the target project is the promotion of a new health insurance product, the multimedia-generated theme might focus on how a "four-day work week followed by five-day weekend" work pattern integrates with the product, emphasizing how the insurance product provides security for users under this new work model. See also... Figure 5 , Figure 5 This is an interface diagram of a multimedia theme generation method provided in one embodiment of this specification.
[0082] After analyzing project topics (such as "4 days on, 5 days off"), the system not only provides trending information but also automatically generates commercially valuable creative inspiration based on insights into social trends and user needs. For example, regarding the phenomenon that the "4 days on, 5 days off" work system may change people's travel habits, the system recommends creative directions for insurance companies: launching new products such as flexible travel insurance and holiday insurance to meet users' needs for travel safety and property protection during long holidays. This suggestion closely integrates trending events with specific industry (insurance) marketing opportunities, providing content creators with clear and feasible creative ideas, effectively guiding them to transform market trends into actual product promotion strategies, and achieving an intelligent upgrade from trending topic capture to commercial insight.
[0083] This step organically integrates project topics with project marketing information to generate a clear multimedia theme. This theme serves as the core direction for content creation, setting clear and specific creative directions and boundaries for subsequent content production. On the one hand, it ensures that the final video or audio output leverages trending topics and attracts traffic; on the other hand, it firmly focuses the content on the product's core selling points, preventing content from deviating or becoming hollow. This mechanism successfully combines market sensitivity with precise brand promotion, ensuring that the automatically generated content has high dissemination potential and effectively serves actual marketing goals, improving content quality and commercial conversion efficiency.
[0084] Step 106: Parse the multimedia data to obtain multimedia reference information.
[0085] Multimedia reference information refers to the multi-dimensional features and patterns extracted from multimedia data corresponding to a project topic through analysis. Multimedia reference information contains the essential elements that constitute successful multimedia data, representing a deep understanding and structured extraction of this data. Multimedia data can indicate the generation process, structure, and text style of each part of the target multimedia data, etc.
[0086] There are several ways to parse multimedia data and obtain multimedia reference information. One possibility is to parse the multimedia data, generate a multimedia script, and then analyze the script to obtain the multimedia reference information. Another possibility is to parse the multimedia data using a pre-trained large language model to obtain the multimedia reference information. The specific method chosen depends on the actual situation, and this specification does not limit the specific approach used in the embodiments.
[0087] In one possible embodiment of this specification, parsing multimedia data to obtain multimedia reference information may include the following steps;
[0088] The multimedia data is parsed to obtain the multimedia script; the multimedia script input information and prompt information are used to generate a model to obtain multimedia reference information. The prompt information is used to guide the information generation model to analyze the multimedia script from the target dimension and generate multimedia reference information. The target dimension includes at least one of the following: structural dimension, language dimension, emotion dimension, viewpoint dimension, value dimension, and rhetoric dimension.
[0089] It's important to note that a multimedia script refers to the textual carrier corresponding to multimedia data. It meticulously records all the information to be presented, the presentation methods, and the timing of presentation. It serves as the intermediary bridge, translating creative ideas and themes into concrete audiovisual language. A multimedia script can include dialogue, visual descriptions of shots, their duration, transition methods, and more. Parsing multimedia data to obtain the multimedia script allows for the establishment of a correspondence between visual content and the information conveyed, enabling the system to "understand" the essential logic of the multimedia data and thus learn its inherent creative principles.
[0090] Generate prompts are a pre-defined set of instructional texts or parameters, often referred to as "cue words," used to guide the analysis of reference multimedia data to obtain multimedia reference information. They instruct the information generation model to analyze and interpret the multimedia script from a specific perspective. Generate prompts can be a set of explicit analytical dimension instructions, such as: "Analyze the following script from a 'structural dimension,' identifying its constituent paragraphs (e.g., opening, conflict, solution)." Specific question lists, such as: "What rhetorical devices are used in this script?" or "Is the overall sentiment positive, negative, or neutral?" Structured tagging or categorization requests, such as: "Label this script with applicable viewpoint tags."
[0091] Generating prompts can analyze multimedia scripts from multiple dimensions, such as at least one of the following: structural dimension, language dimension, emotional dimension, viewpoint dimension, value dimension, and rhetorical dimension. Among them, the structural dimension refers to the overall narrative framework and organizational logic of multimedia content (especially video scripts), which is used to analyze how multimedia data attracts attention at the beginning, how it develops the plot, how it creates a climax, and how it ends, thereby providing a skeleton template for new content.
[0092] The language dimension refers to the specific vocabulary, sentence structure, grammar, and expression style used in the content. It is used to analyze whether the vocabulary used in multimedia data is professional and rigorous or colloquial and conversational, whether the sentence structure is concise and powerful or vivid and complex, and whether specific popular words or internet slang are used, so as to understand and imitate the linguistic appeal and audience relatability of viral content.
[0093] The emotional dimension refers to the main emotions and sentiments that the content aims to convey or evoke. It is used to determine whether the content is humorous, touching, exciting, suspenseful, or professional and credible, and it is crucial to create the same emotional resonance in new content.
[0094] The viewpoint dimension refers to the core argument, stance, or unique insight that the content expresses. It is used to analyze what multimedia data thinks about a certain topic, whether it supports, opposes, or proposes a new perspective. It can help the system understand the thought logic of viral content, thereby integrating a clear viewpoint orientation into the target multimedia data.
[0095] The value dimension refers to the core value and meaning that content provides to the audience. It is used to judge whether the value of content is to provide practical information, bring entertainment, evoke emotional resonance, or provide decision-making reference. The value dimension determines the fundamental purpose and source of attraction of content.
[0096] The rhetorical dimension refers to the various techniques and methods used in the content to enhance the expressive effect. It is used to analyze whether it uses linguistic rhetoric such as metaphor, hyperbole, and parallelism, or audiovisual rhetoric such as fast-paced editing, specific sound effects, and visual impact, revealing the innovative points and eye-catching techniques of the content in terms of form.
[0097] For example, see Figure 6 , Figure 6 This is a schematic diagram of multimedia reference information provided in one embodiment of this specification. As shown in the figure, after analyzing viral videos, the system not only provides content references but also delves into the elements of their success, outputting creative guidance in the form of structured and actionable imitation suggestions. For example, for sports event videos, the system extracts two core writing guidelines: first, structural layout, suggesting a three-part structure of "concise opening – key plot progression – summary and sublimation," enhancing tension and suspense by describing key moments of the game; second, language style, emphasizing the use of concise, powerful, and infectious short sentences, avoiding redundancy, and recommending the use of impactful words such as "sprint," "leading," and "reversal." These specific and practical suggestions provide creators with clear writing templates and language paradigms, helping them quickly master the expression techniques of viral content, thereby efficiently generating high-quality original copy.
[0098] The solution implemented in this specification parses multimedia data into analyzable "multimedia scripts," enabling the system to obtain the foundational text for deep semantic understanding. Generated prompts serve as guiding instructions, driving the AI model to purposefully analyze the script from at least one preset "target dimension," such as "structure," "language," or "emotion." This ensures that the extracted "multimedia reference information" is high-quality, structured, and highly relevant. This allows the system to accurately capture the reasons for the popularity of trending content and apply these refined reasons to the generation of new content. Thus, while ensuring the quality of content creation, it achieves efficient replication and innovative application of market trends, thereby realizing intelligent and precise deconstruction and learning of "viral" content.
[0099] Step 108: Based on the multimedia generation theme and multimedia reference information, generate the target multimedia data corresponding to the target project.
[0100] It should be noted that target multimedia data refers to the final generated multimedia content used for the publicity and promotion of the target project. Target multimedia data integrates project marketing information, project topics, and multimedia reference information to form a complete, publishable promotional work. Target multimedia data can be text, audio, images, or video, such as press releases, posters, or multimedia videos, depending on the specific circumstances. This specification does not limit the specific format of the examples.
[0101] In practical applications, there are several ways to generate target multimedia data corresponding to a target project based on multimedia generation topics and multimedia reference information. One possibility is to input the multimedia generation topics and multimedia reference information into a text generation model to obtain multimedia generated text corresponding to the target project. Another possibility is to display the multimedia generation topics and multimedia reference information to the user through an interactive interface, allowing the user to determine the target multimedia data corresponding to the target project.
[0102] In one possible embodiment of this specification, generating target multimedia data corresponding to a target project based on a multimedia generation theme and multimedia reference information may include the following steps: inputting the multimedia generation theme and multimedia reference information into a text generation model to obtain multimedia generated text corresponding to the target project; and generating target multimedia data corresponding to the target project based on the multimedia generated text.
[0103] It should be noted that a text generation model refers to an artificial intelligence-based algorithm system whose core function is to automatically generate coherent, natural text content that meets specific requirements based on input instructions or contextual information. Text generation models can be large language models that rely on n-grams, RNNs, LSTMs, or Transformers.
[0104] Multimedia generated text (MGM) refers to the core text script used to create the final multimedia content. Automatically generated by a text generation model based on the multimedia theme and reference information, MGM is a crucial intermediate product connecting pre-production planning and post-production. It represents a concentrated reflection of the preliminary analysis results and provides a complete reference for the subsequent generation of target multimedia data. MGM can be a complete video script, containing all necessary narration, dialogue, or voice-over. It can also be a structured storyboard description, including not only dialogue but also simple instructions on visuals, rhythm, or mood. Alternatively, it can be a clean text for audio synthesis, ready to be converted into speech.
[0105] There are several ways to generate target multimedia data corresponding to a target item based on multimedia-generated text. One possibility is to generate the target multimedia data based on a pre-trained multimodal large language model. Another possibility is to generate the target multimedia data based on digital human generation technology. The specific method depends on the actual situation, and the embodiments in this specification do not limit this approach.
[0106] By applying the solutions described in this specification, and inputting both the multimedia generation theme and multimedia reference information into the text generation model, a multimedia generated text—i.e., a draft script for video or audio—that meets project requirements and possesses high dissemination potential is accurately generated. Based on the generated text, the system can automatically execute all subsequent complex target video creation processes. This two-step text-to-video approach enables interpretable target video generation, ensuring consistency in the generated content's thematic and stylistic appeal, thereby achieving large-scale, high-quality output of trending content.
[0107] In one possible implementation of this specification, the target multimedia data includes target audio; generating target multimedia data corresponding to the target project based on multimedia generated text may include the following steps: obtaining audio requirement information, wherein the audio requirement information includes at least one of timbre, speech rate, and temperature; inputting the audio requirement information and multimedia generated text into an audio synthesis model to obtain the target audio corresponding to the target project.
[0108] It's important to note that audio requirements information refers to the specific requirements or parameter settings for the sound characteristics and expressive style of the synthesized speech when generating the target audio. It guides the audio synthesis model on how to convert text into speech that meets the expected specifications. Audio requirements information can be timbre, specifying the type of voice, such as "male voice - steady," "female voice - lively," "child's voice," or "professional announcer," etc. It can also be speech rate: setting the speed of speech playback, such as "slow," "normal," "fast," or a specific word / minute value. Furthermore, it can be temperature (or emotion / tone): describing the emotional coloring of the voice, such as "warm," "calm," "excited," "friendly," etc., to match the emotional dimension of the content. This information typically exists in the form of parameterized labels or instructions.
[0109] It's important to note that audio synthesis models are artificial intelligence algorithms whose core function is to automatically convert input text information into natural, fluent human speech. Audio synthesis models can precisely control various characteristics of sound, quickly generating audio versions in multiple styles to meet diverse audio generation needs. Therefore, they can generate corresponding audio for massive amounts of multimedia text in a short time, satisfying high-intensity generation requirements.
[0110] In practical applications, the method of generating target audio based on audio requirements and multimedia-generated text can involve calling standardized functions, text cleaning tools, or pre-trained cleaning models to convert numerical content such as numbers, amounts, and phone numbers in the multimedia-generated text into Chinese characters. This allows the speech synthesis model to recognize and synthesize the text, avoiding the omission of numbers. The cleaned text is then segmented into multiple speech chunks to accommodate the maximum synthesized audio size limit of the speech synthesis model, such as 30 seconds. Timbre vectors are selected manually or automatically, the speech rate is set, and the temperature coefficient of the sound is determined. The speech synthesis model then synthesizes the audio from these multiple speech chunks, generating corresponding audio segments. These audio segments are then merged to obtain the complete audio corresponding to the multimedia-generated text, i.e., the target audio. See [link to details] for further information. Figure 7 , Figure 7 This is a flowchart illustrating a target audio generation process according to one embodiment of this specification. The generated audio file is as follows: Figure 8 , Figure 8This is a diagram illustrating the interface of a target audio file provided in one embodiment of this specification. The specific content is as follows: Title: The text above reads "Example of a Voice Audio File:", and the file name is "Voice Audio File.wav". The file size is "3 MB". Playback controls: There is a circular play button on the left for previewing the voice audio content. The playback progress bar displays the current playback time as "00:00" and the total duration as "01:05". A black dot on the progress bar indicates the current playback position, currently at the starting point. There are speed control buttons on the right to switch the playback speed.
[0111] The solution described in this specification, by incorporating audio requirement information and flexibly setting parameters according to the characteristics of the target project, precisely controls various acoustic features of the generated speech. This process not only significantly reduces production costs and time but also ensures the uniqueness of the speech output in different scenarios.
[0112] In one optional embodiment of this specification, the target multimedia data includes a target video;
[0113] Generating text based on multimedia, and generating target multimedia data corresponding to the target project, may include the following steps:
[0114] Generate text based on multimedia and generate target audio corresponding to the target project;
[0115] Generate a target video corresponding to the target project based on the reference digital human video and the target audio.
[0116] It's important to note that target videos are multimedia data presented in video format. They are used to showcase project marketing information based on trending topics, thereby expanding the target project's influence. Target videos can be short videos featuring a digital human narrating, where the video screen shows a virtual digital human whose lip movements, facial expressions, and body language are perfectly synchronized with generated audio, narrating promotional content about the target project. They can also be text-and-image + voice-over videos: the video consists of a series of dynamic text and images, infographics, or product images displayed in a slideshow, accompanied by generated "target audio" narration. Alternatively, they can be composite videos: the system automatically retrieves and edits relevant live-action footage, animation effects, or copyrighted video clips from the resource library according to a script, and uses generated audio as narration to create a complete promotional short film. The specific choice can be made based on the actual situation.
[0117] Reference digital human videos are pre-created and stored standard image video clips or templates recorded by a reference person. Reference digital human videos define the appearance, movements, and basic behavioral patterns of the digital human. Therefore, based on reference digital human videos, it is possible to obtain the appearance, behavioral patterns, and morphological changes when speaking of the digital human corresponding to the target video.
[0118] In one possible scenario, the reference digital human video can be pre-recorded, such as a processed pre-recorded video clip of one or more reference figures speaking. Since a one-minute video can capture sufficient feature information without excessive invalid movements, a one-minute video clip of the reference figure speaking is typically recorded. Various video sizes are available, such as aspect ratios of 9:16, 4:3, and 10:16. The 9:16 aspect ratio is widely used due to its versatility and good image quality. Video resolutions can be 360P, 720P, 1080P, 2K, or 4K, with a frame rate of 60 frames per second or higher being preferred. The format can be MP4. The reference digital human should exhibit rich mouth movements and some body motion while speaking, and the submitted video should demonstrate a lifelike performance that meets the client's expectations and the requirements of the corresponding scenario. For specific examples of digital human figures, please refer to [link to relevant documentation]. Figure 9 , Figure 9 This is a diagram illustrating a reference digital human's display interface, provided in one embodiment of this specification.
[0119] Once the reference digital human video is obtained, the target audio of the reference digital human video can be processed to generate the target video corresponding to the target project. For details, please refer to [link to relevant documentation]. Figure 10 , Figure 10 This is a flowchart of a target video generation process provided in one embodiment of this specification.
[0120] The entire process consists of four main steps:
[0121] Input Processing: The system receives a reference digital human video input and performs preliminary processing, including: Video Frame Extraction: Decomposing the continuous video into individual still images. Audio Feature Extraction: Simultaneously, the system analyzes the input target audio, extracting its key acoustic features (such as spectrum, pitch, and energy) to prepare for subsequent animation.
[0122] Face Detection & Processing: Facial Keypoint Extraction: Precisely locates the face in each frame of the image and extracts the key points that constitute the facial contour and expression (such as the coordinates of the eyes, nose, and mouth). Image Cropping and Encoding: Crops the face region from the background and converts it into an encoding format that is easy for computer processing, facilitating subsequent deformation and animation generation.
[0123] Lip-Sync Generation: Cross-modal fusion and inference: The system fuses the audio features extracted in the previous step with facial keypoint information. Using a deep learning model (such as a neural network), it infers the specific lip shape that the face should present under the current speech pronunciation (such as "ah," "oh," "m," etc.). Latent space decoding and post-processing: The model generates a latent space representation, which is then converted into specific facial deformation parameters by a decoder. Post-processing such as smoothing and denoising is then performed to ensure the natural and smooth animation.
[0124] Video Synthesis: Image Pasting and Merging: The facial image with the new lip movements generated in step three is pasted back into the background of the original video. Audio and Video Stream Encapsulation: Finally, the processed video stream is synchronized with the original target audio and packaged into a complete, playable digital human video file, i.e., the target video.
[0125] The solutions described in this specification enable the rapid, batch production of high-quality promotional videos without the need for professional actors, recording studios, or complex post-production editing. This significantly reduces production costs and time, increases the efficiency and scale of content creation, and allows operators to respond to trending topics at lightning speed and continuously release new content, thereby effectively enhancing brand exposure and user engagement.
[0126] In one possible embodiment of this specification, generating a target video corresponding to a target item based on a reference digital human video and target audio data may include the following steps:
[0127] Perform speech recognition on the target audio to obtain the audio text;
[0128] Generate a target video corresponding to the target project based on the audio text, the target audio, and the reference digital human video.
[0129] It should be noted that audio text refers to the text annotation content corresponding to the input target audio. The audio text needs to be accurate to the timestamp so that specific textual auxiliary information can be presented in the corresponding position in the generated target video to help viewers understand the video content.
[0130] Specifically, the audio text recognition process can be found in [link to documentation]. Figure 11 , Figure 11 This is a flowchart illustrating an audio text generation process according to one embodiment of this specification. Specifically, the synthesized audio file calls a speech recognition model to obtain a .srt format subtitle file. The audio text (subtitle file) contains the text to be spoken for each corresponding time period. The text in the subtitle file is segmented to prevent large blocks of text from being displayed below the video; short texts are displayed on one line, while long texts are displayed on two lines.
[0131] The generated audio text (subtitle file) can be referenced. Figure 12 , Figure 12 This is a schematic diagram of an audio text representation provided in one embodiment of this specification. The numbers preceding the subtitles (e.g., 1, 2, 3...) indicate the order of the subtitles. Above the subtitles is a timestamp, in the format 00:00:00,000 --> 00:00:05,000, defining the start and end times of the subtitle's display in the video, in the format "hour:minute:second,millisecond". Below the timestamp is the subtitle text, i.e., the subtitle content to be displayed within the specified time period. Longer subtitle texts, such as "The robot dog has reached a high level," are split into two lines for presentation.
[0132] After generating the audio text, you can use ffmpeg technology to select the WenQuanYi font and load the audio text to obtain a digital human video with subtitles.
[0133] There are several ways to generate a target video for a target project based on audio text, target audio, and reference digital human video. One possible approach is to generate a target video for the target project based on the target audio and reference digital human video, and then append the audio text in segments to the corresponding positions in the target video to obtain the final target video.
[0134] The target video can be found here. Figure 13 , Figure 13 This is a diagram showing the display interface of a target video provided in one embodiment of this specification.
[0135] By applying the solutions in the embodiments of this specification, audio text is obtained by introducing speech recognition of the target audio, and a dual-path parallel display mechanism of audio and text is constructed. This makes the generated video more accurate, smooth and realistic in terms of audio-visual matching, naturalness of expression and overall expressiveness, effectively ensuring the professional quality of automated production content and the audience experience.
[0136] In one possible embodiment of this application, before generating target multimedia data corresponding to the target item based on multimedia-generated text, the following steps may be included:
[0137] The multimedia-generated text is evaluated using quality assessment standards and / or quality assessment models to obtain assessment indicators, which are used to reflect the quality of the multimedia-generated text.
[0138] Based on the evaluation metrics, the target text to be generated is determined;
[0139] Based on multimedia text generation, generate target multimedia data corresponding to the target project, including:
[0140] Based on the target text, generate target multimedia data corresponding to the target project.
[0141] It should be noted that "quality assessment standards" refer to a set of pre-defined rules, indicators, or models used to measure and judge "multimedia-generated text" or "target multimedia data," determining whether the quality of the multimedia-generated text is up to standard and whether it can be used as the text for generating the target video. Quality assessment standards can be automated standards set by code or programs, or they can be manual review standards, depending on the specific circumstances.
[0142] A quality assessment model is an artificial intelligence algorithm used to automatically judge the quality of generated content, specifically for reviewing multimedia-generated text. The model takes the text to be evaluated, relevant multimedia reference information, and the multimedia generation topic as input. Through pre-trained algorithms, it analyzes the text's fluency, logic, relevance to the topic, and whether it meets the characteristics of viral content, ultimately outputting a quality score or a "pass / fail" result.
[0143] There are several ways to evaluate multimedia-generated text using quality assessment standards and / or models. One approach is to use the target project and its marketing information as the standard, evaluating the multimedia-generated text using these standards and models. This method focuses on examining the content of the multimedia-generated text to ensure a close connection between the text and the target project and its marketing information, avoiding model-induced illusions. Another approach is to use quality assessment standards and / or models to evaluate the completeness of the multimedia-generated text, ensuring that the target generated text contains all the content.
[0144] Evaluation metrics are specific standards used to quantify and measure the quality of generated content. They provide a basis for judging whether multimedia-generated text meets the standards, ensuring that the generated multimedia text or target multimedia data reaches the expected level. The criteria for evaluating metrics may include topic relevance, information completeness, language fluency, structural rationality, emotional fit, compliance, and so on.
[0145] In one possible scenario, the target generated text is determined based on evaluation metrics. This could be done by setting scores for the evaluation metrics based on judgment criteria. If the evaluation metrics are higher than the preset scores, the multimedia generated text is selected as the target text. If the evaluation metrics are lower than the preset scores, the multimedia generated text is optimized until the evaluation metrics of the optimized text are higher than the preset scores. The optimized text is then used as the target generated text.
[0146] After the above evaluation, the quality of the target generated text can be guaranteed. At this point, based on the target generated text, the target multimedia data corresponding to the target project is generated, which effectively ensures the content integrity of the target multimedia data.
[0147] For example, the quality assessment process is as follows: "The initial draft is as follows: 'Hello everyone, did you know? Children's health problems always cause parents a lot of headaches. Today I want to introduce you to the best product that can reduce parents' worries and give them more peace of mind. This insurance product not only covers outpatient and emergency care, accidents, medical expenses, and critical illnesses, but also only costs a few tens of yuan per month, making it extremely cost-effective. Outpatient and emergency care is 100% reimbursed, hospitalization is 90% reimbursed, and medication has no deductible, with a 50% reimbursement rate. There is no limit to the number of claims, and the outpatient and emergency medical insurance fund is shared at 50,000 yuan per year. From minor ailments like colds, fevers, vomiting, diarrhea, influenza A and B, and cat or dog bites requiring hospitalization; to major illnesses like pneumonia or even cancer requiring hospitalization, you can apply for reimbursement (within the scope of coverage), and there is no deductible for accidents and medications. Even better, multiple children can enjoy a 5% discount, and future renewals can also enjoy a 5% discount! After successful enrollment, you can also enjoy free dental fluoride treatments, specialist outpatient visits, and green channel services for medical treatment.'" This insurance is
medical insurance
[0148] The final draft has been revised as follows: "Hello everyone, did you know that children's health issues are always a major headache for parents? Today, I'm going to introduce a product that can ease parents' worries and give them peace of mind. This insurance product not only covers outpatient and emergency care, accidents, medical expenses, and critical illnesses, but it also only costs a few dozen yuan per month, making it extremely cost-effective. It offers 100% reimbursement for outpatient and emergency care, 90% reimbursement for hospitalization, and 0 deductible for medications with a 50% reimbursement rate. There are no limits on the number of claims, and the outpatient and emergency medical insurance fund is shared at 50,000 yuan per year (no restrictions on the catalog, with a daily limit of 1,000 yuan). It covers everything from minor ailments like colds, fevers, vomiting, diarrhea, influenza A and B, and cat or dog bites to serious conditions like pneumonia and even cancer requiring hospitalization (within the coverage area). Accidents and medications have 0 deductibles. Even better, multiple children enrolled receive a 5% discount, and future renewals also enjoy a 5% discount! After successful enrollment, you can also enjoy free dental fluoride treatments, specialist outpatient visits, and expedited medical services. This insurance is [Medical Insurance]. Click to join our live stream to inquire!"
[0150] The solution implemented in this specification utilizes preset quality assessment standards and / or more intelligent quality assessment models to quantitatively evaluate the initially generated multimedia text, outputting assessment indicators reflecting its quality level. Based on these indicators, the optimal version is automatically selected or optimized and determined as the target generated text. Target multimedia data is then generated based on the target generated text. This effectively avoids the generation of low-quality, off-topic, or illegal content, significantly improving the professionalism, compliance, and dissemination effect of the final output while achieving automated mass production of content, thus ensuring both efficiency and quality.
[0151] One possible embodiment of this specification, which determines the target content to be published based on evaluation metrics, may include the following steps:
[0152] When multimedia-generated text fails to be generated based on evaluation metrics, suggestions for text modification are obtained.
[0153] Based on the text modification suggestions, generate the target generated text.
[0154] It should be noted that text modification suggestions refer to specific opinions and directions for guiding how to improve the text. These suggestions can be automatically generated by the system based on content that does not meet the evaluation criteria, or they can be determined manually based on the evaluation criteria.
[0155] In practical applications, there are several ways to generate target text based on text modification suggestions. One possible approach is to generate prompts based on the text modification suggestions, then call a correction model to generate text based on the prompts and multimedia, thus generating the target text. Another possible approach is to generate a visual operation page based on the text modification suggestions and the multimedia-generated text, allowing users to modify the multimedia-generated text according to the suggestions, thereby generating the target text.
[0156] The solution implemented in this specification not only identifies quality issues in multimedia-generated text during the evaluation phase but also accurately generates specific text modification suggestions, indicating directions for improvement. Based on these suggestions, target generated text is generated, allowing for targeted optimization of the text based on guided feedback. This significantly improves the efficiency and success rate of producing high-quality target generated text.
[0157] In one possible embodiment of this specification, before generating target multimedia data corresponding to the target project based on multimedia generation themes and multimedia reference information, the following steps may also be included:
[0158] Obtain the multimedia generation framework, which is used to constrain the multimedia data generation structure of the target project;
[0159] Based on the multimedia generation theme and multimedia reference information, target multimedia data corresponding to the target project is generated, including:
[0160] Based on the multimedia generation theme, multimedia reference information, and multimedia generation framework, target multimedia data corresponding to the target project is generated.
[0161] It's important to note that a multimedia generation framework refers to a template or set of rules used to constrain and guide the generation structure of multimedia data for a target project; it can also be called a "creation outline." It defines the organizational form and presentation logic of the content, ensuring that the generated video or audio structurally meets specific requirements. It is the core of achieving content standardization and modularization. The multimedia generation framework specifies narrative structures, modular content segments, style templates, etc., providing a clear generation structure for the target multimedia data, enabling faster content creation and reducing decision-making costs.
[0162] In one possible implementation, the multimedia generation framework is pre-edited and saved in the system by the system administrator. The multimedia generation framework has blank content so that when it is sent to the user, the current user can simultaneously input user-defined content while selecting the multimedia generation framework, thus realizing the integration of user needs and automated generation.
[0163] When generating the final target multimedia data, the core direction and marketing focus of the target multimedia data content are provided by integrating and coordinating three key elements: multimedia generation theme, multimedia reference information, and multimedia generation framework, ensuring that the data stays aligned with project objectives. Success models and attraction principles from viral content are provided to enhance the dissemination effect of the target multimedia data. A standardized content structure and organizational logic are provided to guarantee that the output format conforms to standards.
[0164] Therefore, by applying the solution of the embodiments in this specification, a comprehensive instruction integrating "strategic objectives," "creative inspiration," and "execution blueprint" is set up, realizing the unity of intelligence, standardization, and personalization, thereby generating high-quality multimedia content that not only meets the needs of brand promotion but also has market appeal and a complete and standardized structure.
[0165] By applying the solution outlined in this manual, the system automates the acquisition of project topics, marketing information, and multimedia data for target projects, achieving a precise integration of external trends and internal needs. Combining project topics and marketing information to determine multimedia generation themes provides a clear focus for content creation. In-depth analysis of the multimedia data corresponding to the project topics extracts multimedia reference information covering multiple dimensions such as structure, language, and emotion, enabling the system to learn and reuse successful models validated by the market. Finally, based on the combination of binding multimedia generation themes and guiding multimedia reference information, target multimedia data is collaboratively generated. This significantly improves the speed of responding to trending topics and the efficiency of content production, solving the problems of long cycles and high costs associated with manual creation. More importantly, through its intelligent analysis and generation mechanism, it ensures the quality of creative content and audience appeal while achieving large-scale, standardized content output, significantly enhancing the operator's market competitiveness and brand influence in the rapidly changing internet environment.
[0166] The following is in conjunction with the appendix Figure 14 Taking the application of the multimedia data method provided in this specification in the field of insurance marketing as an example, this paper further explains the multimedia data method. Figure 14 This is a flowchart illustrating the processing steps of a multimedia data generation method according to one embodiment of this specification, specifically including the following steps:
[0167] Step 1402: Use automated data scraping tools to obtain project topics, project marketing information, and high-viewership videos corresponding to the project topics for insurance marketing projects.
[0168] Automated crawling tools can be robotic process automation (RPA) tools, which are technologies that use software robots to simulate human operations on computers to automatically perform repetitive, rule-based tasks.
[0169] In this step, the automated scraping tool can automatically obtain high-viewership and high-popularity topics from various content publishing platforms on the Internet and input them into the classification model for classification.
[0170] The classification model can categorize topics acquired by automated web scraping tools based on user-provided prompts. Examples of prompts are as follows: "You are currently a creative copywriting expert and semantic understanding expert. You excel at categorizing trending topics and providing creative ideas and perspectives from multiple angles within each category. My categories include: Insurance Marketing, Health Science Popularization, Financial Commentary, and Other. I'll give you a trending topic; please categorize it according to the following rules and provide creative ideas:\n" +"Step 1: You need to deeply understand and analyze the trending topic and the deeper meaning of these categories, and be able to elaborate on the topic in detail based on the categories I provide;\n" +"Step 2: Content related to medical resources, medical expenses, early screening, treatment costs, holidays, travel safety, transportation, typhoons, property loss, financial security, and travel can all be categorized under Insurance Marketing,\n" +"For example, travel insurance, medical insurance, critical illness insurance, accident insurance, and property insurance;\n" +"Step 3: Content related to social finance is categorized under Financial Commentary;\n" +"Step 4: Content related to healthy eating, exercise, sleep, mental health, health screening, health management, health records, health monitoring, health knowledge, health courses, health lectures, health promotion, community services, and health education can all be categorized under the Health Science Popularization category;\n" +"Step 5: After categorizing, you need to provide a creative explanation that fits the actual situation and the core viewpoints that can be used for creation. The creative explanation must not exceed 200 words;\n" +"Step 6: You need to think deeply about whether a topic can be creatively combined with multiple categories. If so, you need to return multiple categories. The creative description is a summary of the creative directions of multiple categories;\n" +"Step 7: The output format is: [{"hotTopic":"","analyzer":[{"classify":"","description":"}]}], and the result only returns array format data. There should be no other text or other additional information.\n"
[0172] Based on the above, the system can obtain project topics, project marketing information, and high-viewership videos related to insurance marketing projects.
[0173] Step 1404: Based on the project topic and project marketing information, determine the creative inspiration, which is used to constrain the direction of the document data generation for the insurance marketing project.
[0174] After determining the project topic and marketing information, the creative inspiration generation model can identify creative ideas based on these elements. An example prompt would be: "You are a creative topic writing expert. I now need to create a creative copy about children's insurance based on a given topic. The selling points of children's insurance products are as follows:"
[0175] 1. 100% reimbursement rate for outpatient and emergency visits (within coverage), 90% reimbursement rate for inpatient care, 50% reimbursement rate for medication (no deductible), requirements: - Only output the creative theme; no other irrelevant text is needed. - The given creative theme should be eye-catching. For example: Input product selling points: 1. 100% reimbursement rate for outpatient and emergency visits (within coverage), 90% reimbursement rate for inpatient care, 50% reimbursement rate for medication (no deductible). Output creative themes: ["Isn't there a baby insurance policy that reimburses 100% for outpatient and emergency visits?", "Are there insurance policies that cover a baby's cold and fever?", "If your child has an accident, can you afford it now?", "The cost of a cup of milk tea per month can cover medical expenses without hesitation", "No matter what medication you choose, it will be reimbursed"]
[0176] Step 1406: Analyze high-viewership videos, determine the video script, and obtain imitation suggestions based on the multimedia script.
[0177] Meanwhile, the script generation model can analyze high-viewership videos based on prompt words to determine the video script and provide imitation suggestions based on the multimedia script. Examples of prompt words are as follows:
[0178] Role: Copywriting Deconstruction Master. As a seasoned copywriting deconstruction master, your expertise lies in deeply analyzing various successful video scripts, uncovering the creative techniques and strategies behind them. Your goal is to enable users to quickly grasp the essence of the script through meticulous analysis, and based on this, to quickly master the key points and methods of imitation writing. Skills: Copywriting Analysis: - Proficient in various types of spoken video scripts, including but not limited to advertisements, tutorials, reviews, etc., and able to accurately grasp the uniqueness of each type of script. - Able to conduct comprehensive and in-depth analysis of scripts from multiple dimensions such as structure, language style, and emotional appeal, revealing the key factors for their success. Strategy Extraction: - Accurately extract efficient creative strategies and practical techniques from successful spoken video scripts. - Deeply understand and flexibly apply psychological principles to enhance the appeal of the script to the target audience, improving its persuasiveness and appeal. Imitation Writing Guidance: - Skilled at summarizing and extracting key information, providing users with detailed and highly actionable imitation writing guidance strategies to help users create efficiently.
[0179] Strategy: Copywriting Analysis: A thorough analysis of the provided script will be conducted, breaking it down meticulously from the following aspects, with 3-5 sub-points analyzed in each direction: ... - Structure and Layout Analysis: How does the opening sentence (ending with a period) cleverly attract the audience's attention? How does the development section progressively build upon and expand the argument? How does the ending subtly summarize and echo the preceding text? And how does the overall structure efficiently convey the core message? - Core Message: Deeply explore the main theme the script aims to convey, clarifying its core selling point, concept, or emotional appeal. Analyze how this message permeates the entire script, guiding the audience to understand and resonate. - Language Style: Examine the script's language characteristics, such as whether it is conversational, vivid, or professional and rigorous; whether the word choice is precise and impactful; and how this style aligns with the target audience's preferences and listening habits. - Emotional Appeal: Analyze how the script touches the audience's emotions—whether through heartwarming stories, humorous elements, or an impassioned tone—and how this emotional appeal enhances the script's persuasiveness and communicative power. - Challenging Perspectives: Extract the challenging perspectives presented in the copy, exploring how these perspectives stimulate audience thinking, break conventional perceptions, and inspire audience action and engagement. - Core Appeal of the Copy: Summarize the most attractive parts of the copy, such as unique creativity, novel perspectives, or highly valuable information, analyzing how these elements make the copy stand out and attract audience attention. - Emotional Keywords Analysis: For user-submitted audio copy, carefully select 5-10 emotional keywords that effectively evoke emotional resonance in the audience, adding impact to the copy. - Quotes and Key Phrases Analysis: For user-submitted copy, identify which quotes are key phrases, list at least 5, and provide a concise and powerful commentary for each, highlighting their important role and unique charm in the copy. - Imitation Suggestions: Based on the above analysis results, provide specific imitation suggestions covering aspects such as structure, language style, emotional appeal, challenging perspectives, core appeal of the copy, quotes, and key phrases, helping to create equally compelling audio copy.
[0180] The text is as follows: ...
[0181] Step 1408: Based on creative inspiration, imitation suggestions, and creation outline, generate video scripts corresponding to insurance marketing projects.
[0182] The system can further extract preset or user-generated creative outlines, and based on prompts and a copywriting generation model, combined with creative inspiration, imitation suggestions, and the creative outline, generate video scripts corresponding to insurance marketing projects. Example of prompts:
[0183] Role: Copywriting Master: As a seasoned copywriting expert, you specialize in imitating and reshaping various types of content, accurately grasping and creatively mimicking the style, tone, and intent of given texts to produce high-quality and original copy. You possess keen market insight and solid creative writing skills, enabling you to flexibly adjust copywriting strategies according to different scenarios to adapt to product promotion needs. With years of experience in the insurance industry, you are particularly adept at uncovering the unique selling points of insurance products, cleverly integrating product characteristics into diverse life scenarios to create scripts that are both attractive and aligned with product features, resonating strongly with the target audience. Scenario: Currently, I am working on creating a video broadcast script for an insurance product's video platform, hoping the video will attract users to the live stream, ultimately boosting product sales.
[0184] Skills: Content Analysis: - Ability to deconstruct any article, understand its core message, target audience, and emotional appeal. - Skills include identifying key elements that make the original content effective and engaging. Creative Rewriting: - Expertise in creatively transforming the essence of the original text into novel and refreshing content while maintaining the original intent. - Ability to innovate and apply cross-industry insights to enhance the appeal of content to a wider audience. Language Expression: - Ability to imitate and apply the language style and tone of the original copy, while making appropriate adjustments as needed. - Excellent language organization and editing skills, able to accurately express ideas and ensure the copy is fluent and engaging. Strategy: Based on the imitation writing suggestions and the selling points of the insurance product, carefully create the spoken script, strictly adhering to the product selling points. The text within the selling points does not need to be modified, and no words should be omitted or missing. Subtly integrate the selling points to ensure the content highly aligns with the product characteristics. Product Information: Insurance Product: [(To be filled in). Product Features: 1. (To be filled in).] Script creation requirements: 1. The spoken script needs to be conversational, and the corresponding video length should be within 40 seconds. However, the script must have sufficient word count, with the spoken script containing approximately 250-300 words. 2. The first 0-3 seconds should use a suggested opening, presenting a challenging viewpoint. It must be an explosive opening that showcases the product's uniqueness and distinguishes it from others, laying the groundwork for future developments. This question should be relevant to the product, quick, and eye-catching enough to arouse the user's curiosity and attention. Common techniques include: building climax beforehand, asking questions while reading aloud, and using catchy phrases at the beginning. (To be filled in).
[0185] Script creation restrictions: 1. Avoid using superlative descriptions such as "most" or "comprehensive." 2. Avoid using prompts such as links to or QR codes for live streams. 3. Consider using product features to create new scenarios for promoting the product. 4. The script structure should follow this pattern: emotional engagement + product information + product features + applicable scenarios + purchase guidance. 5. Avoid the following words in the script: 1) Absolute terms 2) Ranking-related words 3) Authoritative words 4) Promises and guarantees 5) Financial investment terms 6) Superstitious and illegal information 7) Sensitive topics and behaviors; content that causes discomfort or easily induces imitation; unsafe or illegal information; content that misleads minors or violates public order and good morals; suggestions for deconstruction and imitation are as follows: ...
[0186] For specific implementation details of 1402-1408 above, please refer to [link / reference]. Figure 15 , Figure 15 This is a flowchart illustrating a video script creation process provided in one embodiment of this specification.
[0187] Step 1410: Based on the multimedia generation theme, multimedia generation framework, and multimedia reference information, generate the target multimedia data corresponding to the target project.
[0188] The system integrates multimedia generation themes, frameworks, and reference information to drive a text generation model to create a complete video script. This ensures that the generated content has clear marketing objectives, appealing user-friendly features, and a clear and standardized structure.
[0189] Step 1412: Based on the video transcript, generate the target audio corresponding to the insurance marketing project.
[0190] The text script is converted into speech. The system inputs the video script generated in the previous step into the audio synthesis model. The model combines preset audio requirements (such as selecting a calm and credible male voice, a moderate speaking speed, and a professional and composed tone) to automatically generate a natural and fluent speech file that matches the tone of the insurance product, i.e., the target audio.
[0191] Step 1414: Generate the video to be labeled corresponding to the target project based on the audio text and the reference digital human video.
[0192] The system first performs speech recognition on the target audio to obtain audio text accurate to a specific time point. Then, using the semantic and temporal information contained in this audio text, it drives a pre-set reference digital human video (such as a virtual insurance advisor) to perfectly match its lip movements, facial expressions, and actions with the audio, thereby generating a preliminary video to be labeled with a broadcast.
[0193] Step 1416: Recognize the audio text to obtain the subtitle file, use the letter file to annotate the target video, and generate the target video.
[0194] The system generates standard subtitle files from the audio text using formats such as SRT. This subtitle file is then overlaid onto the "video to be annotated" generated in the previous step, displaying synchronized scrolling text at the bottom of the screen or in a suitable location. After this process, a complete finished product containing audio, video, and subtitles—the final "target video" ready for publication—is obtained.
[0195] For specific implementation details of 1410-1416 above, please refer to [link / reference]. Figure 16 , Figure 16 This is a flowchart illustrating a target video creation process provided in one embodiment of this specification.
[0196] Corresponding to the above method embodiments, this specification also provides embodiments of a multimedia data generation apparatus. Figure 17 This is a schematic diagram of the structure of a multimedia data generation device provided in one embodiment of this specification. Figure 17 As shown, the device includes:
[0197] The first acquisition module 1702 is configured to acquire the project topic, project marketing information, and multimedia data corresponding to the project topic of the target project.
[0198] The determination module 1704 is configured to determine a multimedia generation theme based on the project topic and the project marketing information, wherein the multimedia generation theme is used to constrain the direction of multimedia data generation for the target project.
[0199] The parsing module 1706 is configured to parse the multimedia data to obtain multimedia reference information.
[0200] The generation module 1708 is configured to generate target multimedia data corresponding to the target project based on the multimedia generation theme and the multimedia reference information.
[0201] Optionally, the first acquisition module 1702 is further configured to acquire the project topic and the project marketing information of the target project; and retrieve the multimedia data corresponding to the project topic from the multimedia publishing platform.
[0202] Optionally, the first acquisition module 1702 is further configured to acquire the project marketing information of the target project; read multiple candidate topics from the multimedia publishing platform and identify the topic categories of the multiple candidate topics; and based on the topic categories, filter out the project topics that match the category of the target project from the multiple candidate topics.
[0203] Optionally, the parsing module 1706 is further configured to parse the multimedia data to obtain a multimedia script; generate a model for generating prompt information and the multimedia script input information to obtain multimedia reference information, wherein the generated prompt information is used to guide the information generation model to analyze the multimedia script from a target dimension and generate the multimedia reference information, wherein the target dimension includes at least one of structural dimension, language dimension, emotion dimension, viewpoint dimension, value dimension, and rhetoric dimension.
[0204] Optionally, the generation module 1708 is further configured to input the multimedia generation topic and the multimedia reference information into a text generation model to obtain multimedia generation text corresponding to the target project; and to generate target multimedia data corresponding to the target project based on the multimedia generation text.
[0205] Optionally, the generation module 1708 is further configured to generate target multimedia data corresponding to the target project based on the multimedia generated text, including: obtaining audio requirement information, wherein the audio requirement information includes at least one of timbre, speech rate, and temperature; inputting the audio requirement information and the multimedia generated text into an audio synthesis model to obtain the target audio corresponding to the target project.
[0206] Optionally, the generation module 1708 is further configured to generate target audio corresponding to the target item based on the multimedia generated text; and to generate target video corresponding to the target item based on the reference digital human video and the target audio.
[0207] Optionally, the generation module 1708 is further configured to perform speech recognition on the target audio to obtain audio text; and generate the target video corresponding to the target project based on the audio text, the target audio, and the reference digital human video.
[0208] Optionally, the multimedia data generation device further includes a quality assessment module, configured to evaluate the multimedia-generated text using quality assessment standards and / or quality assessment models to obtain assessment indicators, wherein the assessment indicators are used to reflect the quality of the multimedia-generated text; and to determine the target generated text based on the assessment indicators.
[0209] Optionally, the generation module 1708 is further configured to generate target multimedia data corresponding to the target item based on the target generated text.
[0210] Optionally, the multimedia data generation device further includes a second acquisition module configured to acquire a multimedia generation framework, wherein the multimedia generation framework is used to constrain the multimedia data generation structure of the target project.
[0211] Optionally, the generation module 1708 is further configured to generate target multimedia data corresponding to the target project based on the multimedia generation theme, the multimedia reference information, and the multimedia generation framework.
[0212] The video generation method provided in this specification involves acquiring the project topic, marketing information, and corresponding multimedia data of a target project; determining a multimedia generation theme based on the project topic and marketing information, whereby the multimedia generation theme constrains the direction of multimedia data generation for the target project; parsing the multimedia data to obtain multimedia reference information; and generating target multimedia data corresponding to the target project based on the multimedia generation theme and multimedia reference information. By automatically acquiring project topics, marketing information, and multimedia data, the system achieves real-time capture of market dynamics. Finally, by combining the constraints of the theme with the guidance of the reference information, target multimedia data is automatically generated. This compresses the manual creation cycle, which originally required hours or even days, to minutes, greatly improving the timeliness and scalability of content production. It effectively enhances user engagement and brand exposure, providing strong technical support for achieving sustainable traffic growth.
[0213] The above is an illustrative scheme of a multimedia data generation apparatus according to this embodiment. It should be noted that the technical solution of this multimedia data generation apparatus and the technical solution of the multimedia data generation method described above belong to the same concept. For details not described in detail in the technical solution of the multimedia data generation apparatus, please refer to the description of the technical solution of the multimedia data generation method described above.
[0214] Figure 18 This is a structural block diagram of a computing device according to one embodiment of this specification. The components of the computing device 1800 include, but are not limited to, a memory 1810 and a processor 1820. The processor 1820 is connected to the memory 1810 via a bus 1830, and a database 1850 is used to store data.
[0215] The computing device 1800 also includes an access device 1840, which enables the computing device 1800 to communicate via one or more networks 1860. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 1840 may include one or more of any type of wired or wireless network interface (e.g., Network Interface Card (NIC)), such as an IEEE 802.11 Wireless Local Area Networks (WLAN) interface, a Wi-MAX (World Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0216] In one embodiment of this specification, the above-described components of the computing device 1800 and Figure 18 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 18 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0217] The computing device 1800 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1800 can also be a mobile or stationary server.
[0218] The processor 1820 is used to execute computer programs / instructions, which, when executed by the processor, implement the steps of the multimedia data generation method described above.
[0219] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the multimedia data generation method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the multimedia data generation method described above.
[0220] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the multimedia data generation method described above.
[0221] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the multimedia data generation method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the multimedia data generation method described above.
[0222] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the multimedia data generation method described above.
[0223] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the multimedia data generation method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the multimedia data generation method described above.
[0224] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0225] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0226] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0227] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0228] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A method for generating multimedia data, characterized in that, include: Acquire the target project's project topics, project marketing information, and multimedia data corresponding to the project topics; Based on the project topic and the project marketing information, a multimedia generation theme is determined, wherein the multimedia generation theme is used to constrain the direction of multimedia data generation for the target project; The multimedia data is parsed to obtain multimedia reference information; Based on the multimedia generation theme and the multimedia reference information, target multimedia data corresponding to the target project is generated.
2. The method according to claim 1, characterized in that, The acquisition of the target project's topic, marketing information, and corresponding multimedia data includes: Obtain the project topic and marketing information of the target project; Multimedia data corresponding to the project topic was retrieved from the multimedia publishing platform.
3. The method according to claim 2, characterized in that, The acquisition of the project topic and the project marketing information of the target project includes: Obtain the project marketing information of the target project; Multiple candidate topics are read from the multimedia publishing platform, and the topic categories of the multiple candidate topics are identified; Based on the topic category, the project topic that matches the target project category is selected from the plurality of candidate topics.
4. The method according to claim 1, characterized in that, The process of parsing the multimedia data to obtain multimedia reference information includes: The multimedia data is parsed to obtain the multimedia script; The multimedia reference information is obtained by generating a model that generates prompt information and multimedia script input information. The prompt information is used to guide the information generation model to analyze the multimedia script from a target dimension and generate the multimedia reference information. The target dimension includes at least one of the following: structural dimension, language dimension, emotion dimension, viewpoint dimension, value dimension, and rhetoric dimension.
5. The method according to claim 1, characterized in that, The step of generating target multimedia data corresponding to the target project based on the multimedia generation theme and the multimedia reference information includes: Input the multimedia generation topic and the multimedia reference information into the text generation model to obtain the multimedia generated text corresponding to the target project; Based on the multimedia generated text, target multimedia data corresponding to the target project is generated.
6. The method according to claim 5, characterized in that, The target multimedia data includes target audio; The step of generating target multimedia data corresponding to the target project based on the multimedia-generated text includes: Obtain audio requirement information, wherein the audio requirement information includes at least one of timbre, speech rate, and temperature; The audio requirement information and the multimedia generated text are input into the audio synthesis model to obtain the target audio corresponding to the target project.
7. The method according to claim 5, characterized in that, The target multimedia data includes the target video; The step of generating target multimedia data corresponding to the target project based on the multimedia-generated text includes: Based on the multimedia generated text, the target audio corresponding to the target item is generated; The target video corresponding to the target project is generated based on the reference digital human video and the target audio.
8. The method according to claim 7, characterized in that, The step of generating the target video corresponding to the target project based on the reference digital human video and the target audio includes: Perform speech recognition on the target audio to obtain audio text; The target video corresponding to the target project is generated based on the audio text, the target audio, and the reference digital human video.
9. The method according to claim 5, characterized in that, Before generating the target multimedia data corresponding to the target project based on the multimedia-generated text, the method further includes: The multimedia-generated text is evaluated using quality assessment standards and / or quality assessment models to obtain assessment indicators, wherein the assessment indicators are used to reflect the quality of the multimedia-generated text. Based on the aforementioned evaluation metrics, the target generated text is determined; The step of generating target multimedia data corresponding to the target project based on the multimedia-generated text includes: Based on the target text, target multimedia data corresponding to the target project is generated.
10. The method according to claim 9, characterized in that, The determination of the target generated text based on the evaluation metrics includes: If the multimedia-generated text fails to be generated based on the evaluation metrics, text modification suggestions are obtained. Based on the text modification suggestions, the target generated text is generated.
11. The method according to any one of claims 1 to 10, characterized in that, Before generating the target multimedia data corresponding to the target project based on the multimedia generation theme and the multimedia reference information, the method further includes: Obtain a multimedia generation framework, wherein the multimedia generation framework is used to constrain the multimedia data generation structure of the target project; The step of generating target multimedia data corresponding to the target project based on the multimedia generation theme and the multimedia reference information includes: Based on the multimedia generation theme, the multimedia reference information, and the multimedia generation framework, target multimedia data corresponding to the target project is generated.
12. A multimedia data generation device, characterized in that, include: The first acquisition module is configured to acquire the project topic, project marketing information and multimedia data corresponding to the project topic of the target project. The determination module is configured to determine a multimedia generation theme based on the project topic and the project marketing information, wherein the multimedia generation theme is used to constrain the direction of multimedia data generation for the target project; The parsing module is configured to parse the multimedia data to obtain multimedia reference information; The generation module is configured to generate target multimedia data corresponding to the target project based on the multimedia generation theme and the multimedia reference information.
13. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The device stores a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.
15. A computer program product, characterized in that, Includes a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.