Multi-modal large model-based one-stop lightweight marketing video generation system

By integrating multimodal large models, a one-stop lightweight marketing video generation system can achieve end-to-end automated generation of marketing videos, solving the problems of cumbersome production processes and high costs in existing technologies, and improving video production efficiency and quality.

CN121309933APending Publication Date: 2026-01-09ZHEJIANG FINANCIAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511575256.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Current marketing video production relies on manual creativity and cumbersome post-production compositing, resulting in fragmented processes, high costs, a lack of personalized expression, and difficulty in quickly responding to market demands.

Method used

This one-stop lightweight marketing video generation system, based on a multimodal big model, integrates learning of trending elements, intent analysis and script generation, material generation and retrieval, video synthesis and editing, and feedback optimization to achieve end-to-end automated generation with single image input and direct video output.

Benefits of technology

It lowers the barrier to entry and time required for producing high-quality marketing videos, achieves a high degree of alignment between copywriting and product characteristics, and is applicable to e-commerce, short video content creation, advertising media, social media marketing, and other fields, providing powerful productivity tools for SMEs and individual creators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121309933A_ABST
    Figure CN121309933A_ABST
Patent Text Reader

Abstract

The invention discloses a one-stop lightweight marketing video generation system based on a multi-modal large model, and relates to the technical field of artificial intelligence. Comprising a blasting element learning module, an intention analysis and script generation module, a material generation and retrieval module, a video synthesis and editing module and a feedback optimization module. According to the method, commodity popular element learning, intention analysis and script generation, material generation and retrieval, video synthesis and editing and feedback optimization are integrated in a lightweight system, end-to-end automatic generation of single-image input and direct video is realized, the production threshold and period of a high-quality marketing video are greatly reduced, and the production efficiency of the high-quality marketing video is improved. The method can be widely applied to the fields of e-commerce, short video content creation, advertisement media, social media marketing and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically to a one-stop lightweight marketing video generation system based on a multimodal large model. Background Technology

[0002] Marketing videos refer to videos that businesses create or produce specifically for their products and services, then distribute them on relevant video platforms to achieve brand exposure and promotion, ultimately leading to user inquiries, purchases, and use of the product or service. If the video is engaging enough and can spark the interest of potential users, it can achieve remarkable results. Users will not only comment and interact with the video but also actively share it with others.

[0003] Video marketing takes many forms, including television commercials, online videos, promotional videos, micro-films, and short videos, which have become increasingly popular with businesses and users in recent years. With the explosive growth of short video platforms, the market's demand for high-quality, high-frequency marketing video content is becoming increasingly urgent.

[0004] Current marketing video production largely relies on manual creativity and tedious post-production compositing, or on single-function tools, requiring switching between different software to process copy, materials, and voice-over. Existing automated solutions also suffer from fragmented processes, high costs, and a lack of personalized expression, making it difficult to quickly respond to market demands.

[0005] Therefore, proposing a one-stop lightweight marketing video generation system based on a multimodal large model to solve the difficulties of existing technologies is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] In view of this, the present invention provides a one-stop lightweight marketing video generation system based on a multimodal large model, which integrates product best-selling element learning, intent parsing and script generation, material generation and retrieval, video synthesis and editing, and feedback optimization into a lightweight system, realizing end-to-end automated generation with single image input and direct video access.

[0007] To achieve the above objectives, the present invention provides the following technical solution: A one-stop lightweight marketing video generation system based on a multimodal large model includes: a module for learning viral elements, a module for intent parsing and script generation, a module for material generation and retrieval, a module for video synthesis and editing, and a feedback optimization module; among which, The trending elements learning module connects to the input end of the intent parsing and script generation module. It is used to receive the original demand text input by the user, continuously analyze and learn from videos using a large model, extract key trending elements, and build a dynamically updated trending elements knowledge base. The intent parsing and script generation module is connected to the input end of the material generation and retrieval module. It is used to parse user intent and match it with the popular element knowledge base to filter out popular paradigms. Based on the matched popular paradigms, a structured video storyboard script is automatically generated using a large model. The material generation and retrieval module is connected to the input end of the video compositing and editing module. It is used to generate or retrieve multi-track materials that match the script containing the viral template based on the script. The video compositing and editing module, connected to the input of the feedback optimization module, receives scripts and multi-track materials and automatically integrates them into a final, coherent marketing video that conforms to the rhythm and visual logic of a viral video. The feedback optimization module is used to provide feedback and verify the effectiveness of popular elements.

[0008] Optionally, the viral element learning module collects data from over 100 highly interactive viral marketing videos with likes, comments, and shares based on the public APIs of social media platforms. Using Natural Language Processing (NLP) technology, it analyzes the copywriting, subtitles, tags, and comment sections of the viral marketing videos to extract key viral elements. The extracted key viral elements are then correlated and structured to form a viral script paradigm that can be understood and applied by computers, thereby constructing a viral element knowledge base.

[0009] Optionally, key elements of a viral hit include structural features, content features, and interactive features; among them, Structural features include opening pattern, video rhythm, and distribution of total video duration; Content-related features include emotional keywords, the way product selling points are presented, and narrative structure; Interactive features include language that guides comments and incentives that encourage users to share.

[0010] Optional, the viral element learning module uses a large language model (LLM) to analyze user intent and extract key viral elements, including product selling points, target audience, video tone, and duration.

[0011] Optionally, the intent parsing and script generation module uses a large model to automatically generate structured video storyboard scripts, including a description of each shot, narration text, background music suggestions, and transition effects.

[0012] Optionally, the material generation and retrieval module includes a video generation submodule, an audio generation submodule, and a cross-modal retrieval submodule connected in sequence; wherein, The video generation submodule generates original video clips or images based on the scene descriptions in the script using the Wensheng image model, and automatically optimizes the generated prompts according to marketing intentions to ensure that the generated content meets the user's original needs. The audio generation submodule uses a text-to-speech (TTS) model to convert narration text into natural human voice dubbing, and adjusts the speech rate, timbre, and emotion. The cross-modal retrieval submodule has a built-in copyright-clear material library. Utilizing cross-modal retrieval technology, it intelligently searches the material library for matching existing video, image, and audio materials based on the user's original text description of their needs.

[0013] Optionally, a feedback optimization module integrates lightweight A / B testing of video effects generated by different viral content paradigms. The test results are fed back to the viral content element knowledge base to verify the effectiveness of the viral content elements.

[0014] As can be seen from the above technical solution, compared with the prior art, the present invention provides a one-stop lightweight marketing video generation system based on a multimodal large model, which has the following beneficial effects: (1) This invention innovatively integrates product best-selling element learning, intent parsing and script generation, material generation and retrieval, video synthesis and editing, and feedback optimization into a lightweight system, realizing end-to-end automated generation with single image input and direct video access; (2) Through large model technology, a deep understanding and conversion of visual information and language semantics was achieved, ensuring a high degree of fit between the copywriting and product characteristics, as well as a precise match between video content and copywriting theme; (3) This invention greatly reduces the production threshold and cycle of high-quality marketing videos. It can be widely used in e-commerce, short video content creation, advertising media, social media marketing and other fields. The market demand is clear and huge, providing small and medium-sized enterprises and individual creators with a powerful tool for creating hit products. It has huge commercial conversion potential. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0016] Figure 1 A schematic diagram of the structure of a one-stop lightweight marketing video generation system based on a multimodal large model provided by the present invention; Figure 2 This is a schematic diagram of the structure of the material generation and retrieval module provided by the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] See Figure 1 As shown, this invention discloses a one-stop lightweight marketing video generation system based on a multimodal large model, including: a popular element learning module, an intent parsing and script generation module, a material generation and retrieval module, a video synthesis and editing module, and a feedback optimization module; wherein, The trending elements learning module connects to the input end of the intent parsing and script generation module. It is used to receive the original demand text input by the user, continuously analyze and learn from videos using a large model, extract key trending elements, and build a dynamically updated trending elements knowledge base. The intent parsing and script generation module is connected to the input end of the material generation and retrieval module. It is used to parse user intent and match it with the popular element knowledge base to filter out popular paradigms. Based on the matched popular paradigms, a structured video storyboard script is automatically generated using a large model. The material generation and retrieval module is connected to the input end of the video compositing and editing module. It is used to generate or retrieve multi-track materials that match the script containing the viral template based on the script. The video compositing and editing module, connected to the input of the feedback optimization module, receives scripts and multi-track materials and automatically integrates them into a final, coherent marketing video that conforms to the rhythm and visual logic of a viral video. The feedback optimization module is used to provide feedback and verify the effectiveness of popular elements.

[0019] Furthermore, the viral element learning module collects data from over 100 highly interactive viral marketing videos with likes, comments, and shares based on the public APIs of social media platforms. Using Natural Language Processing (NLP) technology, it analyzes the copywriting, subtitles, tags, and comment sections of the viral marketing videos to extract key viral elements. These extracted key elements are then correlated and structured to form a viral script paradigm that can be understood and applied by computers, thereby constructing a viral element knowledge base.

[0020] Furthermore, key elements of a hit product include structural features, content features, and interactive features; among them, Structural features include opening pattern, video rhythm, and distribution of total video duration; Content-related features include emotional keywords, the way product selling points are presented, and narrative structure; Interactive features include language that guides comments and incentives that encourage users to share.

[0021] Furthermore, the viral element learning module uses a large-scale language model (LLM) to analyze user intent and extract key viral elements, including product selling points, target audience, video tone, and duration.

[0022] Furthermore, the intent parsing and script generation module uses a large model to automatically generate structured video storyboard scripts, including scene descriptions, narration text, background music suggestions, and transition effects for each shot.

[0023] Furthermore, such as Figure 2 As shown, the material generation and retrieval module includes a video generation submodule, an audio generation submodule, and a cross-modal retrieval submodule connected in sequence; wherein, The video generation submodule generates original video clips or images based on the scene descriptions in the script using the Wensheng image model, and automatically optimizes the generated prompts according to marketing intentions to ensure that the generated content meets the user's original needs. The audio generation submodule uses a text-to-speech (TTS) model to convert narration text into natural human voice dubbing, and adjusts the speech rate, timbre, and emotion. The cross-modal retrieval submodule has a built-in copyright-clear material library. Utilizing cross-modal retrieval technology, it intelligently searches the material library for matching existing video, image, and audio materials based on the user's original text description of their needs.

[0024] Furthermore, the feedback optimization module integrates lightweight A / B testing of video effects generated by different viral content paradigms. The test results are fed back to the viral content element knowledge base to verify the effectiveness of the viral content elements.

[0025] In one specific embodiment, the following is included: Users log in to the system's web interface and enter in the input box: "Create a 15-second short video for a newly launched coffee called 'Delicious,' highlighting its rich aroma and invigorating effect. The style should be warm, healing, and beautiful." After receiving the request, the popular elements learning module first identifies the keywords "delicious coffee", "rich and aromatic taste", "refreshing", "warm, healing and wonderful", and "15 seconds". Then, the module's "popular elements knowledge base" is activated. Knowledge Base Background: Through continuous learning of recent hit videos in the beverage industry, the knowledge base has summarized several effective paradigms. For example, for "new beverage products," a paradigm called "immersive sensory enticement" has been proven to be highly effective. The key elements of this paradigm include: Opening (0-3 seconds): Close-up shot highlighting the product’s most appealing physical attributes, such as coffee beans and milk.

[0026] Mid-section (4-12 seconds): Emphasize the production process or usage scenario, accompanied by strong sensory adjectives such as "listen to this sound" or "feel this refreshing feeling".

[0027] Ending (13-15 seconds): Show satisfaction and give a clear action instruction such as "Come and try it"; Based on the intent parsing and script generation module, a storyboard script containing three shots is generated based on keywords: Scene 1 (5 seconds): Scene description: The morning sunlight shines through the window onto a steaming cup of coffee, with coffee beans beside it. Narration: "Awakening the sleeping morning." Scene 2 (5 seconds): Scene description: slow motion close-up, milk is poured into coffee to form beautiful latte art, narration: "Savor the 'delicious' aroma"; Scene 3 (5 seconds): Scene description: A young man takes a sip of coffee, a satisfied and energetic smile on his face, and looks out the window at the hopeful morning light. Narration: "Full of energy all day long." The video generation submodule uses the descriptive words of the three shots to call the Wensheng video model to generate three 5-second video clips; the audio generation submodule uses TTS technology to synthesize the narration into a warm female voice-over and generate a gentle piano piece as background music. Arrange the three video clips in sequence, embed the voice-over and background music tracks, automatically generate synchronized subtitles, add fade-in and fade-out transition effects, and finally output a 15-second MP4 video. If users are satisfied with the generated video, they can download and use it directly. If they are not satisfied, they can instruct the system to generate another version, such as changing the video style to a creative and fun video, an emotionally resonant video, or a visually stunning video, or manually fine-tune the script and regenerate it. The system can also generate two different opening versions at the same time for subsequent A / B testing.

[0028] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0029] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A one-stop lightweight marketing video generation system based on a multimodal large model, characterized in that, include: The module includes modules for learning trending elements, intent parsing and script generation, material generation and retrieval, video compositing and editing, and feedback optimization. The trending elements learning module connects to the input end of the intent parsing and script generation module. It is used to receive the original demand text input by the user, continuously analyze and learn from videos using a large model, extract key trending elements, and build a dynamically updated trending elements knowledge base. The intent parsing and script generation module is connected to the input end of the material generation and retrieval module. It is used to parse user intent and match it with the popular element knowledge base to filter out popular paradigms. Based on the matched popular paradigms, a structured video storyboard script is automatically generated using a large model. The material generation and retrieval module is connected to the input end of the video compositing and editing module. It is used to generate or retrieve multi-track materials that match the script containing the viral template based on the script. The video compositing and editing module, connected to the input of the feedback optimization module, receives scripts and multi-track materials and automatically integrates them into a final, coherent marketing video that conforms to the rhythm and visual logic of a viral video. The feedback optimization module is used to provide feedback and verify the effectiveness of popular elements.

2. The one-stop lightweight marketing video generation system based on a multimodal large model according to claim 1, characterized in that, The viral marketing element learning module collects data from over 100 highly interactive viral marketing videos with likes, comments, and shares based on the public APIs of social media platforms. Using Natural Language Processing (NLP) technology, it analyzes the copywriting, subtitles, tags, and comment sections of the viral marketing videos to extract key viral elements. These extracted key elements are then correlated and structured to form a viral script paradigm that can be understood and applied by computers, thereby building a viral element knowledge base.

3. The one-stop lightweight marketing video generation system based on a multimodal large model according to claim 2, characterized in that, Key elements of a viral hit include structural features, content features, and interactive features; among them, Structural features include opening pattern, video rhythm, and distribution of total video duration; Content-related features include emotional keywords, the way product selling points are presented, and narrative structure; Interactive features include language that guides comments and incentives that encourage users to share.

4. The one-stop lightweight marketing video generation system based on a multimodal large model according to claim 1, characterized in that, The viral content learning module uses a large-scale language model (LLM) to analyze user intent and extracts key viral elements, including product selling points, target audience, video tone, and duration.

5. The one-stop lightweight marketing video generation system based on a multimodal large model according to claim 1, characterized in that, The intent parsing and script generation module uses a large model to automatically generate structured video storyboard scripts, including a description of each shot, narration text, background music suggestions, and transition effects.

6. The one-stop lightweight marketing video generation system based on a multimodal large model according to claim 1, characterized in that, The material generation and retrieval module includes a video generation submodule, an audio generation submodule, and a cross-modal retrieval submodule connected in sequence; among them, The video generation submodule generates original video clips or images based on the scene descriptions in the script using the Wensheng image model, and automatically optimizes the generated prompts according to marketing intentions to ensure that the generated content meets the user's original needs. The audio generation submodule uses a text-to-speech (TTS) model to convert narration text into natural human voice dubbing, and adjusts the speech rate, timbre, and emotion. The cross-modal retrieval submodule has a built-in copyright-clear material library. Utilizing cross-modal retrieval technology, it intelligently searches the material library for matching existing video, image, and audio materials based on the user's original text description of their needs.

7. The one-stop lightweight marketing video generation system based on a multimodal large model according to claim 1, characterized in that, The feedback optimization module integrates lightweight A / B testing of video effects generated by different viral content paradigms. The test results are fed back to the viral content element knowledge base to verify the effectiveness of the viral content elements.