An AI short video generation method and system for collaborative optimization of narrative rhythm and dissemination transformation double targets
Patent Information
- Application Number
- CN202611282498.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-24
- Publication Date
- 2026-09-29
AI Technical Summary
但是,其优化对象通常是媒体内容整体价值或不同内容版本的评价,并未具体限定如何在AI短视频生成过程中将反馈数据分别回写至叙事模板库、镜头节奏模板库、转化触发片段库和素材权重库,也未公开在生成阶段对镜头级控制参数进行最小改动投影修复的机制
[0023]本发明的有益技术效果在于:本发明通过将自然语言叙事文本解析为起因、冲突、转折、高潮和行动引导五类叙事事件锚点,并将各叙事事件锚点映射至短视频时间轴的相对区间,使AI短视频生成过程不再仅依赖提示词或素材自动拼接,而是具有明确的叙事结构约束,能够减少现有生成视频中常见的高潮前置、转折缺失、行动引导生硬插入以及镜头节奏断裂等问题;同时,本发明引入传播转化评价指标,将预测点击率、有效停留率、完播率、分享率和行动转化率等平台传播因素与叙事节奏评价指标共同纳入生成控制过程,使短视频生成结果不仅具备较好的故事表达和情绪推进效果,还能够更好地适配目标平台的用户观看习惯和传播目标;进一步地,当叙事节奏目标与传播转化目标发生冲突时,本发明并非重新生成完整视频或简单调整综合权重,而是在保持五类叙事事件锚点前后顺序和已选多模态素材主体内容不变的前提下,仅对镜头时长、字幕出现时刻、音频节拍对齐点、关键事件镜头位置和行动引导片段位置等镜头级控制参数进行最小改动投影修复,从而在较小改动范围内同时满足叙事节奏阈值和传播转化阈值,降低重复生成成本和内容漂移风险,提高生成结果的可控性、稳定性和可解释性;此外,本发明还可将实际传播反馈数据分层回写至叙事模板库、镜头节奏模板库、转化触发片段库和素材权重库,使系统能够在后续生成任务中持续优化叙事结构、镜头节奏、行动引导策略和素材选择优先级,形成面向不同目标平台和传播目标的闭环优化能力,由此显著提升AI短视频在叙事完整性、观看流畅性、内容一致性、传播适配性和转化达成率方面的综合表现。
Smart Images

Figure CN122845896A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to AI short video generation methods and systems, and more particularly to an AI short video generation method and system that achieves collaborative optimization of narrative rhythm and dissemination transformation as dual objectives. Background Technology
[0002] With the rapid development of mobile internet, short video platforms, and generative artificial intelligence technologies, short videos have become an important medium for brand communication, e-commerce marketing, knowledge dissemination, and personal content creation. Traditional short video production mainly relies on manual directing, scriptwriting, material selection, shot editing, subtitle dubbing, and release testing. This process is time-consuming and costly, and different creators have varying understandings of narrative rhythm and conversion guidance, leading to significant fluctuations in the dissemination effect of the same material across different platforms. In recent years, technologies such as text-to-video, image-to-video, automatic video editing, intelligent dubbing, and automatic subtitle generation have matured, lowering the barrier to entry for short video production to some extent. Users can automatically obtain video content simply by inputting scripts, prompts, or sets of materials. However, existing AI short video generation technologies mostly emphasize image quality, shot coherence, semantic consistency, or automatic material matching, paying insufficient attention to metrics such as clicks, dwell time, completion time, sharing, and action conversion in actual platform dissemination. This makes it difficult to simultaneously meet both the goals of complete narrative and effective dissemination and conversion.
[0003] For example, US patent application US20240185306A1 discloses a system and method for automatically generating short videos. This system can extract text from text or web pages, select summary sentences, generate short video scenes, access media assets, and combine multiple video scenes into a short video. This solution is suitable for e-commerce or web content converted into short videos, solving the problem of automatically combining text content into short video assets. The document also covers processes such as short video scene selection, media asset selection, and video compilation, reflecting the basic technical approach of "text-driven + material selection + short video generation" in existing technologies. However, the focus of this type of solution is on how to generate scenes based on input text and how to select and combine media assets. It typically does not break down the narrative events within the short video script into event anchors such as cause, conflict, turning point, climax, and action guidance, nor does it set timeline constraints for various narrative anchors during the generation process. Furthermore, it does not address conflicts between narrative rhythm goals and dissemination conversion goals by making minimal modifications to adjust shot duration, subtitle appearance time, audio beat alignment, key event shot positions, and action guidance segment positions. Therefore, although this solution can improve the efficiency of short video generation, the generated results may still have problems such as unreasonable narrative climax positions, premature or late action guidance, and video rhythm that does not match the viewing habits of platform users.
[0004] For example, US Patent 11769528B2 discloses an automated video editing technology that utilizes machine learning algorithms to process video content, understand the video within its semantic and cultural context, identify temporal events within the video, and generate narrative editing sequences by connecting or interleaving different temporal events. This technology demonstrates that machine learning can be used to identify narrative progression, event connections, and segment organization in videos, thereby assisting in generating more narrative editing results. Compared to traditional manual editing, this technology has the advantage of automatically identifying events and forming narrative sequences. However, it primarily focuses on semantic understanding and narrative editing of existing video materials and does not construct dissemination conversion evaluation indicators for short video dissemination scenarios, nor does it coordinate the predicted results such as click-through rate, effective dwell rate, completion rate, sharing rate, or action conversion rate with narrative rhythm indicators.
[0005] Furthermore, US patent application US20210049627A1 discloses a media content evaluation and optimization system that can present media content to online community users, monitor user behavior, comments, votes, or other feedback signals, and evaluate and optimize media content based on user feedback. The document also mentions the use of statistical machine learning and regression methods to estimate the value of media content and supports continuous measurement and improvement of content. This type of solution demonstrates that dissemination feedback and user behavior data can be used for media content evaluation and optimization, providing feedback for content production. However, its optimization targets are usually the overall value of media content or the evaluation of different content versions, without specifically defining how to write feedback data back to the narrative template library, shot rhythm template library, transformation trigger segment library, and material weight library during the AI short video generation process. It also does not disclose a mechanism for minimally modifying the projection repair of shot-level control parameters during the generation stage. Therefore, while this type of technology focuses on user feedback, it lacks refined control over the internal structure of the short video generation process, especially lacking technical means to perform local repairs while maintaining the order of narrative event anchor points and the main content of the material. Summary of the Invention
[0006] The technical objective of this invention is to provide an AI short video generation method and system that achieves collaborative optimization of narrative rhythm and dissemination conversion goals. By dividing natural language narrative text into narrative event anchors such as cause, conflict, turning point, climax, and action guidance, and combining dissemination conversion evaluation indicators for collaborative control of dual goals, when the narrative rhythm goal and the dissemination conversion goal conflict, a shot-level minimal modification projection repair mechanism is adopted to locally adjust the shot duration, subtitle appearance time, audio beat alignment point, key event shot position, and action guidance segment position. This improves the narrative integrity, viewing smoothness, and platform dissemination conversion effect of AI short videos without destroying the original narrative structure and main content.
[0007] To achieve the objectives of this invention, the following technical solution is adopted:
[0008] An AI short video generation method that collaboratively optimizes the dual objectives of narrative rhythm and dissemination conversion includes the following steps:
[0009] S1: Receives natural language narrative text, multimodal materials, target platform type, and propagation target parameters;
[0010] S2: Perform semantic analysis on the narrative text, divide it into five categories of narrative event anchors: cause, conflict, turning point, climax and action guidance, and generate narrative structure vector and narrative rhythm evaluation index;
[0011] S3: Generate dissemination conversion evaluation indicators based on historical dissemination data, user interaction data, and target platform type;
[0012] S4: Construct a dual-objective loss function based on narrative rhythm evaluation index and dissemination conversion evaluation index, and trigger shot-level minimum modification projection repair when a conflict between narrative rhythm objective and dissemination conversion objective is detected to obtain the optimal control parameters;
[0013] S5: Based on the optimal control parameters, control the shot duration, subtitle density, audio beat alignment, key event shot position, and action guidance segment position, and fuse and edit the multimodal materials to generate a short video.
[0014] As a further improvement, in step S2, the five types of narrative event anchors are mapped to different relative intervals on the short video timeline. The cause anchor is located in the first 0% to 20% of the total video duration, the conflict anchor is located in the 15% to 45% interval, the turning point anchor is located in the 35% to 65% interval, the climax anchor is located in the 55% to 85% interval, and the action guidance anchor is located in the 75% to 100% interval. The narrative rhythm evaluation index is obtained by weighting semantic density, emotional fluctuation, key event location matching, and rhythm break penalty. The semantic density index represents the amount of effective semantic information per unit time; the emotional fluctuation index represents the change in emotional intensity between adjacent narrative segments; the key event location matching index represents the degree of matching between each narrative event anchor and the preset time interval; and the rhythm break penalty represents the degree to which the semantic or emotional jump between adjacent shots exceeds a threshold. All weighting coefficients are non-negative.
[0015] As a further improvement, in step S3, the dissemination conversion evaluation index is not determined solely by the click-through rate (CTR), but is a composite index based on clicks, effective dwell time, completion rate, sharing, and action conversion. The dissemination conversion evaluation index is obtained by weighting the predicted CTR, predicted effective dwell time, predicted completion rate, predicted sharing rate, and predicted action conversion rate. Specifically, the predicted CTR represents the predicted probability that a user will click on the short video; the predicted effective dwell time represents the predicted probability that a user will effectively stay and watch the video; the predicted completion rate represents the predicted probability that a user will watch the entire short video; the predicted sharing rate represents the predicted probability that a user will share the short video; and the predicted action conversion rate represents the predicted probability that a user will perform a click, inquiry, purchase, follow, or claim action. Each weight coefficient is a non-negative weight coefficient determined according to the target platform type, and the sum of all weight coefficients is 1.
[0016] As a further improvement, in step S4, the dual-objective loss function includes narrative rhythm loss, propagation conversion loss, narrative event anchor point constraint loss, and minimum modification loss; wherein, narrative rhythm loss is used to represent the deviation between the narrative rhythm evaluation index and the target narrative rhythm range; propagation conversion loss is used to represent the deviation between the propagation conversion evaluation index and the propagation target parameters; narrative event anchor point constraint loss is used to represent the degree to which the five types of anchor points—cause, conflict, turning point, climax, and action guidance—deviate from the preset time interval; minimum modification loss is used to limit the change range of video generation control parameters before and after the repair; each of the above loss terms has a corresponding weight coefficient; step In S4, the conflict between the narrative rhythm objective and the dissemination conversion objective is triggered under the following conditions: when the candidate video scheme meets the following conditions: the narrative rhythm evaluation index is not lower than the narrative rhythm qualification threshold and the dissemination conversion evaluation index is lower than the dissemination conversion qualification threshold, or the dissemination conversion evaluation index is not lower than the dissemination conversion qualification threshold and the narrative rhythm evaluation index is lower than the narrative rhythm qualification threshold, it is determined that there is a dual objective conflict; wherein, the narrative rhythm qualification threshold is used to determine whether the narrative rhythm meets the standard, and the dissemination conversion qualification threshold is used to determine whether the dissemination conversion meets the standard; after determining that there is a dual objective conflict, the entire video is not directly regenerated, but the shot-level minimum modification projection repair process is entered.
[0017] As a further improvement, the shot-level minimal modification projection repair process includes: maintaining the order of narrative event anchor points and the main content of the selected multimodal materials unchanged, and only adjusting the shot duration, shot transition interval, subtitle appearance time, audio beat alignment point, key event shot position, and action guidance segment position; the shot-level minimal modification projection repair process uses the candidate control parameter set before repair as a benchmark to solve for the control parameter set after repair, minimizing the weighted difference between the control parameter set after repair and the candidate control parameter set before repair, while simultaneously satisfying the following constraint: the narrative rhythm evaluation index after repair is not lower than the qualified threshold for narrative rhythm. The post-repair dissemination conversion evaluation index is no lower than the qualified threshold for dissemination conversion, and the constraints of the order and time interval of the five types of narrative event anchor points are all satisfied. Among them, the candidate control parameter set before restoration is the video generation control parameter set before projection restoration; the control parameter set after restoration is the video generation control parameter set after projection restoration; weighted differences are used to represent the change range of different control parameters before and after restoration; the weight matrix is used to represent the change weight of different control parameters; the constraints of the order and time interval of the five types of narrative event anchor points are all satisfied, which means that the order of the five types of narrative event anchor points (cause, conflict, turning point, climax, and action guidance) and their corresponding time intervals all meet the preset requirements.
[0018] As a further improvement, in step S5, the optimal control parameters include at least shot-level control parameters, subtitle-level control parameters, audio-level control parameters, and transition trigger control parameters; the shot-level control parameters include the duration of each shot, the scene switching method, and the shot order; the subtitle-level control parameters include the subtitle appearance time, subtitle density, and keyword highlighting position; the audio-level control parameters include the background music beat points, sound effect insertion points, and voice pause points; the transition trigger control parameters include the insertion position of the action guidance segment, its duration, and the time interval between it and the climax anchor point.
[0019] As a further improvement, after the short video is generated and published, actual dissemination feedback data is collected and written back to the narrative template library, the shot rhythm template library, the conversion trigger segment library, and the material weight library, respectively. Among them, the narrative template library is used to update the recommendation order and time interval of five types of narrative event anchor points, the shot rhythm template library is used to update the shot duration and switching interval, the conversion trigger segment library is used to update the insertion strategy of action guidance segments, and the material weight library is used to update the selection priority of different types of materials on different platforms.
[0020] Another aspect of the present invention provides an AI short video generation system for co-optimizing narrative rhythm and dissemination conversion dual objectives. This system implements the method described above, comprising: an input module for receiving natural language narrative text, multimodal materials, target platform type, and dissemination target parameters; a narrative anchor point analysis module for dividing the narrative text into five categories of narrative event anchor points: cause, conflict, turning point, climax, and action guidance, and generating a narrative structure vector and a narrative rhythm evaluation index; a dissemination conversion evaluation module for generating dissemination conversion evaluation indicators based on historical dissemination data, user interaction data, and target platform type; a dual-objective conflict detection module for determining whether the narrative rhythm objective and the dissemination conversion objective conflict; and a projection repair module for making minimal changes to the shot-level generation control parameters when a conflict is detected. The system includes a projection restoration module; a video generation module for fusing and editing multimodal materials based on the restored optimal control parameters to generate a short video; and a feedback write-back module for writing back actual propagation feedback data to the narrative template library, shot rhythm template library, conversion trigger segment library, and material weight library, respectively. The projection restoration module includes a parameter freezing unit, a local adjustment unit, and a constraint verification unit. The parameter freezing unit is used to freeze the order of narrative event anchor points and the main content of selected multimodal materials. The local adjustment unit is used to adjust shot duration, shot switching interval, subtitle appearance time, audio beat alignment point, key event shot position, and action guidance segment position. The constraint verification unit is used to determine whether the restored control parameters simultaneously meet the narrative rhythm threshold, propagation conversion threshold, and narrative event anchor point constraints.
[0021] A third aspect of the present invention provides a computer device including a processor, a graphics processing unit (GPU), and a memory, wherein the memory stores a computer program that, when executed by the processor and the GPU, causes the computer device to perform the method described thereon.
[0022] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a computer device, causes the computer device to perform the method described thereon.
[0023] The beneficial technical effects of this invention are as follows: By parsing natural language narrative text into five categories of narrative event anchors—cause, conflict, turning point, climax, and action guidance—and mapping each narrative event anchor to a relative interval on the short video timeline, the AI short video generation process no longer relies solely on prompts or automatic splicing of materials, but has clear narrative structure constraints. This reduces common problems in existing generated videos, such as premature climaxes, missing turning points, abrupt insertion of action guidance, and broken camera rhythms. Simultaneously, this invention introduces dissemination conversion evaluation indicators, incorporating platform dissemination factors such as predicted click-through rate, effective dwell time, completion rate, sharing rate, and action conversion rate, along with narrative rhythm evaluation indicators, into the generation control process. This ensures that the generated short video not only possesses better story expression and emotional progression but also better adapts to the viewing habits and dissemination goals of the target platform's users. Furthermore, when narrative rhythm goals conflict with dissemination conversion goals, this invention does not regenerate a complete video or simply adjust... Instead of adjusting the overall weight, this invention maintains the order of the five types of narrative event anchors and the main content of the selected multimodal materials. It only modifies and projects minimally the shot-level control parameters such as shot duration, subtitle appearance time, audio beat alignment, key event shot positions, and action guidance segment positions. This allows for simultaneous satisfaction of narrative rhythm thresholds and dissemination conversion thresholds within a relatively small range of changes, reducing the cost of repeated generation and the risk of content drift, and improving the controllability, stability, and interpretability of the generated results. Furthermore, this invention can hierarchically write back actual dissemination feedback data to the narrative template library, shot rhythm template library, conversion trigger segment library, and material weight library. This enables the system to continuously optimize the narrative structure, shot rhythm, action guidance strategy, and material selection priority in subsequent generation tasks, forming a closed-loop optimization capability for different target platforms and dissemination goals. This significantly improves the overall performance of AI short videos in terms of narrative integrity, viewing smoothness, content consistency, dissemination adaptability, and conversion achievement rate. Attached Figure Description
[0024] Figure 1 This is a schematic diagram of the system structure of the present invention;
[0025] Figure 2 This is a schematic diagram of the method flow of the present invention;
[0026] Figure 3 This is a schematic diagram of the narrative event plotting points and timeline mapping of the present invention;
[0027] Figure 4 This is a flowchart of the dual-target conflict detection and projection repair process of the present invention. Detailed Implementation
[0028] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.
[0029] I. Terminology Explanation
[0030] To facilitate understanding of this invention, the following terms will be explained first.
[0031] Natural Language Narrative Text , refers to the text used to express the theme, storyline, product selling points, user pain points, scene descriptions, conversion prompts, etc. of short videos. It can be scripts directly entered by users, product detail page text, marketing copy, course introductions, knowledge point summaries, or short video script text pre-generated by a large language model.
[0032] Multimodal materials "Short video content" refers to the collection of materials used in short video generation, including but not limited to images, video clips, product images, voiceover clips, background music, sound effects, subtitle templates, stickers, transition templates, brand logos, action guide buttons, or other visual, auditory, and textual materials.
[0033] Target platform type This refers to the platform category on which the short video plan will be released, such as e-commerce platforms, social short video platforms, knowledge-sharing platforms, and embedded video channels on brand websites. Different platforms may have different user viewing behaviors, recommendation rules, duration preferences, completion weight, and interaction weight.
[0034] Propagation target parameters This refers to the dissemination goals set by the user or the system, which may include goals such as clicks, valid dwell time, completion of broadcasts, sharing, favorites, comments, redirects, order placement, reservations, and following. It may also include the priority and threshold of each goal.
[0035] Narrative event anchors refer to key event points with structural functions extracted from narrative text. This invention preferably divides the narrative structure of short videos into five categories of narrative event anchors: cause, conflict, turning point, climax, and action guidance. Cause anchors are used to explain the background or source of the need; conflict anchors are used to showcase pain points, problems, or contrasts; turning point anchors are used to introduce solutions or core selling points; climax anchors are used to present the strongest evidence, key benefits, or visual high points; and action guidance anchors are used to guide users to click, inquire, purchase, follow, or continue watching.
[0036] Narrative Structure Vector It refers to a vectorized expression composed of information such as five types of narrative event anchors, anchor order, anchor confidence, anchor position in the text, anchor corresponding material category, and anchor target interval on the timeline.
[0037] Narrative rhythm evaluation index It refers to a comprehensive evaluation index used to measure whether the narrative structure of a short video is complete, whether the pacing is smooth, whether the key events are placed reasonably, and whether the emotional progression is natural.
[0038] Dissemination and Transformation Evaluation Indicators It refers to a comprehensive evaluation index used to measure the predictive performance of candidate short videos on the target platform in terms of clicks, dwell time, completion time, sharing, or action conversion.
[0039] Optimal control parameters , refers to a set of shot-level parameters used to control the short video generation process, including shot duration, shot order, subtitle appearance time, subtitle density, audio beat alignment point, key event shot position, action guidance segment position, transition method, and material selection weight, etc.
[0040] Shot-level minimal modification projection restoration refers to a local restoration process that, after detecting a conflict between narrative rhythm objectives and propagation conversion objectives, uses the original control parameters of the candidate video scheme as a benchmark and, under the conditions of satisfying narrative anchor point order constraints, narrative rhythm thresholds, and propagation conversion thresholds, minimizes the magnitude of changes to control parameters. This process differs from regenerating a complete video; it only makes local corrections to adjustable shot-level parameters.
[0041] II. System Structure
[0042] Combination Figure 1 The system of this invention may include an input module, a narrative anchor point analysis module, a propagation transformation evaluation module, a dual-objective conflict detection module, a projection repair module, a video generation module, and a feedback write-back module. Each module can be deployed on the same server, or it can be implemented by combining cloud-based model services with a local video editing engine.
[0043] The input module is used to receive natural language narrative text. Multimodal materials Target platform type and propagation target parameters The input module can also perform material preprocessing, such as frame extraction, shot segmentation, sharpness detection, subject detection, and duration annotation for video materials; subject recognition, background segmentation, and composition scoring for image materials; beat detection, volume normalization, and emotion tag extraction for audio materials; and sentence segmentation, noise reduction, sensitive word filtering, and keyword extraction for text materials.
[0044] The narrative anchor analysis module is used to analyze natural language narrative texts. Narrative event anchors are categorized into five types: cause, conflict, turning point, climax, and action guidance, and a narrative structure vector is generated accordingly. Narrative rhythm evaluation indicators This module can include a text encoding unit, an event classification unit, an anchor point localization unit, a timeline mapping unit, and a rhythm scoring unit. The text encoding unit can employ a pre-trained language model, such as a Transformer encoder, a pre-trained Chinese language model, a domain-fine-tuned language model, or a large language model embedding interface; the event classification unit is used to determine which type of narrative event anchor each sentence, phrase, or semantic fragment belongs to; the anchor point localization unit is used to determine the start and end positions and confidence levels of each type of anchor point in the text; the timeline mapping unit is used to map anchor points to relative intervals on the short video timeline; and the rhythm scoring unit is used to calculate semantic density, emotional fluctuation, key event position matching degree, and rhythm break penalty term.
[0045] The communication conversion evaluation module is used to evaluate communication based on historical communication data, user interaction data, and target platform type. Generate dissemination and transformation evaluation indicators This module can include a feature construction unit, a conversion prediction model, a platform weight configuration unit, and a composite scoring unit. The feature construction unit extracts video structure features, text features, material features, platform features, and historical behavior features; the conversion prediction model outputs predicted values such as click-through rate, effective dwell time, completion rate, sharing rate, and action conversion rate; the platform weight configuration unit sets different weights for different dissemination metrics based on the target platform type; and the composite scoring unit generates... .
[0046] The dual-objective conflict detection module is used to determine whether there is a conflict between the narrative pacing objective and the communication conversion objective. It does not simply calculate a composite score, but rather judges each objective separately. and Whether the corresponding threshold has been reached. When one target meets the threshold while the other does not, the system determines that there is a dual-target conflict and enters the projection repair module.
[0047] Combination Figure 4The projection restoration module includes a parameter freezing unit, a local adjustment unit, a projection solving unit, and a constraint verification unit. The parameter freezing unit freezes the order of five types of narrative event anchor points and the main content of the selected multimodal materials, preventing the narrative structure from being disrupted or the main material from shifting during the restoration process. The local adjustment unit adjusts shot duration, shot transition intervals, subtitle appearance times, audio beat alignment points, key event shot positions, and action guidance segment positions. The projection solving unit solves for the minimum modification projection model. The constraint verification unit determines whether the restored control parameters simultaneously satisfy the narrative rhythm threshold, propagation transformation threshold, and narrative event anchor point constraints.
[0048] The video generation module is used to generate videos based on the repaired optimal control parameters. For multimodal materials It performs fusion editing to generate short videos. It can call upon video editing engines, text-to-video models, image-to-video models, subtitle layout modules, speech synthesis modules, music beat alignment modules, and cover generation modules, etc.
[0049] The feedback write-back module is used to collect actual dissemination feedback data after a short video is published. The data is then written back to the narrative template library, the shot rhythm template library, the transformation trigger segment library, and the material weight library, respectively. This layered write-back mechanism allows the system to continuously optimize the narrative structure, shot rhythm, action guidance strategies, and material selection priorities in subsequent generation tasks.
[0050] III. Overall Technical Route for Implementing the Method of the Invention
[0051] Combination Figure 2 The method of the present invention includes five main steps, S1 to S5.
[0052] In S1, the system obtains text, source materials, platform information, and dissemination goals to form a generation task package. In S2, the system parses the text into five types of narrative event anchors and maps these anchors to the short video timeline, thus forming narrative structure constraints. In S3, the system uses a dissemination prediction model to calculate the dissemination conversion index of candidate solutions. In S4, the system considers both narrative rhythm goals and dissemination conversion goals; if there is a conflict between the two, it performs minimal modification projection repair. In S5, the system generates a short video based on the repaired control parameters and can perform layered feedback write-back after publication.
[0053] The following provides a more detailed explanation of each step.
[0054] 1. S1 Input Collection Steps
[0055] In one specific implementation, the user inputs natural language narrative text through an interactive interface. For example, the text reads, "Office workers are prone to drowsiness in the afternoon. Regular coffee is too stimulating and can easily cause palpitations. This product uses decaffeinated coffee beans with herbal aromas, offering a mild taste, perfect for afternoon tea in the office. Click to claim your sample." The system segments this text into several semantic fragments. Each semantic fragment can include a text location number, word count, keywords, sentiment index, and candidate material tags.
[0056] Multimodal materials Materials can be uploaded by users or retrieved from the material library. Each material in the library can be pre-labeled with information such as material type, subject category, mood style, compatible platform, resolution, duration, croppable area, brand compliance mark, and copyright status. For video clips, the system can use a shot segmentation algorithm to divide them into several usable segments; for images, the system can extract the subject outline, composition center, main color tone, and scene category; for audio, the system can extract beat points, mood tags, climax sections, and loopable segments.
[0057] Target platform type Target parameters can be specified by the user or automatically identified based on account binding information. This can include target click-through rate, target completion rate, target share rate, target inquiry rate, target purchase rate, etc., and can also be configured with modes such as "prioritize completion," "prioritize clicks," "prioritize conversion," and "balance dissemination and conversion." The system can convert these targets into thresholds and weight parameters, such as a qualified narrative pacing threshold. and the threshold for qualified propagation and conversion .
[0058] 2. S2 Narrative Rhythm Analysis Steps
[0059] In one specific implementation, the system first processes the natural language narrative text. Text cleaning and semantic segmentation are performed to obtain a set of semantic fragments. Each semantic fragment It can be a sentence, a short phrase, a selling point, or a text unit divided by punctuation and semantic boundaries. Then, the text encoding unit encodes each semantic segment into a semantic vector. The semantic vector can be output by a pre-trained language model or by an embedding interface of a large language model. To adapt to Chinese short video scripts, this implementation method preferably performs domain-specific fine-tuning on Chinese short video scripts, product descriptions, advertising copy, knowledge short video scripts, and platform title data, enabling the model to recognize common pain point expressions, transitional expressions, benefit expressions, and calls to action expressions in short videos.
[0060] The event classification unit categorizes each semantic fragment into five classes, outputting the probability that it belongs to cause, conflict, turning point, climax, and action guidance. Specifically, cause anchors typically correspond to background, scenario, or user needs, such as "office workers are prone to drowsiness in the afternoon"; conflict anchors correspond to pain points, contradictions, or contrasts, such as "regular coffee is too stimulating and can easily cause palpitations"; turning point anchors correspond to solutions, product appearance, or method change, such as "this product uses decaffeinated coffee beans paired with herbal aromas"; climax anchors correspond to the strongest selling point, proof of effectiveness, or emotional high point, such as "mild taste, suitable for afternoon tea in the office"; and action guidance anchors correspond to clicking, purchasing, following, inquiring, or receiving, such as "click to receive a sample pack".
[0061] To avoid missing or disordered anchor points of a certain type within the same text, the system can employ a sequence labeling model combined with rule-based validation. After the sequence labeling model outputs initial anchor point labels, the rule-based validation unit checks for missing, duplicate, or out-of-order anchor points. If a certain type of anchor point is missing, the system can select the candidate segment with the highest confidence from adjacent semantic segments to fill the gap. If multiple segments belong to the same anchor point, the system can select the primary anchor point based on text location, sentiment intensity, and keyword weight, using the other segments as auxiliary descriptions.
[0062] After obtaining the five types of narrative event anchors, the timeline mapping unit maps them to a relative range of the total video duration. Preferably, the causal anchor is located at the beginning of the total video duration. to The interval, the conflict anchor point is located in to The interval, the turning point is located at to The interval, the climax anchor point is located in to The interval, the action guidance anchor point is located at to Intervals. The aforementioned intervals are allowed to partially overlap to accommodate situations such as narrative compression, parallel shot expression, pre-placed subtitles, and pre-embedded action guidance in short videos. For example, in a 15-second e-commerce short video, conflict and turning point can partially overlap between the 4th and 6th seconds; in a 30-second knowledge short video, climax and action guidance can overlap between the 22nd and 26th seconds.
[0063] Based on this, the system generates narrative structure vectors. This vector can be represented as:
[0064] ;
[0065] in, to The location codes of five types of narrative event anchors in the text are respectively: cause, conflict, turning point, climax and action guidance. to These five types of anchor points are mapped to target time intervals on the short video timeline. to The confidence scores for the five anchor point categories are identified respectively. These vectors can then be used as constraint inputs for generating subsequent control parameters.
[0066] Furthermore, narrative rhythm evaluation indicators The following formula can be used for calculation:
[0067] ;
[0068] in, It is a semantic density index used to represent the amount of effective semantic information per unit time. It is an emotional fluctuation index used to represent the change in emotional intensity between adjacent narrative segments; This is a key event location matching indicator, used to represent the degree of matching between five types of narrative event anchors and their corresponding timeline intervals; This is a rhythm break penalty term, used to indicate the degree to which semantic or emotional jumps between adjacent shots exceed a preset threshold; , , , These are non-negative weighting coefficients.
[0069] In practice, This can be obtained based on the number of semantic keywords, information entropy, and the ratio of text length to shot duration. For example, the ratio of keyword count to shot duration can be calculated for the text segment corresponding to each shot, and then averaged over the entire film. It can be calculated from the sentiment intensity sequence output by the sentiment analysis model, for example, by normalizing the sentiment intensity of each semantic segment to... Then calculate the mean or weighted mean of the emotional intensity difference between adjacent segments. The score can be calculated based on the deviation between the actual landing point of the five types of anchor points and the preset time interval. The score is higher when the anchor points fall into the target interval and lower when they deviate from the interval. It can be calculated based on factors such as low semantic similarity between adjacent shots, large emotional changes, discontinuous subtitles, or sudden interruptions in audio beats.
[0070] The model training in this step can be implemented using the following structured approach. The training dataset can include manually annotated short video script text, video timelines, narrative anchor tags, platform playback data, and manual rhythm ratings. Data sources can include enterprise-owned short video samples, publicly available and licensable marketing video scripts, user-authorized uploaded video texts, and manually constructed teaching samples. Annotators label text fragments according to five categories: "cause, conflict, turning point, climax, and action guidance," and annotate the corresponding recommendation intervals in the video timeline. During training, the text fragments are input into a pre-trained language model, which outputs the probabilities of the five categories of tags. Cross-entropy loss is used to train the event classifier; simultaneously, anchor position deviation loss is used to train the timeline mapping head. After model training is complete, anchor classification accuracy, anchor recall, timeline offset error, and narrative rating relevance are calculated on the validation set. Preferably, an anchor classification accuracy of no less than 85% and a timeline interval hit rate of no less than 80% are sufficient for use in the actual generation process.
[0071] 3. S3 Communication and Transformation Evaluation Steps
[0072] S3 is used to analyze historical dissemination data, user interaction data, and target platform type. Generate dissemination and transformation evaluation indicators The key is to incorporate the propagation prediction results and narrative rhythm constraints into subsequent conflict detection and projection repair, rather than using them as separate post-launch evaluation indicators.
[0073] In one implementation, the dissemination conversion evaluation module first constructs a feature set. Features may include text features, video structure features, material features, platform features, and account history features. Text features include title length, keyword type, emotional intensity, location of action prompts, number of benefit points, and proportion of interrogative sentences; video structure features include total duration, number of shots, average shot duration, information density in the first three seconds, climax time, action prompt time, subtitle density, and beat switching frequency; material features include subject category, whether people appear on screen, whether products are clearly displayed, screen brightness, composition score, and brand exposure frequency; platform features include target platform type, recommended time slot, common completion duration preferences, and interaction weight; account history features include the account's past video average click-through rate, completion rate, sharing rate, and conversion rate.
[0074] The propagation prediction model can employ gradient boosting trees, logistic regression, deep neural networks, Wide & Deep models, Transformer sequence models, or multi-task learning models. In a preferred embodiment, the propagation prediction model uses a multi-task learning structure, sharing the underlying feature encoding layer and outputting predicted click rates separately. Predicting effective stay rate Predicted completion rate Predicted sharing rate and predict action conversion rate Multi-task learning can avoid overfitting to a single metric and allow different propagation metrics to share text, shot, and footage features.
[0075] Dissemination and Transformation Evaluation Indicators The following formula can be used for calculation:
[0076] ;
[0077] in, To predict click-through rate, To predict the effective stay rate, To predict the completion rate, To predict the sharing rate, To predict action conversion rates; , , , , To be based on the target platform type Determined non-negative weight coefficients, and satisfying .
[0078] For example, in platforms where e-commerce conversion is the primary goal, it is possible to improve and Knowledge-sharing platforms can improve... and In brand exposure scenarios, it can improve and Platform weights can be preset based on human experience or optimized through regression analysis or Bayesian methods based on the actual target contribution of historical samples. In e-commerce conversion platforms, the weights for predicted click-through rate, predicted effective dwell time, predicted completion rate, predicted share rate, and predicted action conversion rate can be set to 0.25, 0.15, 0.20, 0.10, and 0.30, respectively; in knowledge dissemination platforms, these weights can be set to 0.15, 0.25, 0.30, 0.20, and 0.10, respectively.
[0079] When training the propagation prediction model, training samples can include video segment features, release time, platform type, account type, user behavior data, and actual conversion results. Tags can include whether a video was clicked, the viewing duration, whether it was completed, whether it was shared, whether an action button was clicked, and whether a purchase or inquiry was completed. To avoid bias caused by differences in the number of followers of different accounts, the number of views, clicks, shares, and conversions can be normalized, for example, by using the average of the same account over the past 30 days as a benchmark to calculate the relative improvement rate. After training, the prediction effect can be evaluated using metrics such as AUC, mean squared error, mean absolute error, and ranking correlation coefficient.
[0080] The output of this step It does not directly determine the final generated result, but rather serves as an important input for subsequent S4 conflict detection and projection repair. Therefore, this invention differs from systems that only statistically analyze propagation effects after publication; instead, it embeds propagation prediction into the generation control process before generation.
[0081] 4. S4 Dual-Target Conflict Detection and Lens-Level Minimal Projection Repair Steps
[0082] Traditional AI short video generation methods typically generate videos directly based on text prompts and source materials, or select the highest-scoring video from multiple candidates. Even when using multi-objective optimization, a simple weighted summation method is often employed. This invention differs in that, when narrative rhythm objectives conflict with dissemination and conversion objectives, it does not regenerate a complete video, nor does it simply sacrifice one objective. Instead, while maintaining the sequential order of the five narrative event anchor points and the main content of the selected multimodal source materials, it only makes minimal changes to the controllable parameters at the shot level for projection repair.
[0083] In one specific implementation, the system first obtains candidate video schemes based on S2 and S3. and The system sets a threshold for acceptable narrative pacing. and the threshold for qualified propagation and conversion .in, It can be determined based on manually reviewed samples or percentiles of historical high-quality videos, for example, taking the 60th percentile of the narrative rhythm score of historical high-quality videos. It can be determined based on the target platform type and dissemination goals, such as taking the target percentile of the predicted conversion score for similar videos.
[0084] The conditions for triggering a dual-target conflict are as follows:
[0085] ;
[0086] or:
[0087] ;
[0088] in, The threshold for acceptable narrative rhythm The threshold for successful conversion is defined as follows: When a candidate video meets any of the above conditions, the system determines that there is a conflict between two objectives. The first scenario indicates that the video has a good narrative rhythm but insufficient conversion, which may be due to factors such as late action guidance, insufficient appeal in the first three seconds, insufficient product exposure, or unprominent keywords in the subtitles. The second scenario indicates that the video has a good conversion prediction but insufficient narrative rhythm, which may be due to factors such as abrupt action guidance, excessive emphasis on selling points, premature climax, lack of conflict, or excessively abrupt camera transitions.
[0089] After determining that a conflict exists, the system enters a lens-level minimal modification projection repair process. This repair process uses a set of candidate control parameters. Based on this, solve for the set of control parameters after the repair. The set of control parameters may include:
[0090] ;
[0091] in, The number of shots in the short video; For the first The duration of each shot; For the first The switching interval or transition intensity between individual shots and adjacent shots; For the first Each shot corresponds to the moment when the subtitle appears; For the first Alignment point between the camera and the audio beat; Location of key event shots; This determines the location of the action-guided segments. The variables mentioned above can be expanded based on the specific video editing engine, such as adding subtitle font size, subtitle keyword highlighting time, product close-up duration, and background music volume envelope.
[0092] The projection restoration model can be represented as:
[0093] ;
[0094] And satisfy:
[0095] ;
[0096] in, This is the set of candidate control parameters before repair. This is the set of control parameters after the repair. The weighted L2 distance is used to represent the weighted change in control parameters before and after the repair. This is a weight matrix for different control parameters, used to reflect the cost of modifying different parameters; The constraints on the sequential order and timeline interval of the five types of narrative event anchor points are satisfied.
[0097] The weighted 2-norm distance can be further expressed as:
[0098] ;
[0099] in, This is a positive semi-definite weight matrix used to control the adjustment priority of different parameters. For example, if you want to minimize changes to the camera position of key events, you would increase the weight matrix. Corresponding weight; if moderate adjustment of the subtitle appearance time is allowed, then reduce it. Corresponding weights; if the audio beat points have a significant impact on viewing smoothness, then increase the weight. Corresponding weights.
[0100] In practical solutions, projection repair models can be implemented using constrained optimization, sequential quadratic programming, coordinate descent, heuristic search, Bayesian optimization, or reinforcement learning strategies. For industrial applications, a hierarchical heuristic approach can be prioritized to improve speed: the first layer adjusts the position of the action guidance fragment. and key event shot positions To address insufficient dissemination and conversion; the second layer involves adjusting the duration of shots. and camera switching interval To resolve rhythmic breaks; the third layer adjusts the timing of subtitle appearance. Alignment point with audio beat Improve viewing smoothness. Recalculate after each adjustment. and This continues until both metrics meet the threshold, or the maximum number of iterations is reached.
[0101] For example, in a 15-second product video, the narrative anchors of the candidate solutions are in the correct order, but the action-guided clip is located at the 14th second, resulting in a lower predicted click-through rate. The system detected this. but The projector then proceeded to the projection repair module. This module maintained the original order of cause, conflict, turning point, climax, and action guidance, and preserved the selected product display video clips and the main content of the voiceover. The only changes were moving the action guidance clip from 14 seconds to 11.5 seconds, shortening the product close-up shot corresponding to the climax anchor point from 2 seconds to 1.6 seconds, advancing the subtitle keywords by 0.4 seconds, and aligning the background music beat with the product close-up shot. After the repair, the action guidance remained within the 75% to 100% time range, the narrative structure was not disrupted, but the predicted conversion rate improved, reaching [a certain value]. .
[0102] For example, a short educational video, in an effort to increase clicks, piled on action prompts and conclusions in the first three seconds, but this resulted in a lack of conflict and plot twists, leading to a low narrative pacing score. The system detected this. but Therefore, while keeping the main content of the material unchanged, the question captions and shots corresponding to the conflict anchor points were inserted from the 3rd to the 5th second, the explanatory shots corresponding to the turning point anchor points were placed from the 6th to the 9th second, and the action guidance was reserved after the 13th second. By adjusting the timeline distribution locally, the video narrative became complete, while retaining the conversion guidance effect.
[0103] The loss function can also be expressed as:
[0104] ;
[0105] in, For narrative rhythm loss, used to represent Deviation from the target narrative rhythm range; To represent the propagation transformation loss, it is used to express... With propagation target parameters Deviation between; The narrative event anchor point constraint loss is used to represent the degree to which the five types of narrative event anchor points deviate from the preset time axis interval; To minimize the loss of modification, the parameters used to limit the changes in video generation control parameters before and after projection restoration are used. , , , These are the weight coefficients for the corresponding loss terms. It can be calculated based on the distance between the actual landing point of each anchor point and the target area; The weighted L2 distance described above can be used; It can be calculated based on the degree to which the narrative pacing score falls below the target range; It can be calculated based on the degree to which the predicted propagation conversion value is lower than the propagation target. This is achieved by introducing... and This invention transforms the optimization process from a simple multi-objective weighted process into a local repair mechanism with narrative structure constraints and minimum modification constraints.
[0106] The technical benefits of this step are twofold: firstly, it avoids content drift, material replacement, loss of brand information, and increased generation costs caused by complete regeneration; secondly, it simultaneously satisfies the narrative rhythm threshold and the dissemination conversion threshold through local parameter repair, enabling the generated short videos to achieve a more stable balance between story expression and platform dissemination.
[0107] 5. S5 Video Generation Steps
[0108] S5 is used based on optimal control parameters. For multimodal materials The key to generating short videos lies in executing the shot-level control parameters obtained from S4, rather than simply calling the generation model.
[0109] In one implementation, the video generation module first determines the video based on... and Construct a timeline. For each narrative event anchor, the system selects materials that match its semantics and material tags. For example, the cause anchor can prioritize scene setup shots; the conflict anchor can select user pain point scenes, comparison scenes, or problem captions; the turning point anchor can select product appearance, method explanation, or solution scenes; the climax anchor can select close-ups, result displays, evidence scenes, or emotionally charged shots; and the action guidance anchor can select buttons, QR codes, verbal calls to action, caption prompts, or snippets of promotional information.
[0110] Subsequently, the system followed The video clips are arranged on the timeline based on their shot duration, shot order, and shot position. The subtitle generation module then generates subtitles based on their appearance time. The system generates subtitles based on subtitle density parameters and can highlight key words. The audio module aligns with the audio beat. Align the peak beats of the background music with the timing of key shots or subtitles. Action-guided clips should be based on... Insert it within the specified interval and maintain a preset time interval between it and the climax anchor point to avoid the narrative break caused by the action guidance being too early, and also to avoid the conversion guidance being too late so that users do not see it.
[0111] During the generation process, if a text-to-video or image-to-video model is invoked, the system can convert the semantic fragments corresponding to each anchor point into generated prompts, and add shot style, subject, duration, motion type, and emotion tags. If an existing material editing mode is used, the system completes cropping, splicing, transitions, subtitles, dubbing, and export through the editing engine. After generation, the system can recalculate. and If the threshold is not met, the process can return to step S4 for secondary repair; if the threshold is met, the final image is output.
[0112] 6. Feedback Layered Write-back Implementation Method
[0113] After a short video is published, the system can collect actual dissemination feedback data. . This can include actual impressions, click-through rate, three-second dwell time, effective dwell time, completion rate, like rate, comment rate, share rate, favorite rate, action button click rate, inquiry rate, and purchase rate. The system will... With the generation time , , and Associated storage.
[0114] Unlike simply updating a single global model weight, this invention hierarchically writes feedback data back to four categories of libraries. First, it writes back to the narrative template library to update the recommendation order and time intervals of five types of narrative event anchors across different platforms and themes. For example, if a certain type of e-commerce video performs well with "conflict preceding, climax following, and action guidance in the last 20%," then the corresponding template is updated. Second, it writes back to the shot rhythm template library to update shot duration, switching intervals, the rhythm of the first three seconds, and the rhythm of the climax. Third, it writes back to the conversion trigger segment library to update action guidance scripts, button appearance methods, discount prompt locations, and conversion segment durations. Fourth, it writes back to the material weight library to update the selection priority of different types of materials on different platforms.
[0115] Through this layered write-back mechanism, the system of this invention can continuously improve in subsequent short video generation tasks, forming a closed-loop optimization capability. For example, if actual data shows that users on a certain platform respond better to the structure of "proposing a conflict in the first 3 seconds + providing action guidance in the 10th second," the system will increase the recommendation weight of this structure in the next similar task; if a certain type of background music increases clicks but decreases completion time, the system will reduce its material weight in the corresponding scene.
[0116] IV. Model Structure, Training Methods, Parameter Selection, and Dataset Disclosure
[0117] To enable those skilled in the art to actually implement this invention, the following provides a structured disclosure of the model structure, training method, parameter selection, and dataset usage.
[0118] First, the narrative anchor point analysis model can adopt a structure of "pre-trained language model encoder + sequence labeling layer + anchor point classification layer + time axis regression layer". The input is the narrative text after sentence segmentation, and the output is the probability of five types of anchor points and the recommended time axis interval for each semantic segment. The training data can consist of manually annotated short video scripts, and each sample includes a text segment, anchor point label, segment order, manually recommended time interval, and manually rated rhythm. The training loss can include anchor point classification cross-entropy loss and time axis interval regression loss. After training, the model can be used in step S2.
[0119] Second, the dissemination conversion prediction model can adopt a multi-task learning structure. Input features include text features, video structure features, material features, platform features, and account history features; output includes... , , , and The training labels are derived from the dissemination data of actual published videos. To reduce the influence of different accounts, labels can be normalized to the same account or the same category. The training loss can be a weighted sum of the losses from multiple tasks. For example, binary cross-entropy is used for click tasks, mean squared error or cross-entropy is used for stay and completion tasks, and weighted cross-entropy is used for conversion tasks to handle sample imbalance.
[0120] Third, regarding parameter selection, to It can be obtained by fitting based on artificial rhythm scores, or it can be initialized as... , , , Then adjust based on the validation set; to Based on the target platform type Settings, such as those for e-commerce conversion platforms, can be configured. , , , , The knowledge dissemination platform can be set up , , , , The above values are merely examples, and those skilled in the art can adjust them based on platform data.
[0121] Fourth, regarding threshold selection, and This can be determined based on the quantiles of historical high-quality samples. For example, by sorting historical samples according to manual rhythm scoring, the lowest quantile of the top 30% of samples can be used. As ; Sort the historical samples according to their actual conversion performance, and take the lowest value of the top 30% of samples. As It can also be dynamically determined based on the user-defined dissemination goals.
[0122] Fifth, regarding projection repair parameters, the weight matrix... It can be set as a diagonal matrix, where the diagonal elements correspond to the cost of modifying different control variables. If you want to keep the camera positions of key events as constant as possible, then increase... Corresponding weight; if minor adjustments to the subtitles are allowed, then reduce the weight. Corresponding weights; if the audio beat significantly impacts the viewing experience, then increase the weight. Corresponding weights. The maximum number of iterations can be set from 10 to 50. The step size for each adjustment can be set proportionally to the total video duration. For example, the single adjustment of shot duration should not exceed 5% of the total duration, and the single adjustment of subtitle timing should not exceed 0.5 seconds.
[0123] V. Specific Application Examples and Experimental Data
[0124] To verify the technical effectiveness of the AI short video generation method and system of the present invention, which achieves collaborative optimization of narrative rhythm and dissemination conversion, the applicant built an AI short video generation verification platform and conducted experiments on three typical short video tasks: e-commerce product promotion, knowledge course promotion, and brand service introduction. The experiments focused on verifying whether the present invention, through five types of narrative event anchor point constraints, narrative rhythm-dissemination conversion dual-objective conflict detection, shot-level minimal modification projection repair, and feedback layered rewrite mechanism, can improve the narrative integrity, viewing smoothness, completion rate, click-through rate, and action conversion rate of short videos without destroying the original narrative structure and main content.
[0125] 1. Experimental Environment and Verification Process
[0126] This experiment uses Figure 1 The system architecture shown was validated. The system includes an input module, a narrative anchor point analysis module, a propagation transformation evaluation module, a dual-objective conflict detection module, a projection repair module, a video generation module, and a feedback write-back module. The experimental procedure is as follows: Figure 2 As shown: First, input the natural language narrative text. Multimodal materials Target platform type and propagation target parameters ; then according to Figure 3 The method shown divides the text into five categories of narrative event anchors: cause, conflict, turning point, climax, and action guidance, and maps them to the short video timeline; then, a narrative rhythm evaluation index is calculated. and dissemination transformation evaluation indicators When a conflict is detected between the two, proceed according to... Figure 4 The process shown performs minimal lens-level projection repair; finally, a short video is generated and actual dissemination feedback data is collected in the test environment. The feedback data is then written back to the narrative template library, the lens rhythm template library, the transformation trigger segment library, and the material weight library.
[0127] The experimental server was configured with a 16-core CPU, a 24GB GPU, 128GB of RAM, and a Linux operating system. The text encoding model used a pre-trained Chinese language model for fine-tuning, the propagation and conversion prediction model employed a multi-task learning structure, and the video generation was completed using a combination of automatic material editing, subtitle generation, audio beat alignment, and template synthesis. In the experiment, each short video was kept between 15 and 30 seconds in total length, with a resolution of 1080×1920 and a frame rate of 30fps.
[0128] The test data used in this experiment includes three types of tasks: short videos of e-commerce products, short videos of knowledge courses, and short videos of brand services. For each type of task, 30 sets of input text and corresponding materials were prepared, totaling 90 test tasks. Each task includes a natural language narrative text, 5 to 12 images or video clips, 1 to 3 background audio clips, a set of subtitle templates, and one action guidance material. Short videos for each task were generated using both the method of this invention and the comparative method, and were then evaluated automatically and through small-scale user testing.
[0129] 2. Evaluation Indicators
[0130] To objectively evaluate the technical effects of this invention, the following indicators were used in the experiment.
[0131] 2.1 Narrative Rhythm Evaluation Indicators This metric is used to measure the narrative completeness, anchor point placement rationality, emotional progression, and shot fluidity of short videos. Its formula is:
[0132] ;
[0133] in, As a semantic density index, As an indicator of emotional fluctuations, Matching metrics to key event locations For rhythm break penalty terms, , , , These are non-negative weighting coefficients. In this experiment, , , , .
[0134] 2.2 Evaluation Indicators for Dissemination and Transformation This is used to measure the predictive performance of short videos in terms of dissemination and conversion on a target platform. Its calculation formula is:
[0135] ;
[0136] in, To predict click-through rate, To predict the effective stay rate, To predict the completion rate, To predict the sharing rate, To predict action conversion rates, to These are weighting coefficients set based on the target platform type and communication objectives. In this experiment, e-commerce tasks place greater emphasis on clicks and action conversions, and therefore... , , , , Knowledge-based courses place greater emphasis on effective engagement and completion rates, and are designed with... , , , , Brand service tasks are allocated in a balanced manner.
[0137] 2.3. The measured click-through rate (CTR) represents the percentage of test users who click on video details, links, or action buttons.
[0138] 2.4. Actual completion rate, which represents the proportion of users who watch the entire short video.
[0139] 2.5 Actual action conversion rate, which represents the proportion of actions that convert into behaviors such as clicking to receive, inquiring, purchasing, making an appointment, or following.
[0140] 2.6 Parameter Change Rate, used to measure the degree of change in control parameters before and after projection restoration. Its calculation formula is:
[0141] ;
[0142] in, This is the set of candidate control parameters before repair. This is the set of control parameters after the repair. To control the parameter weight matrix, The smaller the value, the smaller the repair / modification.
[0143] 3. Scale settings
[0144] To demonstrate the technical effects of the present invention, the following comparative examples are provided.
[0145] Comparative Example 1 illustrates a standard AI short video generation method. This method directly generates short videos based on input text and materials, without classifying them into five categories of narrative event anchor points or calculating... and It also does not perform dual-target conflict detection and projection repair.
[0146] Comparative Example 2 is a narrative rhythm optimization method only. This method analyzes the narrative structure of the text and tries to ensure the complete sequence of cause, conflict, turning point, climax, and action guidance, but does not introduce dissemination conversion evaluation indicators. It also does not optimize based on the data disseminated on the target platform.
[0147] Comparative Example 3 is a method that only optimizes the conversion rate. This method optimizes the video structure based on dissemination metrics such as click-through rate, completion rate, and action conversion rate, but it does not impose timeline constraints on the five types of narrative event anchors, which can easily lead to premature action guidance, excessive stacking of selling points, or narrative breaks.
[0148] Comparative Example 4 illustrates a conventional bi-objective weighted optimization method. This method simultaneously calculates... and The system uses a standard weighted loss function for optimization. However, when the two conflict, it does not perform the minimum shot-level projection repair that "keeps the anchor point order and the main content of the material unchanged". Instead, it handles the issue by regenerating or adjusting the weights as a whole.
[0149] Examples 1 to 3 all employ the method of this invention, namely: first, five types of narrative event anchor points are divided and timeline mapping is performed, and then propagation and transformation prediction is conducted. and When conflicts exist, keep the anchor point order and the main content of the material unchanged, only adjust the shot duration, the time when the subtitles appear, the audio beat alignment, the position of the key event shots and the position of the action guidance clips, and perform minimal changes to the projection repair.
[0150] 4. Application Example 1: Generation of Short Videos for E-commerce Beverages
[0151] 4.1 Input Data
[0152] This example demonstrates how to generate a 15-second short video for e-commerce beverages. Input text "Feeling sleepy in the afternoon but worried about coffee being too strong? This decaffeinated herbal coffee has a mild taste and refreshing aroma, perfect for afternoon tea in the office. Click now to claim your sample."
[0153] Input materials Includes: 2 office scene videos, 3 beverage packaging images, 1 video of coffee being poured into a cup, 1 close-up video of a beverage, 2 background music tracks, 1 set of subtitle templates, and 1 action guide button. Target platform type. For e-commerce short video platforms, target parameters for dissemination Prioritize clicks and conversions.
[0154] 4.2 Narrative Anchor Point Analysis
[0155] The system categorizes the text into five types of narrative event anchors, as shown in Table 1 below:
[0156]
[0157] In the initial candidate videos, the action guidance clip was located between 14.3 and 15 seconds, resulting in a lower predicted click-through rate. The system calculated... Reaching the acceptable threshold for narrative pacing ;but Below the acceptable threshold for propagation and conversion Therefore, the system determined that there was a dual-objective conflict: the narrative pacing met the requirements, but the dissemination and conversion were insufficient.
[0158] 4.3 Projection Restoration Process
[0159] The system has entered a minimal-modification, lens-level projection repair phase. During the repair, the order of the five narrative event anchor points remains unchanged, as do the main elements of the beverage packaging, coffee being poured into the cup, the office scene, and the background music. Only the following parameters are adjusted: the position of the action guidance segment... The duration of the climax scene was shortened from 14.3 seconds to 12.1 seconds; the duration of the climax scene was extended. Compressed from 3.1 seconds to 2.6 seconds; the timing of the decaffeinated herbal coffee subtitle appearance was also shortened. 0.3 seconds in advance; align the product close-up with the peak of the background music beat. Alignment; Position the beverage packaging in a close-up shot. The time has been adjusted from 9.8 seconds to 8.9 seconds.
[0160] After the repair, the order of narrative anchors remains unchanged, action guidance is still located within the last 25% of the video, and the main content of the source material has not been replaced, thus meeting the requirements. .
[0161] 4.4 The experimental results are shown in Table 2 below.
[0162]
[0163] As shown in Table 2, compared with Comparative Example 4, the present invention achieves [the desired effect] with a parameter modification rate of only 7.9%. The measured click-through rate increased from 5.1% to 6.4%, and the measured action conversion rate increased from 2.5% to 3.3%, while maintaining a high narrative pacing score. These results demonstrate that the lens-level minimal-modification projection repair of this invention can improve dissemination and conversion effects without significantly altering the video structure.
[0164] 5. Application Example 2: Generation of Short Videos for Knowledge Courses
[0165] 5.1 Input Data
[0166] This example demonstrates how to generate a 30-second promotional video for a knowledge course. Input text The course explains: "Many people learn Excel by only memorizing formulas, but don't know how to connect data cleaning, pivot analysis, and automatic updates. This course uses three real-world business cases to guide you from messy spreadsheets to automated reports. It's suitable for office workers with no prior experience. Click to view a sample lesson."
[0167] Target platform type As a knowledge-sharing platform, it disseminates target parameters. Prioritize effective viewing and completion of the presentation. Materials include teacher-led presentation clips, screen recordings of table operations, course cover images, case study screenshots, subtitle templates, and background music.
[0168] 5.2 Initial Conflict Situation
[0169] The system identifies five types of anchor points as follows: The cause is "Many people only memorize formulas when learning Excel"; the conflict is "But they don't know how to connect data cleaning, pivot analysis, and automatic updates"; the turning point is "This course uses three real business cases"; the climax is "From messy tables to automated reports"; and the action guide is "Click to view the trial lesson".
[0170] In the initial candidate videos, to increase click-through rates, the system-generated action guide subtitles appeared too early, around the 5-second mark. This resulted in a conversion prompt appearing before the course content was fully developed, leading to a low narrative pacing score. (Calculations were made...) higher than ;but lower than The system determined that there was a conflict between the two objectives: "the communication conversion met the target, but the narrative pace was insufficient."
[0171] 5.3 Projection Restoration Process
[0172] The projection repair module keeps the teacher's narration, the table recording, and the course cover material unchanged, moves the action guidance segment from the 5th second to the 25th to 30th second interval; extends the subtitles corresponding to the conflict anchor points by 0.8 seconds; advances the course case screen corresponding to the turning point anchor point to the 12th second; places the climax anchor point "automated report generation effect" between the 20th and 24th seconds; and aligns the peak of the background music beat with the screen where the report generation is complete.
[0173] After the fix, the action guidance is in the recommended range, the sequence of conflict, turning point, and climax is complete, and both the predicted completion value and the actual completion performance have improved.
[0174] 5.4 The experimental results are shown in Table 3 below.
[0175]
[0176] As shown in Table 3, compared with Comparative Example 4, this invention, with a parameter modification rate of less than 10%, increased the measured completion rate from 44.2% to 51.8% and the effective dwell time rate from 57.1% to 63.5%. This result demonstrates that this invention does not simply guide viewers to clicks in advance, but rather restores the narrative structure while maintaining the dissemination goals, thus improving the viewing integrity of knowledge-based videos.
[0177] 6. Application Example 3: Generation of Short Videos Introducing Brand Services
[0178] 6.1 Input Data
[0179] This example demonstrates how to generate a 20-second short video introducing a brand's services. Input text "Businesses often struggle with topic selection, editing, and post-production analysis when creating short videos. We offer a one-stop service from script planning and footage shooting to data analysis, helping brands consistently produce stable content. Feel free to schedule a consultation."
[0180] Target platform type As a brand showcase platform, disseminating target parameters To achieve a balance between brand trust and consultation conversion, the materials include videos of office scenes, team work footage, screenshots of client case studies, data dashboards, the brand logo, and a consultation appointment button.
[0181] 6.2 Initial Conflict Situation
[0182] Typical AI-generated content tends to piece together materials in sequence, resulting in the brand logo and appointment / consultation button appearing at the end. However, the first half lacks conflict and transition, leading to high viewer drop-off. Conversion optimization methods, on the other hand, move the appointment / consultation button forward to the 4th second, causing users to see the consultation prompt before they understand the service's value, resulting in a clear narrative break.
[0183] The method of this invention divides the text into: Cause: "Enterprises make short videos"; Conflict: "Often don't know how to choose topics, edit, and review"; Turning point: "We provide a one-stop service from script planning and material shooting to data review"; Climax: "Helping brands continuously produce stable content"; Action guidance: "Welcome to make an appointment for consultation".
[0184] After the system initially generates candidate solutions, it detects... Approaching the threshold, The value is below the threshold, thus triggering projection repair.
[0185] 6.3 Projection Restoration Process
[0186] The system retains the brand logo, team work scenes, and client case screenshots, only moving the data dashboard display from 15 seconds to 12 seconds to enhance the climax; moving the appointment / consultation button from 19 seconds to 16.8 seconds; aligning the "Stable Content Production" text with the client case screenshots; and adjusting the camera transition interval from an average of 1.9 seconds to 1.6 seconds to improve the smoothness of the rhythm.
[0187] 6.4 The experimental results are shown in Table 4 below.
[0188]
[0189] As shown in Table 4, this invention can also improve narrative pace and conversion performance in brand service videos while maintaining a low parameter modification rate. Compared with ordinary AI generation, the measured completion rate increased from 36.4% to 48.9%, and the consultation click-through rate increased from 1.7% to 3.1%.
[0190] 7. Overall Experimental Results
[0191] The results of the 90 test tasks are summarized in Table 5 below:
[0192]
[0193] As can be seen from the overall results in Table 5, compared with ordinary AI generation methods, the present invention, on average... The average value increased from 0.61 to 0.80. The average completion rate increased from 0.58 to 0.75, the average click-through rate increased from 36.9% to 52.5%, the average click-through rate increased from 3.4% to 5.6%, and the average conversion rate increased from 1.5% to 3.1%. Compared with the conventional dual-objective weighted method, the average parameter change rate of this invention decreased from 19.2% to 8.6%, while the average completion rate, click-through rate, and conversion rate were all improved.
[0194] The foregoing description of embodiments of the present invention, through which those skilled in the art are able to implement or use the present invention, will be readily apparent to those skilled in the art. Various modifications to these embodiments will be readily apparent to those skilled in the art. The general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novelty disclosed herein.
Claims
1. An AI short video generation method that collaboratively optimizes the dual objectives of narrative rhythm and dissemination transformation, characterized in that, Includes the following steps: S1: Receives natural language narrative text, multimodal materials, target platform type, and propagation target parameters; S2: Perform semantic analysis on the narrative text, divide it into five categories of narrative event anchors: cause, conflict, turning point, climax and action guidance, and generate narrative structure vector and narrative rhythm evaluation index; S3: Generate dissemination conversion evaluation indicators based on historical dissemination data, user interaction data, and target platform type; S4: Construct a dual-objective loss function based on narrative rhythm evaluation index and dissemination conversion evaluation index, and trigger shot-level minimum modification projection repair when a conflict between narrative rhythm objective and dissemination conversion objective is detected to obtain the optimal control parameters; S5: Based on the optimal control parameters, control the shot duration, subtitle density, audio beat alignment, key event shot position, and action guidance segment position, and fuse and edit the multimodal materials to generate a short video.
2. The method according to claim 1, characterized in that, In step S2, the five types of narrative event anchors are mapped to different relative intervals on the short video timeline. The cause anchor is located in the first 0% to 20% of the total video duration, the conflict anchor is located in the 15% to 45% interval, the turning point anchor is located in the 35% to 65% interval, the climax anchor is located in the 55% to 85% interval, and the action guidance anchor is located in the 75% to 100% interval. The narrative rhythm evaluation index is obtained by weighting the semantic density index, emotional fluctuation index, key event position matching index, and rhythm break penalty term.
3. The method according to claim 1, characterized in that, In step S3, the dissemination conversion evaluation index is not determined solely by the click-through rate, but is a composite index based on clicks, effective dwell time, completion rate, sharing, and action conversion. The dissemination conversion evaluation index is obtained by weighting the predicted click-through rate, predicted effective dwell rate, predicted completion rate, predicted sharing rate, and predicted action conversion rate; each weight coefficient is a non-negative weight coefficient determined according to the target platform type, and the sum of all weight coefficients is 1.
4. The method according to claim 1, characterized in that, In step S4, the dual-objective loss function includes narrative rhythm loss, propagation conversion loss, narrative event anchor point constraint loss, and minimum modification loss; each of the above loss terms has a corresponding weight coefficient; the conflict between the narrative rhythm objective and the propagation conversion objective is triggered according to the following conditions: when the candidate video scheme satisfies that the narrative rhythm evaluation index is not lower than the narrative rhythm qualified threshold and the propagation conversion evaluation index is lower than the propagation conversion qualified threshold, or satisfies that the propagation conversion evaluation index is not lower than the propagation conversion qualified threshold and the narrative rhythm evaluation index is lower than the narrative rhythm qualified threshold, it is determined that there is a dual-objective conflict.
5. The method according to claim 4, characterized in that, The shot-level minimal modification projection repair process includes: maintaining the order of narrative event anchor points and the main content of the selected multimodal materials unchanged, and only adjusting the shot duration, shot switching interval, subtitle appearance time, audio beat alignment point, key event shot position, and action guidance segment position; the shot-level minimal modification projection repair process uses the candidate control parameter set before repair as a benchmark to solve the control parameter set after repair, so that the weighted difference between the control parameter set after repair and the candidate control parameter set before repair is minimized, and at the same time, the following constraints are satisfied: the narrative rhythm evaluation index after repair is not lower than the qualified threshold of narrative rhythm, the dissemination conversion evaluation index after repair is not lower than the qualified threshold of dissemination conversion, and the constraints of the order and time interval of the five types of narrative event anchor points are all satisfied.
6. The method according to claim 1, characterized in that, In step S5, the optimal control parameters include at least shot-level control parameters, subtitle-level control parameters, audio-level control parameters, and transition trigger control parameters. The shot-level control parameters include the duration of each shot, the scene switching method, and the shot order. The subtitle-level control parameters include the subtitle appearance time, subtitle density, and keyword highlighting position. The audio-level control parameters include the background music beat points, sound effect insertion points, and voice pause points. The transition trigger control parameters include the insertion position of the action guidance segment, its duration, and the time interval between it and the climax anchor point.
7. The method according to claim 1, characterized in that, After the short video is generated and published, actual dissemination feedback data is collected and written back to the narrative template library, the shot rhythm template library, the conversion trigger segment library, and the material weight library, respectively. The narrative template library is used to update the recommendation order and time interval of five types of narrative event anchors; the shot rhythm template library is used to update the shot duration and switching interval; the conversion trigger segment library is used to update the insertion strategy of action guidance segments; and the material weight library is used to update the selection priority of different types of materials on different platforms.
8. An AI short video generation system for co-optimizing narrative rhythm and dissemination transformation, the system being used to implement the method described in any one of claims 1 to 7, characterized in that, include: The input module is used to receive natural language narrative text, multimodal materials, target platform type, and dissemination target parameters; the narrative anchor point analysis module is used to divide the narrative text into five categories of narrative event anchor points: cause, conflict, turning point, climax, and action guidance, and generate narrative structure vectors and narrative rhythm evaluation indicators. The dissemination conversion evaluation module is used to generate dissemination conversion evaluation indicators based on historical dissemination data, user interaction data, and target platform type; the dual-objective conflict detection module is used to determine whether the narrative rhythm objective and the dissemination conversion objective conflict. The projection repair module is used to repair the projection with minimal changes to the lens-level generation control parameters when a conflict is detected; the video generation module is used to fuse and edit multimodal materials based on the repaired optimal control parameters to generate a short video. The feedback write-back module is used to write back the actual propagation feedback data to the narrative template library, the shot rhythm template library, the conversion trigger segment library, and the material weight library, respectively. The projection repair module includes a parameter freezing unit, a local adjustment unit, and a constraint verification unit. The parameter freezing unit is used to freeze the order of narrative event anchor points and the main content of selected multimodal materials. The local adjustment unit is used to adjust shot duration, shot switching interval, subtitle appearance time, audio beat alignment point, key event shot position, and action guidance segment position. The constraint verification unit is used to determine whether the repaired control parameters simultaneously meet the narrative rhythm threshold, propagation conversion threshold, and narrative event anchor point constraints.
9. A computer device, characterized in that, The computer device includes a processor, a graphics processing unit (GPU), and a memory, wherein the memory stores a computer program that, when executed by the processor and the GPU, causes the computer device to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, the computer program, when executed by a computer device, causing the computer device to perform the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Systems and methods for automating video editing
US11769528B2
System and method for evaluating and optimizing media content
US20210049627A1
Text-driven ai-assisted short-form video creation in an ecommerce environment
US20240185306A1