Book short video multi-stage generation method based on hotspot fusion

By employing a multi-stage generation method, the problems of low production efficiency and content homogenization in short videos have been solved, enabling efficient and accurate integration of trending information and content generation, thereby improving the quality and dissemination effect of short videos.

CN122489799APending Publication Date: 2026-07-31CITIC UNITED CLOUD TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CITIC UNITED CLOUD TECH CO LTD
Filing Date
2026-05-11
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing short video production methods rely on manual planning and editing, which is inefficient and cannot meet the needs of large-scale production. The generated results are highly homogenized, lack a deep semantic understanding of the original content, and fail to effectively integrate real-time trending information, resulting in content that lacks timeliness and dissemination potential.

Method used

The book short video generation method based on hot topic integration is a multi-stage process that includes content analysis and semantic modeling, content quality assessment, hot topic information modeling, semantic matching and content integration, marketing copy generation, script structure design, shot modeling and multimodal content generation. This method achieves accurate matching and efficient generation of content and hot topics.

Benefits of technology

It significantly improves the scientific nature and stability of content selection, increases generation efficiency, enhances the accuracy of matching content with trending topics and the ability to control the timing of dissemination, and forms high-quality, diversified short video content with good scalability and commercial value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122489799A_ABST
    Figure CN122489799A_ABST
Patent Text Reader

Abstract

This multi-stage method for generating short books based on trending topic fusion takes the original book text as input and constructs a unified feature representation through semantic decomposition and vectorized encoding. It then performs multi-dimensional quantitative scoring of the book content, selecting high-quality segments with narrative tension and dissemination potential. Trending topic data is collected from multiple platforms, and a comprehensive weighting of trending topics is calculated using cross-platform signal normalization and time decay mechanisms. A dual-channel fusion strategy based on semantic similarity and keyword overlap achieves precise matching between trending topics and book segments. On this basis, a joint scoring model of content quality and trending topic relevance is constructed to optimize candidate content and automatically generate marketing copy. Finally, the method designs the script structure, divides the scene, and allocates the duration, generating a shot sequence. Combined with a multimodal model, it completes the automatic generation of short videos. This method automates the entire process from content understanding to video generation, improving the generation efficiency and dissemination effect of marketing short videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and more specifically to a multi-stage method for generating short videos of books based on hotspot fusion. Background Technology

[0002] With the rapid development of short video platforms, the demand for content production has exploded, especially in scenarios such as marketing promotion, knowledge dissemination, and content monetization, where the demand for high-quality short videos continues to rise. Current short video production methods mainly rely on manual planning and editing, which are complex, inefficient, and unable to meet the needs of large-scale production. Meanwhile, some automated tools based on templates or simple generation models often lack a deep semantic understanding of the original content, resulting in highly homogenized outputs that fail to showcase content differentiation. Furthermore, existing methods generally fail to effectively integrate real-time trending information, leading to a lack of timeliness and dissemination potential in the generated content. In terms of multimodal generation, modules such as images, videos, and audio are often independent, lacking unified semantic constraints and collaborative mechanisms, easily resulting in inconsistent content or fragmented expression. Therefore, how to construct an automated short video generation method that combines content understanding, trending information perception, and multimodal collaborative generation to improve generation efficiency, content quality, and dissemination effects has become a key issue that urgently needs to be addressed in the current technological field.

[0003] To address the above issues, this invention provides a multi-stage method for generating short book videos based on hotspot fusion. Summary of the Invention

[0004] To address the shortcomings of existing technologies, a multi-stage generation method for book short videos based on hotspot fusion is proposed, characterized by:

[0005] Step S1, Content Parsing and Semantic Modeling: The input content is segmented and divided into semantic units. Key events, relationships between characters, and topic information are extracted through structural analysis, and context-related semantic representations are constructed. At the same time, a cross-segment association mechanism is introduced to model the contextual dependencies in long texts, ensuring that content units maintain overall semantic consistency while being locally independent.

[0006] Step S2, Content Quality Assessment and Candidate Screening: Based on a multi-dimensional feature extraction mechanism, the degree of plot conflict, intensity of emotional fluctuation, information density, and visual expression potential indicators are calculated for each content unit, and a comprehensive score is generated through a weighted model. Combined with redundancy removal constraints, high-scoring content is screened to form a diverse set of candidate content with dissemination potential.

[0007] Step S3: Hotspot information modeling and dynamic updating. Obtain multi-source hotspot data, represent hotspot information from different sources in a unified manner, and calculate the corresponding popularity weights. Dynamically update hotspots through a time decay mechanism, and combine cross-platform data for fusion correction to obtain a hotspot representation that reflects the current dissemination trend.

[0008] Step S4: Hotspot semantic matching and content fusion. Match the semantic representation of candidate content with the semantic representation of hotspots. Establish the mapping relationship between content and hotspots through the joint calculation of semantic similarity and keyword relevance. Weight the content according to the matching strength to generate an enhanced content representation that integrates hotspot information.

[0009] Step S5: Candidate content selection and ranking. Based on the content quality score and hot topic matching results, the candidate content is comprehensively ranked, and the weight relationship between content quality and hot topic popularity is balanced through a normalization strategy, so as to select the target content most suitable for short video generation.

[0010] Step S6: Marketing copy generation and constraint optimization. Based on the fused content representation, short video marketing copy is generated through a conditional generation model. A multi-objective constraint mechanism is introduced during the generation process to jointly optimize the copy structure, semantic expression, and user appeal. Specifically, structural constraints control the copy to include an introductory opening, core conflict, and result-oriented content. Appeal constraints enhance the information density and emotional intensity of the opening. The copy content is also dynamically adjusted based on hot topic relevance, thereby generating short video copy that combines information expression and dissemination appeal.

[0011] Step S7: Script structure generation and rhythm planning. Based on the generated text content, construct the short video script structure, divide the video into scenes and design the narrative order, and model the importance of each scene based on the semantic strength of the content, the criticality of the plot, and the relevance of the hot topics. On this basis, determine the duration of each scene through a normalized allocation mechanism, and introduce a position weight adjustment strategy to make the opening and climax scenes occupy higher weights in the overall structure, thereby realizing the structured planning and dynamic optimization of the video rhythm.

[0012] Step S8, Shot Modeling and Fine-Grained Design, further breaks down the scene-level structure into shot sequences, generates visual descriptions, action information and audio text for each shot, and constructs shot-level feature representations; calculates shot importance based on visual complexity, action intensity and emotional expression factors, and allocates shot duration through normalization, thereby achieving fine-grained video rhythm control;

[0013] Step S9: Multimodal content generation and consistency constraints. Generate corresponding images, video clips and audio content based on the shot description, and align different modalities through a unified semantic space. Introduce cross-modal consistency constraints during the generation process to optimize the semantic deviation between visual and textual content, so as to ensure the consistency and coherence of video content in multimodal expression.

[0014] Step S10, Video Synthesis and Dissemination Effect Optimization: The multimodal content corresponding to each shot is spliced ​​sequentially, and a complete short video is generated through transition control, audio-visual synchronization, and rhythm scheduling. During the synthesis process, the transition method is adaptively selected based on the semantic continuity of adjacent shots, and the overall video rhythm is corrected for consistency. Furthermore, a dissemination effect evaluation mechanism is constructed to comprehensively analyze the theme attractiveness, emotional expression intensity, hot topic relevance, and rhythm characteristics to predict the video's dissemination potential. Based on the evaluation results, the video is screened or optimized at the parameter level, thereby outputting short video content with high dissemination effect.

[0015] The present invention also discloses a non-volatile storage medium, characterized in that the non-volatile storage medium includes a stored program, wherein the program, when running, controls the device where the non-volatile storage medium is located to execute the above-described method.

[0016] The present invention also discloses an electronic device, characterized in that it comprises a processor and a memory; the memory stores computer-readable instructions, and the processor is used to execute the computer-readable instructions, wherein the computer-readable instructions execute the method described above.

[0017] Beneficial effects

[0018] This invention, through semantic decomposition and multi-dimensional quantitative scoring of book content, can screen high-quality content with both narrative appeal and dissemination potential from the source, significantly improving the scientific rigor and stability of content selection. By introducing cross-platform hot topic signal fusion and time decay mechanisms, it achieves unified modeling of the "popularity and timeliness" of hot topics, effectively improving the matching accuracy between content and hot topics and the ability to control dissemination timing. Through a dual-channel matching strategy of semantic similarity and keyword overlap, it enhances the ability to capture implicit semantic associations, enabling book content to establish more precise connections with diverse hot topics. Through a joint optimization mechanism of content quality scoring and hot topic association strength, it achieves a synergistic improvement in "content quality" and "dissemination potential." Through a multi-stage generation process, including copywriting generation, script design, storyboard construction, and multimodal synthesis, it forms a complete video generation chain, significantly improving the structural integrity and visual expressiveness of the generated video. Simultaneously, by automating processes to replace manual operations, this invention can significantly reduce content production costs, improve generation efficiency, and possesses good scalability and commercial value. Attached Figure Description

[0019] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0020] Example 1

[0021] A multi-stage method for generating short videos of books based on hot topic fusion includes the following steps:

[0022] Step S1: Automatic Decomposition and Semantic Encoding of Book Content: Receive the original text B of the book as input, decompose it into several semantically complete story fragments, and convert each fragment into a fixed-dimensional vector representation through a pre-trained semantic encoding model, providing a unified feature space for subsequent content scoring and hotspot matching;

[0023] This step includes the following two sub-steps:

[0024] Step S11: Story Fragment Division. Using a sliding window segmentation strategy, the book text B is divided according to semantic boundary paragraphs and chapters to obtain n semantically complete story fragments:

[0025]

[0026] in:

[0027] Book: A collection of fragments of book text, containing all n story fragments;

[0028] C i The i-th story segment (i = 1, 2, …, n) is divided into segments with the paragraph boundaries as the smallest unit.

[0029] n: Total number of segments, dynamically determined by the book length and window size;

[0030] Step S12, Segment semantic vectorization: For each segment C i A pre-trained semantic encoding model is used to generate semantic vector representations, which serve as features for subsequent similarity calculations.

[0031]

[0032] in:

[0033] Fragment C i The semantic vector, with dimension d, is typically 768 or 1024 and is output by the Encoder model;

[0034] Encoder(·): A pre-trained semantic encoding model that maps text of arbitrary length to a fixed-dimensional vector space. It can use either Sentence-BERT or BERT models.

[0035] d: Semantic vector dimension, which depends on the selected encoding model;

[0036] This step also extracts each fragment C. i Keyword set K(C) i ), using TF-IDF or TextRank algorithms:

[0037] K(C i = TopK_TF-IDF(C i )

[0038] in:

[0039] K(C i Fragment C i The keyword set is composed of the top K_w words with the highest weights selected by the TF-IDF algorithm.

[0040] This step also counts the number of sentences in each segment |sent(C i | and word count | tokens(C i )|,|sent(C i )| is fragment C i The number of sentences is calculated using periods, question marks, and exclamation marks as clause boundaries; |tokens(C i )| is fragment C i The word count is calculated after word segmentation;

[0041] Step S2: Multi-dimensional content scoring of story fragments, assigning C to each fragment. i Multi-dimensional quantitative scoring was conducted to select high-quality segments that possess "narrative tension, emotional resonance, and visualization potential."

[0042] Step S21: Calculate features of each dimension, and calculate segment C sequentially. i Five dimensions of characteristics:

[0043] The plot conflict level, F_conflict, reflects the narrative tension of a segment. A conflict word dictionary, C_dict, is defined, including transition words, contrasting words, and conflict verbs. A size of 500-1000 words is recommended. The average number of times conflict words appear per sentence in the segment is calculated.

[0044] F_conflict(C i ) = |{w j ∈ C i : wj ∈ C_dict}| / |sent(C i )|

[0045] in:

[0046] C_dict: A conflict term dictionary, pre-annotated by domain experts, covering transitional and conflict expressions such as "but," "however," "confrontation," and "betrayal."

[0047] {w j ∈ C i : w j ∈ C_dict}: Fragment C i The word w in the conflict word dictionary j A set, |·| represents the number of its elements;

[0048] The emotion intensity F_emotion reflects the degree of emotional fluctuation in a segment. For each sentence in the segment, a pre-trained sentiment analysis model is used to output an emotion polarity score e. s ∈ [-1, +1], take the mean of absolute values:

[0049] F_emotion(C i ) = (1 / |sent(C i )|) · Σs |e s |

[0050] in:

[0051] e s The sentiment polarity score of the s-th sentence is output by the sentiment analysis model, with a value range of [-1, +1], where negative values ​​represent negative sentiment and positive values ​​represent positive sentiment.

[0052] |e s |:e s The absolute value of the value is used to include both positive and negative emotions in the statistics, with a high absolute value representing strong emotions.

[0053] Σs: for fragment C i All in the middle | sent(C i Summing the sum of the sentences;

[0054] Character complexity F_character reflects the richness of relationships between characters in a segment. It is calculated by extracting the set of named entities E(C) from the segment using named entity recognition. i ), calculate the average number of personal name entities per sentence:

[0055] F_character(C i ) = |E(C i )| / |sent(C i)|

[0056] in:

[0057] E(C i Fragment C i The set of personal names identified by the NER model, where |·| represents the number of elements;

[0058] Information density F_density measures the degree of information condensation in a fragment, using the keyword set K(C i ), calculate the percentage of keywords in the total number of words in the segment:

[0059] F_density(C i ) = |K(C i )| / |tokens(C i )|

[0060] in:

[0061] |K(C i |: The set of keywords K(C) extracted by TF-IDF in step S1 i The number of words in ( );

[0062] |tokens(C i )|:Fragment C i Total word count;

[0063] Visualization potential F_visual measures the ability of a fragment to generate visuals. It defines an action word dictionary A_dict (e.g., "running, fighting, exploding") and a scene word dictionary S_dict (e.g., "forest, castle, street"), and calculates the combined percentage of these two types of words in the fragment.

[0064] F_visual(C i ) = |{w j ∈ C i : w j ∈ A_dict ∪ S_dict}| / |tokens(C i )|

[0065] in:

[0066] A_dict: A dictionary of action verbs, containing verbs that express specific actions; a size of 300-500 words is recommended.

[0067] S_dict: A scene-specific dictionary containing nouns representing visual scenes; a size of 300-500 words is recommended.

[0068] A_dict ∪ S_dict: The union of the dictionaries of action words and scene words, with no duplicate entries for the two categories.

[0069] Step S22: Multidimensional weighted scoring. The feature values ​​of the five dimensions are weighted and summed to obtain segment C. i Overall content score (C) i ):

[0070] Score(C i ) = α · F_conflict + β · F_emotion + γ · F_character + δ · F_density + η · F_visual

[0071] in:

[0072] α: Weighting coefficient for plot conflict, default value 0.25. Conflict is the core element that attracts the audience, so it has the highest weight.

[0073] β: Weighting coefficient for emotional intensity, default value 0.25. Emotional resonance directly affects the dissemination rate.

[0074] γ: Weighting coefficient for character complexity, default value 0.20, increases the memorability of character clues;

[0075] δ: The weighting coefficient for information density, with a default value of 0.15. Appropriate density ensures the integrity of the content.

[0076] η: The weighting coefficient for visualization potential, with a default value of 0.15, which determines the quality of the generated image;

[0077] The five weighting coefficients satisfy normalization constraints to ensure that the score range is consistent with the magnitude of each dimension:

[0078] α + β + γ + δ + η = 1, and α, β, γ, δ, η > 0

[0079] Step S3: Hotspot data collection, weight calculation, and time-sensitive fusion. Hotspot data is collected in real time from multiple platforms, and cross-platform signal normalization and weighting, as well as time attenuation correction are performed sequentially. The final comprehensive weight W_final(H) of each hotspot is output. j This provides a quantitative basis for hotspot matching;

[0080] Step S31: Cross-platform hotspot signal normalization and weighting; real-time collection of m hotspot data points to form a hotspot set:

[0081] Hotspot = {H1, H2, …, Hm}

[0082] in:

[0083] H j : The j-th hot topic (j = 1, 2, …, m), each hot topic includes text content, source platform, and publication time t. j property;

[0084] m: Total number of hotspots collected in this batch;

[0085] hotspot H j The original weights of each platform are calculated from the three types of signals. To eliminate dimensional differences, each signal must first be normalized to [0, 1] using Min-Max.

[0086] = [S_search(H j ) - S_search_min] / [S_search_max - S_search_min]

[0087] , Normalize in the same way, then perform weighted fusion after normalization to obtain hotspot H. j The original weights of the single platform W_raw(H) j ):

[0088] W_raw(H j ) = λ1 · + λ2 · S̃_social + λ3 ·

[0089] Where S_search(H j ) is the hotspot H j Search volume on search engine platforms; S_social(H j ) is the hotspot H j Interaction volume on social media platforms; S_video(H j ) is the hotspot H j Views on short video platforms; , , These are the scores of the three types of signals after Min-Max normalization, with values ​​ranging from [0, 1]. The subscripts min / max represent the extreme values ​​of all hotspots within the current batch; λ1 is the search volume signal weight, with a default value of 0.30; λ2 is the social interaction signal weight, with a default value of 0.40; λ3 is the video playback signal weight, with a default value of 0.30; the three weight coefficients satisfy the normalization constraint:

[0090] λ1 + λ2 + λ3 = 1, and λ1, λ2, λ3 > 0

[0091] Step S32: Time decay correction, introducing an exponential decay factor to adjust W_raw(H) j Timeliness adjustments are made to ensure that hot topics are both "hot and new," resulting in a comprehensive weight W_final(H). j ):

[0092] W_final(H j = W_raw(H j ) · exp(-λ_t · Δt j )

[0093] in:

[0094] Δt j : Current time t_now and the time t_now when the hot topic was published j The difference, in hours, is Δt. j = t_now -t j A larger value indicates that the hotspot is more outdated;

[0095] λ_t: Time decay coefficient, with a value range of [0.05, 0.20]; λ_t = 0.05 corresponds to a half-life of approximately 14 hours, suitable for slow-burning topics; λ_t = 0.20 corresponds to a half-life of approximately 3.5 hours, suitable for fast-paced, sudden hot topics; the default value is 0.10.

[0096] exp(·): Natural exponential function, ensuring that the weights decrease monotonically over time;

[0097] Step S4: Hotspot-book fragment semantic matching, using fragment semantic vector v si and keyword set K(C i ), and the comprehensive weight of hotspots W_final(H j ), calculate the overall matching score between each segment and each hotspot;

[0098] Step S41: Dual-channel similarity calculation, employing a dual-channel fusion strategy of "semantic similarity + keyword overlap" to comprehensively measure the relevance of the segment to the hot topic;

[0099] Semantic Channel: Based on Semantic Vector v si For hotspot H j Similarly, the semantic vector is obtained by encoding using the Encoder model. Calculate the cosine similarity:

[0100] Sim_sem(C i Hj ) = (v si · v hj ) / (‖v si ‖ · ‖v hj ‖)

[0101] Where v hj For hotspot H j The semantic vector of the text is encoded using the same Encoder model as in step S1, with dimension d; (·): vector dot product operation; ‖·‖: vector L2 norm; The value ranges from -1 to 1; the closer the value is to 1, the more semantically relevant it is.

[0102] Keyword channel: Using a set of fragment keywords K(C) i ) and the set of hot keywords K(H) j ), calculate the Jaccard coefficient:

[0103] Sim_kw(C i H j ) = |K(C i ) ∩ K(H j )| / |K(C i ) ∪ K(H j )|

[0104] in:

[0105] K(H j Hotspot H j The keyword set is extracted using the same TF-IDF method as in step S1.

[0106] K(C i ) ∩ K(H j ): The intersection of two keyword sets, representing the keywords that appear in each other's names.

[0107] K(C i ) ∪ K(H j ): The union of two keyword sets, representing all unique keywords.

[0108] Sim_kw(C i H j ): Value range [0, 1], the larger the value, the more keyword overlap.

[0109] Dual-channel fusion:

[0110] Sim(C i H j = α_s · Sim_sem(C iH j ) + (1 - α_s) · Sim_kw(C i H j )

[0111] in:

[0112] α_s: Semantic channel weight, ranging from [0, 1], with a default value of 0.70; (1 - α_s) is the keyword channel weight, with a default value of 0.30; the semantic channel weight is higher because it can capture implicit semantic associations;

[0113] Step S42: Calculate the weighted matching score. Multiply the dual-channel fusion similarity by the hotspot comprehensive weight to obtain the weighted matching score Score_match(C) for the segment-hotspot pair. i H j ):

[0114] Score_match(C i H j ) = Sim(C i H j ) · W_final(H j )

[0115] Take fragment C i The maximum matching score among all m hotspots is taken as the hotspot association strength of that segment:

[0116] MaxMatch(C i ) = max j=1,…,m Score_match(C i H j )

[0117] in:

[0118] max j=1,…,m : Take the maximum value of j from 1 to m

[0119] MaxMatch(C i Fragment C i The strongest hotspot correlation strength is used as a quality indicator of the matching between the segment and the current hotspot environment;

[0120] Step S5: Comprehensive selection of candidate content, and score the content quality using the Score (C). i The correlation strength between the hotspot and the MaxMatch(C) function. i Multiply the results to form a comprehensive score, and select the segment with the highest score to proceed to the next generation process;

[0121] Step S51: Calculate the overall score. The overall score of segment Cᵢ is defined as the product of the content quality score and the strength of its relevance to trending topics. A high score is achieved when both scores are high.

[0122] FinalScore(C i ) = Score(C i MaxMatch(C) i )

[0123] FinalScore(C i ) is fragment C i The overall score quantifies both content quality and relevance to trending topics;

[0124] Step S52, Diversity Constraints and Candidate Selection: Sort by FinalScore in descending order and select the top K (default K = 3~5) segments as candidates; To avoid high content repetition, a diversity constraint is introduced: If the semantic similarity between two candidate segments exceeds the threshold θ_div (default 0.90), only the one with the higher FinalScore is retained;

[0125] Step S6: Marketing copy is automatically generated, based on the book excerpt C selected in step S5. i and the optimal hotspot H corresponding to this segment in step S4 j * It calls the conditional language generation model to automatically generate marketing copy with the characteristics of "attracting clicks and emotional resonance";

[0126] Step S61: Conditional copy generation, using a conditional language generation model, with fragment semantic vector v si and hotspot semantic vector v hj* Given the condition, the autoregressive text generation Text = {w1, w2, …, w T}:

[0127] P(Text | C i H j *) = ∏ t=1 T P(w t | w1, …, w t-1 , v si , v hj* )

[0128] in:

[0129] Text: The generated marketing copy consists of T tokens;

[0130] w t The t-th word element in the text;

[0131] w1, …, w t-1 : The sequence of historical lexical units generated at step t (autoregressive condition);

[0132] T: Maximum number of metawords in the text, defaults to 100-200, to adapt to the character limit of short video cover / subtitle;

[0133] Step S62: Joint optimization objective. The model training objective is to jointly minimize the following loss function L, taking into account language fluency, click appeal, and sentiment consistency:

[0134] L = L_LM + λ_attr · L_attract + λ_emo · L_emotion

[0135] in:

[0136] L_LM: Cross-entropy loss of standard language models, ensuring the linguistic validity of the text;

[0137] L_attract: Attractiveness loss: Introducing a pre-trained click-through rate prediction model CTR(·), L_attract = 1 - CTR(Text); CTR(·) scores candidate text, and the negative gradient direction encourages the generation of more attractive text;

[0138] L_emotion: Loss of emotional consistency: L_emotion = |F_emotion(Text) - F_emotion(C) i F_emotion(·) is the emotion intensity function defined in step S21, which penalizes the deviation between the generated copy and the emotional tendency of the original book excerpt.

[0139] λ_attr: Attraction loss weight, default value 0.30, controls the degree to which the generated copy optimizes the click-through rate;

[0140] λ_emo: Emotional consistency loss weight, default value 0.20, controls the degree of consistency between the copy and the book's emotion;

[0141] Step S7: Video script structure design. The book segment selected in step S5 is decomposed into scenes, and the duration of each scene is automatically allocated based on the importance of the scenes to complete the video rhythm design.

[0142] Step S71: Scene sequence construction, converting book fragments C i Divide the video script into K scenes based on plot dividing points:

[0143] Script = {Scene1, Scene2, …, Scene K}

[0144] in:

[0145] Scene k The k-th scene text block (k = 1, 2, …, K) corresponds to a plot unit in a book excerpt;

[0146] K: Total number of scenes, automatically determined based on segment length and plot dividing points;

[0147] Step S72: Scene importance scoring, for each scene... k The importance score is calculated by considering four dimensions:

[0148] Imp(Scene k = w_e · F_emo(Scene) k ) + w_c · F_con(Scene k ) + w_p · Pos(k, K) + w_h · Hot(Scene k )

[0149] in:

[0150] F_emo(Scene k Scene k The intensity of emotion is calculated at the scene level using the F_emotion formula defined in step S21;

[0151] F_con(Scene k Scene k The conflict intensity is calculated at the scene granularity using the F_conflict formula defined in step S21;

[0152] Pos(k, K): Scene position weight: 1.5 for the first scene (k = 1) and the last scene (k = K), and 1.0 for the middle scenes; the first scene is responsible for grabbing attention, and the last scene is responsible for leaving suspense.

[0153] Hot (Scene) k ): Scene hotspot correlation strength, taken from the entire segment C in step S4. i MaxMatch(C i The Hot(Scene) value is the same for all scenes within the same segment. k = MaxMatch(C i );

[0154] w_e: Emotion intensity weight, default value 0.30;

[0155] w_c: Conflict intensity weight, default value 0.30;

[0156] w_p: Position weight coefficient, default value 0.20;

[0157] w_h: Hotspot association weight, default value 0.20;

[0158] The four weighting coefficients satisfy the normalization constraint:

[0159] w_e + w_c + w_p + w_h = 1, and w_e, w_c, w_p, w_h > 0

[0160] Step S73: Scene duration allocation. The duration of each scene is proportionally allocated from the total duration T_total according to its importance.

[0161] Duration(Scene k ) = [Imp(Scene k ) / · T_total

[0162] in: The sum of the importance of all K scenes is used for normalization so that the sum of the duration of all scenes equals T_total;

[0163] T_total: Total video duration, recommended range 15-60 seconds, dynamically set according to the target platform;

[0164] Constraints: The minimum duration of each scene shall not be less than 1.5 seconds to ensure readability; if a scene falls below this minimum after normalization, the duration shall be supplemented by the minimum and the duration of the remaining scenes shall be renormalized.

[0165] Step S8: Scene generation, generating each scene output from step S7. k Further breakdown is performed to generate shot sequences containing visual cues, voice-over text, and duration, for use in subsequent multimodal material generation;

[0166] Step S81: Construct the shot sequence, which involves creating a scene sequence for each scene. K Break it down into several shots based on semantic cut points:

[0167]

[0168] in:

[0169] Scenek The i-th shot (i = 1, 2, …, N) k )

[0170] N k Scene k The total number of shots included is automatically determined based on the length of the scene text;

[0171] Each lens is defined as a triplet:

[0172]

[0173] in:

[0174] Image generation prompts for the shot include subject description (Who / What), scene description (Where), mood (Mood), and style keywords (Style), which are used by the image generation model in step S9;

[0175] The voiceover text for the scene is extracted from the content of the book corresponding to the scene or the text generated in step S6, and used for speech synthesis in step S9.

[0176] Shot duration (seconds), allocated according to importance in step S82;

[0177] Step S82: Shot importance rating and duration allocation, for each shot. Importance is calculated by combining three dimensions: visual, action, and emotion.

[0178]

[0179] in:

[0180] The visualization potential of lens text is calculated at the lens granularity using the F_visual formula;

[0181] Action word density in the scene text, calculated as the ratio of words belonging to A_dict to the total number of words;

[0182] The emotional intensity of the shot text is calculated at the shot granularity using the F_emotion formula;

[0183] Visual rating weight, default value 0.40

[0184] Action word density weight, default value 0.30

[0185] Sentiment intensity weight, default value 0.30

[0186] The weights satisfy the normalization constraint:

[0187] + + = 1, and , , > 0

[0188] Shot duration is the same as scene duration. k The content is allocated according to its importance.

[0189]

[0190] in:

[0191] Scene k The sum of the importance of all shots is used for normalization;

[0192] Duration(Scene k ): Scene duration (seconds) allocated in step S73;

[0193] Constraint: The minimum duration of each shot must be no less than 0.5 seconds;

[0194] Step S83, Visual cue word generation specifications, The following specifications must be followed during the construction process to ensure the quality of subsequent image generation;

[0195] It must include four elements: subject description (Who / What), scene description (Where), mood, and style keywords (Style);

[0196] All shots within the same Scene share the same StyleSeed to ensure a consistent visual style.

[0197] If a specific character appears in a clip, the character's appearance description must be consistent throughout the entire video to avoid character design drift.

[0198] Step S9: Multimodal material synthesis. Based on the storyboard triplet, the image generation, video generation, and speech synthesis modules are called respectively to generate images, video clips, and dubbing audio. Semantic alignment constraints are used to ensure consistency between images and text.

[0199] Step S91, Image generation, based on the image generated in step S83. Using StyleSeed as input, the textural image model is called to generate lens images:

[0200]

[0201] in:

[0202] (·): Text-generated image model, which generates images based on prompt words;

[0203] : The generated lens image;

[0204] Step S92: Video clip generation. Based on the generated static images, the video generation model is driven to generate video clips with camera movement effects.

[0205]

[0206] in:

[0207] (·): Image-driven video generation model that generates video clips based on static images and camera movement parameters;

[0208] The camera movement type is automatically selected from {push-in, pull-out, pan, tilt-up, still} based on the mood of the scene; high-conflict scenes prioritize push-in shots, while calm scenes prioritize still shots.

[0209] The generated video clips are used for video compositing in step S10;

[0210] Step S93: Speech synthesis. Using AudioText_{k,i} defined in step S81 as text input, the TTS model is called to generate voice-over audio.

[0211]

[0212] in:

[0213] TTS(·): Text-to-speech synthesis model that converts text into natural speech;

[0214] VoiceStyle: Voice style parameters, including timbre (gender / age perception), speech rate (recommended 1.0~1.2×normal speech rate), and emotional style, which should be consistent with the mood of the scene;

[0215] The generated audio for the scene dubbing is used for audio-visual synchronization in step S10;

[0216] Step S94, Visual-Text Semantic Alignment Verification: To ensure semantic consistency between the generated image and the prompt words, the CLIP model is introduced to calculate the image-text alignment loss.

[0217]

[0218] in:

[0219] (·): The CLIP model image encoder, which converts the image... Mapped to a multimodal semantic vector space, the output vector dimension is equal to... Consistent; (·): The text encoder of the CLIP model will Mapped to the same multimodal semantic vector space;

[0220] cos(·, ·): Cosine similarity function, range [0, 1], the larger the value, the more consistent the semantics of the image and text;

[0221] Image-text alignment loss, range [0, 1]; = 0 indicates perfect alignment. = 1 indicates a complete deviation;

[0222] like If the threshold τ (alignment threshold, default τ = 0.25) is exceeded, regeneration will be triggered; a maximum of 3 retries will be made. If the threshold is still exceeded, the system will be marked as requiring manual review.

[0223] τ: CLIP alignment threshold, default value 0.25; the smaller the value, the stricter the requirements, suitable for high-precision marketing scenarios;

[0224] Step S95, Video Synthesis and Quality Optimization: The video clips and dubbing audio are spliced ​​together to synthesize a complete video, and the quality is optimized through transition design, audio-visual synchronization and rhythm control.

[0225] Step S10, Propagation Probability Prediction and Final Screening: The Video_final output in step S95 is used to predict the probability of viral spread. A final score is given based on the comprehensive content features and the intensity of the hot topic. The result is used to determine whether to output directly, manually review, or regenerate.

[0226] Extract five types of propagation-related features from Video_final:

[0227] F_topic: Topic popularity feature: Take segment C from step S4 i Optimal hotspot H j * of W_final(H j*), normalized to [0, 1] using Min-Max;

[0228] F_emo_v: Video emotion intensity feature, emotion intensity of each scene in step S7. Based on scene duration (Duration) k ) represents the weighted average of the weights;

[0229] F_novel: Content novelty feature, calculated by using the semantic vector v of the book segment corresponding to Video_final. si The complement of the maximum cosine similarity with the published video semantic vector library; the higher the value, the more novel the content.

[0230] F_rhythm: Rhythm matching feature, which is the ratio of the actual average shot switching frequency (times / second) of the video to the switching frequency recommended by the target platform, mapped to [0, 1] by Sigmoid;

[0231] W_final(H j *) Directly used as a hotspot intensity feature for propagation prediction;

[0232] The probability of viral spread, P_viral, is predicted using a propagation probability prediction model that employs a linear weighted average of five features followed by Sigmoid activation. P_viral = σ(θ0+ ​​θ1· F_topic + θ2· F_emo_v + θ3· F_novel + θ4· F_rhythm + θ5· W_final(H j *));

[0233] in:

[0234] σ(x): Sigmoid function, σ(x) = 1 / (1 + e^{-x}), maps linear combinations to the interval (0, 1), representing the probability of a hit product;

[0235] θ0: Bias term (intercept), obtained during model training;

[0236] θ1: Coefficient corresponding to F_topic (topic popularity);

[0237] θ2: Coefficient corresponding to F_emo_v (video emotion intensity);

[0238] θ3: Coefficient corresponding to F_novel (content novelty);

[0239] θ4: Coefficient corresponding to F_rhythm (rhythm matching degree);

[0240] θ5: W_final(H j *) Corresponding coefficient for (Comprehensive weighting of hot topics);

[0241] The parameter vector θ = {θ0, θ1, θ2, θ3, θ4, θ5} is obtained by performing logistic regression training on historical short video dissemination data (number of likes, number of shares, completion rate), and supports periodic updates to adapt to changes in platform algorithms;

[0242] The final comprehensive score and selection decision will be based on the final score (C) of the candidate content. i Multiplying this by the propagation probability P_viral yields the final comprehensive score:

[0243] Score_final = FinalScore(C i ) · P_viral

[0244] in:

[0245] FinalScore(C i Step S5 calculates the comprehensive score of candidate segments, which measures the content quality and relevance to trending topics;

[0246] P_viral: The predicted probability of a viral hit;

[0247] Score_final: The final overall score, reflecting content quality, relevance to trending topics, and potential for dissemination. A three-tiered selection process is based on Score_final.

[0248] If Score_final ≥ 0.60: Directly output Video_final, and the generation process ends;

[0249] 0.40 ≤ Score_final < 0.60: Mark as "Suggest manual review", and push Video_final and the quality report to the review queue;

[0250] Score_final < 0.40: Trigger regeneration, return to step S3 to reselect hotspots, or return to step S5 to select suboptimal candidate segments, with a maximum of 3 iterations.

[0251] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.

Claims

1. A multi-stage generation method for book short videos based on hot topic fusion, characterized by: Step S1: Content parsing and semantic modeling The input content is segmented and divided into semantic units to extract key events, relationships between people and themes, and to construct context-related semantic representations. A cross-segment association mechanism is introduced to model the context dependency of long texts, ensuring local independence and overall semantic consistency. Step S2: Content Quality Assessment and Candidate Screening The system calculates indicators such as the degree of plot conflict, intensity of emotional fluctuation, information density, and visual expression potential, and generates a comprehensive score through a weighted model. High-scoring content is then selected by combining redundancy removal constraints to form a diverse candidate set with strong dissemination potential. Step S3: Hotspot Information Modeling and Dynamic Updates Acquire multi-source hotspot data, represent it uniformly, and calculate the popularity weight; dynamically update hotspots through a time decay mechanism, and combine cross-platform data fusion correction to obtain a hotspot representation that reflects the current dissemination trend; Step S4: Hotspot semantic matching and content fusion By jointly calculating semantic similarity and keyword relevance, a mapping relationship between content and hot topics is established. The content is weighted according to the matching strength to generate an enhanced content representation that integrates hot topic information. Step S5: Candidate Content Selection and Sorting Based on a comprehensive ranking of content quality scores and trending topic matching results, a normalization strategy is used to balance the weights of content quality and trending topic popularity, and the target content most suitable for short video generation is selected. Step S6: Marketing copy generation and constraint optimization Based on the fused content representation, the text is generated through a conditional generation model. A multi-objective constraint mechanism is introduced to jointly optimize the text structure, semantic expression, and attractiveness, generating short video text that combines information expression and dissemination appeal. Step S7: Script Structure Generation and Rhythm Planning The script structure is constructed based on the text, the scenes are divided and the narrative order is designed, the importance of the scenes is modeled based on semantic strength, the keyness of the plot and the relevance of hot topics, and the structured planning of the video rhythm is achieved through time allocation and position weight adjustment. Step S8: Storyboard Modeling and Fine-Grained Design The scene is broken down into a sequence of shots, and visual descriptions, action information and audio text are generated for each shot. The importance is calculated and the duration is allocated based on visual complexity, action intensity and emotional expression to achieve fine-grained rhythm control. Step S9: Multimodal content generation and consistency constraints Based on the storyboard description, images, video clips and audio content are generated. Multimodal alignment is performed through a unified semantic space, and cross-modal consistency constraints are introduced to optimize the semantic deviation between visual and textual content, ensuring content consistency and coherence. Step S10: Optimization of video synthesis and propagation effects Multimodal content is spliced ​​sequentially, and a complete video is generated through transition control, audio-visual synchronization, and rhythm scheduling. A dissemination effect evaluation mechanism is constructed to predict the dissemination potential of videos and perform screening or parameter-level optimization to output short videos with high dissemination effect.

2. The multi-stage generation method for book short videos based on hotspot fusion according to claim 1, characterized in that: Step S1 further includes the following: Step S11: Story Fragment Division. Using a sliding window segmentation strategy, the book text B is divided according to semantic boundary paragraphs and chapters to obtain n semantically complete story fragments: ; in: Book: A collection of fragments of book text, containing all n story fragments; C i The i-th story segment (i = 1, 2, …, n) is divided into segments with the paragraph boundaries as the smallest unit. n: Total number of segments, dynamically determined by the book length and window size; Step S12, Segment semantic vectorization: For each segment C i A pre-trained semantic encoding model is used to generate semantic vector representations, which serve as features for subsequent similarity calculations. ; in: Fragment C i The semantic vector, with dimension d, is typically 768 or 1024 and is output by the Encoder model; Encoder(·): A pre-trained semantic encoding model that maps text of arbitrary length to a fixed-dimensional vector space. It can use either Sentence-BERT or BERT models. d: Semantic vector dimension, which depends on the selected encoding model; This step also extracts each fragment C. i Keyword set K(C) i ), using TF-IDF or TextRank algorithms: K(C i ) = TopK_TF-IDF(C i ); in: K(C i Fragment C i The keyword set is composed of the top K_w words with the highest weights selected by the TF-IDF algorithm. This step also counts the number of sentences in each segment |sent(C i | and word count | tokens(C i )|,|sent(C i )| is fragment C i The number of sentences is calculated using periods, question marks, and exclamation marks as clause boundaries; |tokens(C i )| is fragment C i The word count is calculated after word segmentation.

3. The multi-stage generation method for book short videos based on hotspot fusion according to claim 1, characterized in that: Step S2 further includes the following: Step S21: Calculate features of each dimension, and calculate segment C sequentially. i Five dimensions of characteristics: The plot conflict level, F_conflict, reflects the narrative tension of a segment. A conflict word dictionary, C_dict, is defined, including transition words, contrasting words, and conflict verbs. A size of 500-1000 words is recommended. The average number of times conflict words appear per sentence in the segment is calculated. F_conflict(C i ) = |{w j ∈ C i : w j ∈ C_dict}| / |sent(C i )|; in: C_dict: A conflict term dictionary, pre-annotated by domain experts, covering transitional and conflict expressions such as "but," "however," "confrontation," and "betrayal." {w j ∈ C i : w j ∈ C_dict}: Fragment C i The word w belongs to the conflict word dictionary j A set, |·| represents the number of its elements; The emotion intensity F_emotion reflects the degree of emotional fluctuation in a segment. For each sentence in the segment, a pre-trained sentiment analysis model is used to output an emotion polarity score e. s ∈ [-1, +1], take the mean of absolute values: F_emotion(C i ) = (1 / |sent(C i )|) · Σs |e s |; in: e s The sentiment polarity score of the s-th sentence is output by the sentiment analysis model, with a value range of [-1, +1], where negative values ​​represent negative sentiment and positive values ​​represent positive sentiment. |e s |:e s The absolute value of the value is used to include both positive and negative emotions in the statistics, with a high absolute value representing strong emotions. Σs: for fragment C i All in the middle | sent(C i Summing the sum of the sentences; Character complexity F_character reflects the richness of relationships between characters in a segment. It is calculated by extracting the set of named entities E(C) from the segment using named entity recognition. i ), calculate the average number of personal name entities per sentence: F_character(C i ) = |E(C i )| / |sent(C i )|; in: E(C i Fragment C i The set of personal names identified by the NER model, where |·| represents the number of elements; Information density F_density measures the degree of information condensation in a fragment, using the keyword set K(C i ), calculate the percentage of keywords in the total number of words in the segment: F_density(C i ) = |K(C i )| / |tokens(C i )|; in: |K(C i |: The set of keywords K(C) extracted by TF-IDF in step S1 i The number of words in ( ); |tokens(C i )|:Fragment C i Total word count; Visualization potential F_visual measures the ability of a fragment to generate visuals. It defines an action word dictionary A_dict (e.g., "running, fighting, exploding") and a scene word dictionary S_dict (e.g., "forest, castle, street"), and calculates the combined percentage of these two types of words in the fragment. F_visual(C i ) = |{w j ∈ C i : w j ∈ A_dict ∪ S_dict}| / |tokens(C i )|; in: A_dict: A dictionary of action verbs, containing verbs that express specific actions; a size of 300-500 words is recommended. S_dict: A scene-specific dictionary containing nouns representing visual scenes; a size of 300-500 words is recommended. A_dict ∪ S_dict: The union of the dictionaries of action words and scene words, with no duplicate entries for the two categories. Step S22: Multidimensional weighted scoring. The feature values ​​of the five dimensions are weighted and summed to obtain segment C. i Overall content score (C) i ): Score(C i ) = α · F_conflict + β · F_emotion + γ · F_character + δ · F_density + η · F_visual; in: α: Weighting coefficient for plot conflict, default value 0.

25. Conflict is the core element that attracts the audience, so it has the highest weight. β: Weighting coefficient for emotional intensity, default value 0.

25. Emotional resonance directly affects the dissemination rate. γ: Weighting coefficient for character complexity, default value 0.20, increases the memorability of character clues; δ: The weighting coefficient for information density, with a default value of 0.

15. Appropriate density ensures the integrity of the content. η: The weighting coefficient for visualization potential, with a default value of 0.15, which determines the quality of the generated image; The five weighting coefficients satisfy normalization constraints to ensure that the score range is consistent with the magnitude of each dimension: α + β + γ + δ + η = 1, and α, β, γ, δ, η > 0.

4. The multi-stage generation method for book short videos based on hotspot fusion according to claim 1, characterized in that: Step S3 further includes the following: Step S31: Cross-platform hotspot signal normalization and weighting; real-time collection of m hotspot data points to form a hotspot set: Hotspot = {H1, H2, …, Hm}; in: H j : The j-th hot topic (j = 1, 2, …, m), each hot topic includes text content, source platform, and publication time t. j property; m: Total number of hotspots collected in this batch; hotspot H j The original weights of each platform are calculated from the three types of signals. To eliminate dimensional differences, each signal must first be normalized to [0, 1] using Min-Max. = [S_search(H j ) - S_search_min] / [S_search_max - S_search_min]; , Normalize in the same way, then perform weighted fusion after normalization to obtain hotspots. The original weights of the single platform W_raw(H) j ): ; in Hot topic Search volume on search engine platforms; S_social(H j (This is a hot topic) Interaction volume on social media platforms; S_video(H j (This is a hot topic) Views on short video platforms; , , These are the scores of the three types of signals after Min-Max normalization, with values ​​ranging from [0, 1]. The subscripts min / max represent the extreme values ​​of all hotspots within the current batch; λ1 is the search volume signal weight, with a default value of 0.30; λ2 is the social interaction signal weight, with a default value of 0.40; λ3 is the video playback signal weight, with a default value of 0.30; the three weight coefficients satisfy the normalization constraint: λ1 + λ2 + λ3 = 1, and λ1, λ2, λ3 > 0; Step S32: Time decay correction, introducing an exponential decay factor to adjust W_raw(H) j Timeliness adjustments are made to ensure that hot topics are both "hot and new," resulting in a comprehensive weight W_final(H). j ): W_final(H j ) = W_raw(H j ) · exp(-λ_t · Δt j ); in: Δt j : Current time t_now and the time t_now when the hot topic was published j The difference, in hours, is Δt. j = t_now-t j A larger value indicates that the hotspot is more outdated; λ_t: Time decay coefficient, with a value range of [0.05, 0.20]; λ_t = 0.05 corresponds to a half-life of approximately 14 hours, suitable for slow-burning topics; λ_t = 0.20 corresponds to a half-life of approximately 3.5 hours, suitable for fast-paced, sudden hot topics; the default value is 0.

10. exp(·): Natural exponential function, ensuring that the weights decrease monotonically over time.

5. The multi-stage generation method for book short videos based on hotspot fusion according to claim 1, characterized in that: Step S4 further includes the following: Step S41: Dual-channel similarity calculation, employing a dual-channel fusion strategy of "semantic similarity + keyword overlap" to comprehensively measure the relevance of the segment to the hot topic; Semantic Channel: Based on Semantic Vector v si For hotspot H j Similarly, the semantic vector is obtained by encoding using the Encoder model. Calculate the cosine similarity: Sim_sem(C i , H j ) = (in si in hj ) / (‖in si ‖ · ‖in hj ‖); Where v hj For hotspot H j The semantic vector of the text is encoded using the same Encoder model as in step S1, with dimension d; (·): vector dot product operation; ||·||: vector L2 norm; Sim_sem(C i H j ): Values ​​range from -1 to 1; the closer the value is to 1, the more semantically relevant it is. Keyword channel: Using a set of fragment keywords K(C) i ) and the set of hot keywords K(H) j ), calculate the Jaccard coefficient: Sim_kw(C i , H j ) = |K(C i ) ∩ K(H j )| / |K(C i ) ∪ K(H j )|; in: K(H j Hotspot H j The keyword set is extracted using the same TF-IDF method as in step S1; K(C i ) ∩ K(H j ): The intersection of two keyword sets, representing the keywords that appear together; K(C i ) ∪ K(H j ): The union of two keyword sets, representing all unique keywords; Sim_kw(C i H j ): The value ranges from [0, 1]. The larger the value, the more overlapping the keywords. Dual-channel fusion: Yes(C i H j ) = α_s · Sim_sem(C i H j ) + (1 - α_s) · Sim_kw(C i H j ); in: α_s: Semantic channel weight, ranging from [0, 1], with a default value of 0.70; (1-α_s) is the keyword channel weight, with a default value of 0.30; the semantic channel weight is higher because it can capture implicit semantic associations; Step S42: Calculate the weighted matching score. Multiply the dual-channel fusion similarity by the hotspot comprehensive weight to obtain the weighted matching score Score_match(C) for the segment-hotspot pair. i H j ): Score_match(C i , H j ) = Sim(C i , H j ) · W_final(H j ); Take fragment C i The maximum matching score among all m hotspots is taken as the hotspot association strength of that segment: MaxMatch(C i ) = max j=1,…,m Score_match(C i , H j ); in: max j=1,…,m : Take the maximum value of j from 1 to m MaxMatch(C i Fragment C i The strongest hotspot correlation strength is used as a quality indicator of the matching between the segment and the current hotspot environment.

6. The multi-stage generation method for book short videos based on hotspot fusion according to claim 1, characterized in that: Step S6 further includes the following: Step S61: Conditional copy generation, using a conditional language generation model, with fragment semantic vector v si and hotspot semantic vector v hj* Given the condition, the autoregressive text generation Text = {w1, w2, …, w T }: P(Text | C i , H j *) = ∏ t=1 T P(w t | w1, …, w t-1 , v si , v hj* ); in: Text: The generated marketing copy consists of T tokens; w t The t-th word element in the text; w1, …, w t-1 : The sequence of historical lexical units generated at step t (autoregressive condition); T: Maximum number of metawords in the text, defaults to 100-200, to adapt to the character limit of short video cover / subtitle; Step S62: Joint optimization objective. The model training objective is to jointly minimize the following loss function L, taking into account language fluency, click appeal, and sentiment consistency: L = L_LM + λ_attr · L_attract + λ_emo · L_emotion; in: L_LM: Cross-entropy loss of standard language models, ensuring the linguistic validity of the text; L_attract: Attractiveness loss: Introducing a pre-trained click-through rate prediction model CTR(·), L_attract = 1 - CTR(Text); CTR(·) scores candidate text, and the negative gradient direction encourages the generation of more attractive text; L_emotion: Loss of emotional consistency: L_emotion = |F_emotion(Text) - F_emotion(C) i F_emotion(·) is the emotion intensity function defined in step S21, which penalizes the deviation between the generated copy and the emotional tendency of the original book excerpt. λ_attr: Attraction loss weight, default value 0.30, controls the degree to which the generated copy optimizes the click-through rate; λ_emo: Emotional consistency loss weight, default value 0.20, controls the degree of consistency between the copy and the book's emotion.

7. The multi-stage generation method for book short videos based on hotspot fusion according to claim 1, characterized in that: Step S7 further includes the following: Step S71: Scene sequence construction, converting book fragments C i Divide the video script into K scenes based on plot dividing points: Script = {Scene1, Scene2, …, Scene K }; in: Scene k The k-th scene text block (k = 1, 2, …, K) corresponds to a plot unit in a book excerpt; K: Total number of scenes, automatically determined based on segment length and plot dividing points; Step S72: Scene importance scoring, for each scene... k The importance score is calculated by considering four dimensions: Imp(Scene k ) = w_e · F_emo(Scene k ) + w_c · F_con(Scene k ) + w_p · Pos(k,K) + w_h · Hot(Scene k ); in: F_emo(Scene k Scene k The intensity of emotion is calculated at the scene level using the F_emotion formula defined in step S21; F_con(Scene k Scene k The conflict intensity is calculated at the scene granularity using the F_conflict formula defined in step S21; Pos(k, K): Scene position weight: 1.5 for the first scene (k = 1) and the last scene (k = K), and 1.0 for the middle scenes; the first scene is responsible for grabbing attention, and the last scene is responsible for leaving suspense. Hot (Scene) k ): Scene hotspot correlation strength, taken from the entire segment C in step S4. i MaxMatch(C i The Hot(Scene) value is the same for all scenes within the same segment. k = MaxMatch(C i ); w_e: Emotion intensity weight, default value 0.30; w_c: Conflict intensity weight, default value 0.30; w_p: Position weight coefficient, default value 0.20; w_h: Hotspot association weight, default value 0.20; The four weighting coefficients satisfy the normalization constraint: w_e + w_c + w_p + w_h = 1, and w_e, w_c, w_p, w_h > 0; Step S73: Scene duration allocation. The duration of each scene is proportionally allocated from the total duration T_total according to its importance. ; in: The sum of the importance of all K scenes is used for normalization so that the sum of the duration of all scenes equals T_total; T_total: Total video duration, recommended range 15-60 seconds, dynamically set according to the target platform; Constraints: The minimum duration of each scene shall not be less than 1.5 seconds to ensure readability; if a scene falls below this minimum after normalization, the duration shall be supplemented by the minimum and the duration of the remaining scenes shall be renormalized.

8. The multi-stage generation method for book short videos based on hotspot fusion according to claim 1, characterized in that: Step S8 further includes the following: Step S81: Construct the shot sequence, which involves creating a scene sequence for each scene. K Break it down into several shots based on semantic cut points: ; in: Scene k The i-th shot (i = 1, 2, …, N) k ); N k Scene k The total number of shots included is automatically determined based on the length of the scene text; Each lens is defined as a triplet: ; in: Image generation prompts for the shot include subject description (Who / What), scene description (Where), mood (Mood), and style keywords (Style), which are used by the image generation model in step S9; The voiceover text for the scene is extracted from the content of the book corresponding to the scene or the text generated in step S6, and used for speech synthesis in step S9. Shot duration (seconds), allocated according to importance in step S82; Step S82: Shot importance rating and duration allocation, for each shot. Importance is calculated by combining three dimensions: visual, action, and emotion. ; in: The visualization potential of lens text is calculated at the lens granularity using the F_visual formula; Action word density in the scene text, calculated as the ratio of words belonging to A_dict to the total number of words; The emotional intensity of the shot text is calculated at the shot granularity using the F_emotion formula; Visual scoring weight, default value 0.40; Action word density weight, default value 0.30; Emotional intensity weight, default value 0.30; The weights satisfy the normalization constraint: + + = 1, and , , > 0; Shot duration is the same as scene duration. k The content is allocated according to importance: ; in: Scene k The sum of the importance of all shots is used for normalization; Duration(Scene k ): Scene duration (seconds) allocated in step S73; Constraint: The minimum duration of each shot must be no less than 0.5 seconds; Step S83, Visual cue word generation specifications, The following specifications must be followed during the construction process to ensure the quality of subsequent image generation; It must include four elements: subject description (Who / What), scene description (Where), mood, and style keywords (Style); All shots within the same Scene share the same StyleSeed to ensure a consistent visual style. If a specific character appears in a clip, the character's appearance description must be consistent throughout the entire video to avoid character design shifts.

9. The multi-stage generation method for book short videos based on hotspot fusion according to claim 1, characterized in that: Step S9 further includes the following: Step S91, Image generation, based on the image generated in step S83. Using StyleSeed as input, the textural image model is called to generate lens images: ; in: (·): Text-generated image model, which generates images based on prompt words; : The generated lens image; Step S92: Video clip generation. Based on the generated static images, the video generation model is driven to generate video clips with camera movement effects. ; in: (·): Image-driven video generation model that generates video clips based on static images and camera movement parameters; The camera movement type is automatically selected from {push-in, pull-out, pan, tilt-up, still} based on the mood of the scene; high-conflict scenes prioritize push-in shots, while calm scenes prioritize still shots. The generated video clips are used for video compositing in step S10; Step S93: Speech synthesis. Using AudioText_{k,i} defined in step S81 as text input, the TTS model is called to generate voice-over audio. ; in: TTS(·): Text-to-speech synthesis model that converts text into natural speech; VoiceStyle: Voice style parameters, including timbre (gender / age perception), speech rate (recommended 1.0~1.2×normal speech rate), and emotional style, which should be consistent with the mood of the scene; The generated audio for the scene dubbing is used for audio-visual synchronization in step S10; Step S94, Visual-Text Semantic Alignment Verification: To ensure semantic consistency between the generated image and the prompt words, the CLIP model is introduced to calculate the image-text alignment loss. ; in: (·): The CLIP model image encoder, which converts the image... Mapped to a multimodal semantic vector space, the output vector dimension is equal to... Consistent; (·): The text encoder of the CLIP model will Mapped to the same multimodal semantic vector space; cos(·, ·): Cosine similarity function, range [0, 1], the larger the value, the more consistent the semantics of the image and text; Image-text alignment loss, range [0, 1]; = 0 indicates perfect alignment. =1 indicates a complete deviation; like If the threshold τ (alignment threshold, default τ = 0.25) is exceeded, regeneration will be triggered; a maximum of 3 retries will be made. If the threshold is still exceeded, the system will be marked as requiring manual review. τ: CLIP alignment threshold, default value 0.25; the smaller the value, the stricter the requirements, suitable for high-precision marketing scenarios; Step S95, Video Synthesis and Quality Optimization: The video clips and dubbing audio are spliced ​​together to synthesize a complete video, and the quality is optimized through transition design, audio-visual synchronization and rhythm control.

10. The multi-stage generation method for book short videos based on hotspot fusion according to claim 1, characterized in that: Step S10 further includes the following: Extract five types of propagation-related features from Video_final: F_topic: Topic popularity feature: Take segment C from step S4 i Optimal hotspot H j * of W_final(H j *), normalized to [0, 1] using Min-Max; F_emo_v: Video emotion intensity feature, emotion intensity of each scene in step S7. Based on scene duration (Duration) k ) represents the weighted average of the weights; F_novel: Content novelty feature, calculated by using the semantic vector v of the book segment corresponding to Video_final. si The complement of the maximum cosine similarity with the published video semantic vector library; the higher the value, the more novel the content. F_rhythm: Rhythm matching feature, which is the ratio of the actual average shot switching frequency (times / second) of the video to the switching frequency recommended by the target platform, mapped to [0, 1] by Sigmoid; W_final(H j *) Directly used as a hotspot intensity feature for propagation prediction; The probability of viral spread, P_viral, is predicted using a propagation probability prediction model that employs a linear weighted average of five features followed by Sigmoid activation. P_viral = σ(θ0+ ​​θ1· F_topic + θ2· F_emo_v + θ3· F_novel + θ4· F_rhythm+ θ5· W_final(H j *)); in: σ(x): Sigmoid function, Mapping the linear combination to the (0, 1) interval represents the probability of a hit product; θ0: Bias term (intercept), obtained during model training; θ1: Coefficient corresponding to F_topic (topic popularity); θ2: Coefficient corresponding to F_emo_v (video emotion intensity); θ3: Coefficient corresponding to F_novel (content novelty); θ4: Coefficient corresponding to F_rhythm (rhythm matching degree); θ5: W_final(H j *) Corresponding coefficient for (Comprehensive weighting of hot topics); The parameter vector θ = {θ0, θ1, θ2, θ3, θ4, θ5} is obtained by performing logistic regression training on historical short video dissemination data (number of likes, number of shares, completion rate), and supports periodic updates to adapt to changes in platform algorithms; The final comprehensive score and selection decision will be based on the final score (C) of the candidate content. i Multiplying this by the propagation probability P_viral yields the final comprehensive score: Score_final = FinalScore(C i ) · P_viral; in: FinalScore(C i Step S5 calculates the comprehensive score of candidate segments, which measures the content quality and relevance to trending topics; P_viral: The predicted probability of a viral hit; Score_final: The final overall score, reflecting content quality, relevance to trending topics, and potential for dissemination. A three-tiered selection process is based on Score_final. If Score_final ≥ 0.60: Directly output Video_final, and the generation process ends; 0.40 ≤ Score_final < 0.60: Mark as "Suggest manual review", and push Video_final and the quality report to the review queue; Score_final < 0.40: Trigger regeneration, return to step S3 to reselect hotspots, or return to step S5 to select suboptimal candidate segments, with a maximum of 3 iterations.