A Personalized Micro-Drama Writing and Storyboard Generation Method and System Based on a Large Model

By generating micro-drama content through structured processing based on user interest categories and cross-modal consistency, the problem of low content preference matching and limited automation in micro-drama creation has been solved, the production process has been optimized, and audiovisual logic gaps have been reduced.

CN122138024AInactive Publication Date: 2026-06-02SHANGHAI GUANCHI CULTURE COMM CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI GUANCHI CULTURE COMM CO LTD
Filing Date
2026-04-14
Publication Date
2026-06-02
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The creation of micro-dramas suffers from problems such as low content preference alignment, limited production automation, and audiovisual logical discontinuity.

Method used

By determining user interest categories based on historical user behavior data, target content text is generated and structured to generate scene structure. This is then combined with cross-modal consistency processing to generate micro-drama content.

Benefits of technology

It improved the alignment between creative content and user preferences, reduced human intervention, optimized the production cycle, and lowered the probability of audiovisual logical inconsistencies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122138024A_ABST
    Figure CN122138024A_ABST
Patent Text Reader

Abstract

This application relates to the field of information technology and discloses a method and system for personalized micro-drama scriptwriting and storyboard generation based on a large model. The method includes: generating target content text based on user interest categories determined from historical user behavior data; performing structured processing on the target content text to obtain a scene structure containing multiple scene units; generating a visual scheme based on the scene structure; performing cross-modal consistency processing on the visual scheme and the target content text; and generating micro-drama content based on the processing results. This can at least solve the technical problems faced by micro-drama creation, such as low content preference fit, limited production automation, and audiovisual logical discontinuity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information technology, and in particular to a method and system for personalized micro-drama scriptwriting and storyboard generation based on a large model. Background Technology

[0002] With the explosive growth of mobile internet and short video platforms, micro-dramas, as a high-frequency, fragmented new content format, have become a core product of digital entertainment consumption. Due to their characteristics such as short production cycles, fast-paced storylines, and high audience verticality, the industry has placed extremely high demands on the ability to personalize content and the efficiency of automated production.

[0003] In related technologies, the automated creation of micro-dramas or short videos mainly relies on a general-purpose large language model to execute end-to-end text generation logic. Typically, the system generates literary scripts based on a pre-set general-purpose prompt word-driven model. These scripts are usually presented in the form of continuous, unstructured narrative text. In the visual output stage, existing solutions often involve inputting the entire script into the video generation model or using keyword matching technology to retrieve relevant segments from a media library. In this creative mode, text generation and visual compositing are usually considered two independent execution modules. The system focuses on the surface-level semantic mapping, and the generated audiovisual content mainly relies on post-production manual proofreading, logical filtering, and editing encapsulation. Alignment and quality control between different modalities are achieved through manual intervention.

[0004] However, the inventors have discovered at least the following technical problems in the related technologies: the creation of micro-dramas faces low content preference matching, limited production automation, and audiovisual logic discontinuity. Summary of the Invention

[0005] One objective of this application is to provide a method and system for personalized micro-drama scriptwriting and storyboard generation based on a large model, at least to solve the technical problems faced by micro-drama creation in related technologies, such as low content preference fit, limited production automation, and audiovisual logic discontinuity.

[0006] To achieve the above objectives, some embodiments of this application provide the following aspects:

[0007] Firstly, some embodiments of this application provide a method for personalized micro-drama scriptwriting and storyboard generation based on a large model. The method includes: generating target content text based on user interest categories determined from user historical behavior data; performing structured processing on the target content text to obtain a scene structure containing multiple scene units; the structured processing of the target content text to obtain the scene structure containing multiple scene units includes: calculating a plot complexity index based on the target content text by statistically analyzing the number of plot points, plot branches, and semantic transitions; determining scene splitting granularity parameters based on a comparison of the plot complexity index with a preset threshold; performing sequence labeling processing on the target content text based on the scene splitting granularity parameters to obtain scene labeling results; extracting scene element information based on the scene labeling results to generate a scene structure containing the multiple scene units; generating a visual scheme based on the scene structure; performing cross-modal consistency processing on the visual scheme and the target content text, and generating micro-drama content results based on the processing results.

[0008] Secondly, some embodiments of this application also provide a system comprising: one or more processors; and a memory storing computer program instructions that, when executed, cause the processor to perform the steps of the method described above.

[0009] Compared with related technologies, the solution provided in this application establishes a connection between text description and audience characteristics at the source of content creation by determining user interest categories based on historical user behavior data and generating target content text accordingly. This allows scriptwriting to shift towards a data-driven approach to some extent, helping to alleviate the problem of low content preference matching and thus improving the alignment between created content and potential user preferences, reducing the randomness of content generation. By structuring the target content text to obtain a scene structure containing multiple scene units, and generating a visual scheme based on the scene structure, unstructured literary descriptions can be transformed into logically hierarchical business units. This provides basic data guidance for subsequent visual generation and reduces reliance on manual storyboarding and material selection to some extent, helping to alleviate the problem of limited production automation and assisting in reducing the impact of short scripts. Lowering the barrier to entry for short drama visuals can optimize the production cycle to some extent. By structuring the target content text to obtain a scene structure containing multiple scene units, and generating a visual scheme based on the scene structure, unstructured literary descriptions can be transformed into business units with logical hierarchy. This provides basic data guidance for subsequent visual generation, thereby reducing reliance on manual storyboarding and material selection to some extent, helping to alleviate the problem of limited production automation, lowering the barrier to entry for short drama visuals, and optimizing the production cycle. By performing cross-modal consistency processing on the visual scheme and the target content text, and generating short drama content results based on the processing results, a semantic comparison-based verification mechanism can be introduced between the text modality and the image modality to help identify and adjust feature offsets that may occur during cross-modal transformation. This helps to alleviate the problem of audiovisual logical discontinuity and reduce the probability of audio-visual disconnect. Attached Figure Description

[0010] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0011] Figure 1 An exemplary flowchart of a personalized micro-drama writing and storyboard generation method based on a large model is provided for some embodiments;

[0012] Figure 2 An exemplary structural diagram of a personalized micro-drama writing and storyboard generation system based on a large model is provided for some embodiments. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0014] In this embodiment of the disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good morals.

[0015] First Embodiment

[0016] The first embodiment relates to a personalized micro-drama scriptwriting and storyboard generation method based on a large model. This embodiment utilizes the large model's ability to understand user preferences to automatically complete the entire creation chain from text conception to visual scheme generation, aiming to improve the efficiency and personalization of micro-drama creation. Figure 1 As shown, the method may include the following steps:

[0017] Step S101: Generate target content text based on the user interest categories determined by the user's historical behavior data;

[0018] Step S102: Perform structured processing on the target content text to obtain a scene structure containing multiple scene units;

[0019] Step S103: Based on the scene structure, generate a screen layout;

[0020] Step S104: Perform cross-modal consistency processing on the screen scheme and the target content text, and generate the micro-drama content result based on the processing result.

[0021] The following sections will provide a detailed explanation of each of the above steps.

[0022] Regarding step S101, for example, in some cases, the user's interest category can be determined first based on the user's historical behavior data. The user's historical behavior data refers to the original interaction records generated by the user in the past when using relevant platforms (such as short videos, social media, e-commerce, etc.). The user's historical behavior data may include, but is not limited to: browsing behavior (such as viewing time, completion rate), interaction behavior (such as likes, comments, forwarding, collection), search behavior (such as search keywords, order of clicking search results), etc.

[0023] The user interest categories refer to generalized user profile tags or probability distributions obtained by analyzing the user's historical behavior data. These user interest categories may include, but are not limited to: genre preferences (e.g., suspense, workplace), pace preferences (e.g., high-conflict, fast-paced, slow-paced, healing), visual styles (e.g., traditional Chinese style, American retro), and emotional needs (e.g., stress relief, brain-teasing).

[0024] In some examples, determining user interest categories based on historical user behavior data may specifically include: extracting features from the historical user behavior data to obtain a behavior feature vector; performing cluster analysis based on the behavior feature vector to obtain a user behavior pattern; and determining the user interest category based on the user behavior pattern.

[0025] Specifically, in the process of feature extraction from the user's historical behavior data, preprocessing operations can be performed on the data first. Multi-dimensional numerical indicators, including but not limited to interaction frequency, behavior intensity, content semantic features, and category preferences, can be extracted through feature engineering to construct a high-dimensional spatial vector representing user behavior characteristics, i.e., a behavior feature vector. Further, the behavior feature vector can be mapped to a preset feature space, and iterative clustering can be performed using clustering algorithms (such as K-means, Gaussian mixture models, or density clustering algorithms). The feature vectors with centripetal spatial distribution are then grouped into specific behavior clusters, and user behavior patterns representing the core commonalities of these clusters are extracted. Then, a weighted mapping model between user behavior patterns and a preset interest tag matrix can be established. By calculating the semantic activation intensity of specific behavior patterns on different interest dimensions, the abstract behavior patterns are decoded into structured user interest categories that can be parsed by large models.

[0026] Furthermore, target content text can be generated based on the user interest categories. This target content text represents a personalized script prototype or story outline generated by the large model based on the user interest categories, serving as the underlying logical input for subsequent structured processing and storyboard generation.

[0027] Specifically, the user interest categories can be converted into feature labels recognizable by a large model, and combined with a pre-set scriptwriting template to form structured generation instructions. These instructions are then input into a pre-trained large model, which utilizes its semantic generation capabilities to create text within the thematic scope constrained by the user interest categories, generating preliminary script content. By performing a logical consistency check on the preliminary script content, the target content text can be output.

[0028] Specifically, regarding step S102, the scene unit refers to the smallest logical segment constituting the script of a short drama. Each scene unit may contain specific spatiotemporal information, character relationships, and core action descriptions. The scene structure refers to a structured data sequence composed of multiple scene units arranged in a logical order according to the plot.

[0029] In the specific implementation process, semantic analysis can be performed on the target content text to identify spatiotemporal transition features and character change features. Based on these features, the target content text is segmented to obtain multiple physical segments. For each physical segment, element extraction is performed to identify the location, characters, and core events. The extracted element information is then encapsulated into multiple scene units and combined according to the plot sequence to form the scene structure.

[0030] Specifically, regarding step S103, the visual scheme refers to a set of visual images or a sequence of frames corresponding to the scene structure, used to represent the script content.

[0031] Specifically, visual descriptive elements can be extracted from each scene unit of the scene structure. These visual descriptive elements may include scene objects, action relationships, and environmental attributes. Based on these visual descriptive elements, corresponding image generation control parameters are constructed. These parameters guide the image generation model to generate visual content that conforms to the scene description. Furthermore, the image generation control parameters can be input into a preset image generation model to obtain a corresponding frame sequence, and the scene scheme can be constructed based on this frame sequence.

[0032] Specifically, regarding step S104, the cross-modal consistency processing refers to the process of ensuring that the visual presentation and narrative logic achieve a high degree of consistency in semantics and time by performing multi-dimensional correlation analysis on the visual scheme and the target content text.

[0033] Specifically, the visual features of the image scheme and the semantic features of the target content text can be extracted separately, and the degree of content matching between the two can be quantitatively evaluated by establishing a cross-modal mapping relationship. During this process, the alignment status of image elements and text descriptions can be monitored in real time to identify whether there are audiovisual content disconnects or semantic deviations due to the randomness of generation.

[0034] Furthermore, the cross-modal consistency processing can also include temporal synchronization of visual sequences and literary descriptions, that is, by setting a unified time base, discrete screen frames are logically anchored to their corresponding script content. Based on the feedback results of the consistency processing, necessary closed-loop corrections can be made to the visual scheme or the target content text until a complete video content containing text, visuals, and corresponding binding relationships is generated, i.e., the micro-drama content result is output.

[0035] For steps S101-S104 above, for example, the system can obtain user A's raw interaction logs for the past month and extract that the click frequency of the keywords "deep palace" and "mystery solving" exceeds 60%. Based on this, the user's interest category is determined as a tag set {theme: ancient style, style: suspense, core element: locked room}. After receiving this tag set, the large model generates a target content text T1 containing 500 words, the core plot of which is: "The painter finds bloodstained silk by a dry well in the cold palace and tracks down the missing concubine by comparing the silk's texture." Then, semantic scanning can be performed on text T1 to identify two key points: "by the dry well (location change)" and "entering the secret passage (space change)," thereby automatically dividing text T1 into three scene units: U1 (discovery by the well), U2 (infiltration of the secret passage), and U3 (revelation of the truth). These three units are arranged in logical chronological order, forming a structured scene structure S1.

[0036] Furthermore, for scene unit U1, visual feature words "moonlight, dry well, blood-stained silk" can be extracted. These features are converted into image generation parameters, and the model is called to generate a scene scheme V1 containing 30 keyframes. In V1, each frame contains a cool-toned background consistent with the "ancient style" and the core prop "blood-stained silk". The keyword "red silk" extracted from text T1 is compared with the pixel features in scene scheme V1. The semantic consistency score of the two is calculated to be 0.92, which is higher than the preset threshold of 0.85, and is judged as a match. Subsequently, the system aligns the narration in text T1 with the scene V1 on the timeline, and finally synthesizes and outputs the short drama content result W1.

[0037] It is not difficult to see that, compared with related technologies, the solution provided in this application, by determining user interest categories based on user historical behavior data and generating target content text accordingly, can establish a connection between text description and audience characteristics at the source of content creation. This allows scriptwriting to shift towards a data-driven approach to some extent, helping to alleviate the problem of low content preference fit, thereby improving the fit between created content and potential user preferences and reducing the blindness of content generation. By structuring the target content text to obtain a scene structure containing multiple scene units, and generating a screen scheme based on the scene structure, unstructured literary descriptions can be transformed into business units with logical hierarchies. This provides basic data guidance for subsequent screen generation and reduces reliance on manual storyboarding and material selection to some extent, helping to alleviate the problem of limited production automation and assisting in... Lowering the threshold for acquiring footage for micro-dramas can optimize the production cycle to some extent. By structuring the target content text to obtain a scene structure containing multiple scene units, and generating a scene scheme based on the scene structure, unstructured literary descriptions can be transformed into business units with logical hierarchy. This provides basic data guidance for subsequent scene generation, thereby reducing reliance on manual storyboarding and material selection to some extent, helping to alleviate the problem of limited production automation, and further lowering the threshold for acquiring footage for micro-dramas and optimizing the production cycle. By performing cross-modal consistency processing on the scene scheme and the target content text, and generating the micro-drama content result based on the processing result, a semantic comparison-based verification mechanism can be introduced between the text modality and the image modality to help identify and adjust feature offsets that may occur in cross-modal transformation. This helps to alleviate the problem of audiovisual logical discontinuity and reduce the probability of audio-visual disconnect.

[0038] Second Embodiment

[0039] The second embodiment relates to a personalized micro-drama scriptwriting and storyboard generation method based on a large model. The second embodiment is an improvement upon the first embodiment, specifically in that it provides a concrete implementation method for generating target content text based on user interest categories determined from historical user behavior data.

[0040] Specifically, step S101, which involves generating target content text based on user interest categories determined from user historical behavior data, may include:

[0041] Step S1011: Determine the target content topic based on the user's interest category;

[0042] Step S1012: Based on the target content topic, extract sentiment feature vectors as sentiment elements using a sentiment classification model or a semantic embedding model;

[0043] Step S1013: Based on the emotional elements, construct the plot structure through a preset plot template or plot state transition rules;

[0044] Step S1014: Generate initial content text based on the plot structure;

[0045] Step S1015: Generate the target content text based on the initial content text.

[0046] Specifically, regarding step S1011, the system can retrieve a matching core intent from a preset theme library based on the user's interest category, thereby determining the target content theme. The target content theme defines the macro-narrative scope and basic tone of the micro-drama. For example, if the user's interest category is "ancient style suspense," then the determined target content theme is "ancient locked-room mystery solving."

[0047] Regarding step S1012, the system can perform deep semantic mining on the target content theme, use a sentiment classification model or a semantic embedding model to encode the features of the target content theme, extract sentiment feature vectors that can represent the emotional tone of the script, and identify the corresponding set of keywords based on the sentiment feature vectors, thereby obtaining sentiment elements.

[0048] The emotional element can be represented as the mapping result of the emotional feature vector in a preset emotional label space, which is used to depict the emotional features or psychological motivations hidden within the theme that can resonate with users.

[0049] For example, for the theme of "nostalgia for growing up", the system extracts the corresponding emotional feature vector through a semantic embedding model and matches it in the emotional label space. The emotional elements obtained can include "making up for childhood regrets" and "feeling the passage of time".

[0050] Specifically, regarding step S1013, the emotional elements can be used as a logical driver to construct a plot structure through the preset plot template or the plot state transition rules. The plot state transition rules describe the semantic transition relationships between different plot nodes, thereby generating a plot structure that includes a beginning, development, climax, and ending. The plot structure can refer to an ordered structure composed of multiple logical nodes and their connections, where each logical node corresponds to a different plot stage, and the connections between nodes are determined by the plot state transition rules, used to establish the structure of the target content text.

[0051] For example, based on the emotional element of "making up for childhood regrets", the system can generate a plot structure with the logical sequence of "encountering old objects - starting a journey to find one's roots - uncovering old misunderstandings - achieving self-reconciliation" according to a preset plot template or state transition path.

[0052] Optionally, in some embodiments, the preset plot template may be represented as a structured template containing multiple plot stages, wherein the plot stages include at least a beginning, development, climax and ending, and each plot stage corresponds to at least one plot slot, and each plot slot is used to fill semantic content related to the emotional element.

[0053] Furthermore, each of the aforementioned preset plot templates can be represented by the following structure:

[0054] ;

[0055] in, This refers to the preset plot template; Indicates the first Each plot stage, This indicates the total number of plot stages contained in the preset plot template; Indicates the first Each plot stage corresponds to a stage type, and the stage type may include at least one of the following: introductory event, escalation of conflict, climax revelation, or emotional release; Indicates the first The semantic constraints corresponding to each plot stage may include at least one of the following: emotional tags, character relationship constraints, or event type constraints.

[0056] For example, for the theme of "nostalgic childhood," the corresponding preset plot templates could include:

[0057] S1 (Beginning): Introduces the triggering event ( (Contains the semantic tag "memory trigger");

[0058] S2 (Development): Character Action Deployment ( (Includes the semantic tag "exploring the past").

[0059] S3 (Climax): Conflict Revealed (Contains semantic tags of "regret" or "misunderstanding").

[0060] S4 (Ending): Emotional Release (Contains semantic tags such as "reconciliation" or "reconciliation").

[0061] In some embodiments, the plot state transition rule can be represented as a state transition graph or a state transition matrix, used to describe the transition relationships between different plot nodes. The plot state transition rule can be represented as:

[0062] ;

[0063] in, This represents a set of plot states, with each state corresponding to a plot node. This represents the set of transition relationships between different plot states; This represents the transition probability distribution, used to characterize the likelihood of transitioning from the current plot state to the next plot state.

[0064] For example, regarding the emotional element of "making up for childhood regrets," the following set of states can be constructed:

[0065] V = {Triggering memories, exploring clues, revealing conflicts, releasing emotions}

[0066] The corresponding transition relationships and transition probabilities may include:

[0067] (Trigger memory → Explore clues) = 1.0;

[0068] (Exploring clues → Revealing conflicts) = 0.8;

[0069] (Exploring clues → Emotional release) = 0.2.

[0070] In some embodiments, the emotional elements are used to weight and modify the plot state transition rules to adjust the selection probability of different plot paths. Specifically, the system can adjust the transition probability distribution based on the emotional polarity or semantic tags in the emotional elements. The system can make dynamic adjustments. For example, when the emotional element contains the "regret" label, the system can increase the probability of transitioning to the "conflict disclosure" state; when the emotional element contains the "healing" label, the system can increase the weight of transitioning to the "emotional release" state.

[0071] The preset plot template or plot state transition rule can be obtained based on historical script data statistics, or by training sequence modeling on the script corpus.

[0072] Specifically, regarding step S1014, the initial content text refers to a preliminary draft of the script, containing specific dialogue and action descriptions, initially generated based on the plot skeleton. Specifically, the plot structure can be input into a pre-trained large model, and the model's text generation capabilities can be used to fill in specific content based on each node of the plot structure, thereby generating the initial content text.

[0073] Specifically, in step S1015, consistency checks and language polishing can be performed on the initial content text to correct logical breaks or redundant expressions, ultimately generating the target content text that can be used in subsequent steps. For example, by polishing the initial content text B1, redundant dialogue that is too straightforward can be removed and suspense and twists can be enhanced to generate the final target content text T1.

[0074] The consistency check refers to the logical judgment process that assesses whether the intent conveyed by the text content matches the expected visual presentation constraints.

[0075] Optionally, in some embodiments, the step of generating the target content text based on the initial content text, i.e., step S1015, may include:

[0076] Step S10151: Generate an initial scene prediction representation based on the semantic features of the initial content text;

[0077] Step S10152: Extract the storyboard constraint parameters based on the initial storyboard prediction representation;

[0078] Step S10153: Generate multiple plot branches based on the large model for the initial content text;

[0079] Step S10154: In the process of generating the plot branch, the storyboard constraint parameter is introduced to constrain and control the plot branch, and the modulated plot branch is obtained.

[0080] Step S10155: Weight allocation is performed on the modulated plot branches, and sequence splicing or linear combination processing is performed on the weighted plot branches to generate the target content text.

[0081] Specifically, in step S10151, deep semantic analysis can be performed on the initial content text to extract the semantic features. The semantic features may include at least one of the following: spatial orientation, subject action, and lighting atmosphere.

[0082] Furthermore, a pre-defined text-visual association model can be used to transform the semantic features into preliminary shot description information, thereby generating an initial storyboard prediction representation. The generation of the initial storyboard prediction representation refers to a preliminary visual mapping containing shot language features, pre-deduced from the text description. Specifically, the initial storyboard prediction representation is used to characterize the preliminary storyboard description data deduced from the semantic information in the initial content text. This storyboard prediction representation can be represented in a structured data form, such as: (scene number, shot type, perspective parameters, main object, action description, environmental atmosphere). For example, for the text fragment "A swordsman walks slowly through a bamboo forest," the system can generate the following initial storyboard prediction representation: (Scene 1, medium shot, eye level, swordsman, walking slowly, nighttime bamboo forest environment).

[0083] It is worth mentioning that the preset text-visual association model is a well-known technology to those skilled in the art or can be trained using existing deep learning frameworks (such as cross-modal mapping models based on CLIP or Transformer architectures). The main function of this model is to establish an association matrix between the text semantic space and the visual parameter space (such as shot size, camera movement, composition, etc.). This application focuses on using the results produced by this model for subsequent plot modulation and verification. Therefore, this will not be elaborated further.

[0084] Optionally, in some embodiments, the text-visual association model can be implemented using a cross-modal Transformer architecture, which includes a text encoder and a visual parameter prediction network. The text encoder performs semantic encoding on the input text and can generate text semantic vectors using a pre-trained language model (e.g., BERT or a Transformer encoder). The visual parameter prediction network predicts corresponding scene parameters based on the text semantic vectors. This network can include multiple fully connected layers and outputs the following visual parameters: shot type probability distribution (e.g., long shot, medium shot, close-up), camera viewpoint parameters, shot motion type, and scene composition.

[0085] In some embodiments, the training data for the text-visual association model may include data pairs consisting of script excerpts and corresponding storyboards. For example: Input text: "The detective walked into the dimly lit warehouse"; Corresponding storyboard parameters: Lens type: Medium shot Viewpoint: Eye level Environment: Low light.

[0086] The text-visual association model can be trained using supervised learning, enabling the network to predict corresponding scene parameters based on text semantics, thereby generating an initial scene prediction representation.

[0087] For step S10152, for example, the initial storyboard prediction representation can be parameterized to identify dimensions with high visual difficulty or conflict, and storyboard constraint parameters can be generated based on the semantic dimensions. The storyboard constraint parameters are used to characterize the set of storyboard control parameters extracted from the initial storyboard prediction representation.

[0088] In some embodiments, the storyboard constraint parameters are used to constrain and control narrative scheduling-related content. These parameters may include, but are not limited to, scene switching frequency, character appearance order, camera movement type parameters, and shot size limitation parameters. For example, when the initial storyboard prediction indicates multiple consecutive close-up shots, the system can generate corresponding shot continuity constraint parameters to limit large-scale scene jumps during plot generation.

[0089] For step S10153, for example, the multi-sample sampling mechanism of a large model can be used. Taking the initial content text as input, multiple candidate plot branches can be generated in the probability space by adjusting the generation strategy parameters or random sampling factors. The plot branch refers to a sequence of candidate texts that differ in local plot details or expression methods while maintaining the consistency of the main narrative logic. For example, for an initial content text describing "the protagonist lighting a torch to explore in a dark cave," the large model can generate three plot branches: the first plot branch describes "the torch flame being extinguished by a sudden draft"; the second plot branch describes "the torchlight illuminating the ancient totems on the stone wall"; and the third plot branch describes "the crackling sound of the burning torch startling the birds in the depths."

[0090] Regarding step S10154, for example, in some embodiments, during the plot branch generation process, the storyboard constraint parameters can be introduced as conditional inputs into the generation process of the large model and participate in the generation control and selection process of candidate plot segments.

[0091] Specifically, in the process of generating each plot branch, the candidate plot segments are subjected to constraint screening or generation probability modulation based on the storyboard constraint parameters, so that the generated plot branches meet the constraint conditions corresponding to the storyboard constraint parameters.

[0092] Regarding the scene switching frequency, the number of scene switching times per unit plot length can be limited based on the storyboard constraint parameters. During the generation of candidate plot segments, candidate plot segments that exceed the scene switching frequency constraint can be suppressed or downweighted.

[0093] Regarding the order of character appearances, the appearance sequence of different characters can be constrained based on the storyboard constraint parameters. During the generation of candidate plot segments, candidate plot segments that do not conform to the character appearance order constraints can be filtered or reordered.

[0094] For the type of camera movement, based on the constraints of the storyboard constraint parameters, the camera movement pattern in the plot description is constrained. During the generation of candidate plot segments, candidate plot segments that do not conform to the constraints of the camera movement type are suppressed or replaced.

[0095] In some embodiments, after the plot branches are generated, a consistency check can be performed on each plot branch based on the storyboard constraint parameters, and plot branches that do not meet the storyboard constraint parameters can be eliminated or corrected to obtain modulated plot branches.

[0096] For step S10155, for example, weights can be assigned to each modulated plot branch according to the excitement of the plot, logical coherence and visual adaptability, and the text fragments can be reorganized, trimmed or linearly fused according to the temporal logic to generate the target content text.

[0097] The sequence splicing mentioned here refers to connecting specific segments from multiple different plot branches end-to-end according to the development logic or chronological order of the storyline, constructing a complete narrative chain. This method is suitable for different branches that describe different stages or aspects of the story. For example, based on the logical labels of each plot branch (such as "beginning," "conflict point," and "twist point"), the local segments with the highest weight can be selected from different branches, and deduplication and smooth connection processing can be performed.

[0098] The linear combination processing refers to the reorganization, fusion, or feature overlay of the semantic content of multiple plot branches within the same spatiotemporal context, thereby generating a comprehensive text with higher information density and more precise expression. This approach is suitable for scenarios where multiple branches describe the same plot but with different emphases (e.g., one side emphasizes dialogue, the other action). For example, it can identify semantically overlapping parts in multiple branches and, based on weight allocation ratios, interweave and fuse keywords, rhetoric, or details from each branch, eliminating low-weight or redundant expressions.

[0099] Optionally, in some embodiments, the step of weighting the modulated plot branches and performing sequence concatenation or linear combination processing on the weighted plot branches to generate the target content text, i.e., step S10155, may further include the following steps:

[0100] Step A1: Perform sequence splicing or linear combination processing based on the weighted plot branches to generate candidate content text;

[0101] Step A2: Generate a corresponding posterior segmentation representation based on the candidate content text, and extract validation constraint parameters based on the posterior segmentation representation;

[0102] Step A3: Perform consistency verification on the candidate content text based on the verification constraint parameters;

[0103] Step A4: When the candidate content text does not meet the consistency check, the large model is triggered to correct and adjust the candidate content text until the consistency condition is met, and then the target content text is generated.

[0104] Specifically, for step A1, based on the weight allocation results of each plot branch, a calculation method of sequence splicing or linear combination can be used to aggregate the discrete branch fragments into a logically coherent complete text sequence, which serves as the basic input for subsequent verification.

[0105] Specifically, for step A2, deep semantic parsing technology can be used to identify spatial location words, action verbs, and environmental adjectives in the candidate content text. A preset visual deduction algorithm is then used to simulate the camera positioning and lighting arrangement within the text context, thereby generating the posterior storyboard representation. The posterior storyboard representation refers to the storyboard data obtained through reverse visual deduction based on the text content after the candidate content text is generated, used to characterize the visual presentation structure corresponding to the candidate content text. The posterior storyboard representation can use the same data structure as the initial storyboard prediction representation for consistency comparison.

[0106] Subsequently, the system can extract multi-dimensional verification values ​​from the posterior storyboard representation. Specifically, it can obtain a shooting difficulty value by calculating the frequency and amplitude of shot transitions, a scene coherence index by analyzing the overlap of visual elements between adjacent scenes, and a visual conflict intensity index by detecting whether the text description violates physical laws (such as contradictory lighting directions or spatial jumps). These values ​​collectively constitute the verification constraint parameters, which are a set of quantitative indicators characterizing the visual rationality of the candidate content text. Examples include shot transition frequency, scene coherence index, visual conflict intensity index, and shooting complexity index. The system can determine whether the candidate content text meets the preset visual generation conditions based on the above verification constraint parameters, and trigger text correction or regeneration accordingly.

[0107] For steps A3 and A4, in some examples, the semantic features of the candidate content text can be matched and calculated with the verification constraint parameters to determine whether the plot described in the text is inconsistent with physical laws, visual logic, or shooting cost overruns in the actual scene generation. When the verification result is unsuccessful, the unsuccessful verification indicators can be fed back to the large model, instructing it to make targeted modifications to the conflicting parts of the candidate content text, and repeating the generation and verification process until the text fully meets the visual constraint requirements.

[0108] It is not difficult to see that in this embodiment, by transforming user interest categories into specific target content themes and further utilizing sentiment classification models or semantic embedding models to extract sentiment feature vectors as sentiment elements, the subsequently generated script can maintain the consistency of themes while possessing a narrative tone that is closer to the user's psychological expectations and perceptions. By constructing a plot structure using preset plot templates or plot state transition rules, a clear logical framework can be provided for the text generation process, helping to avoid common problems in content creation such as narrative chaos or logical gaps. By filling in the initial content text based on the plot structure and further optimizing and shaping it to obtain the target content text, this progressive generation mechanism can improve the professionalism and controllability of script generation to a certain extent. This progressive processing from macro themes to micro emotions, and then to structured text, helps to improve the overall performance of micro-drama scripts in terms of narrative tension and logical completeness while ensuring generation efficiency.

[0109] Third Embodiment

[0110] The third embodiment relates to a personalized micro-drama scriptwriting and storyboard generation method based on a large model. The third embodiment is an improvement upon the first embodiment, specifically in that it provides a concrete implementation method for structuring the target content text to obtain a scene structure containing multiple scene units.

[0111] Specifically, the step S102, which involves performing structuring processing on the target content text to obtain a scene structure containing multiple scene units, may include:

[0112] Step S1021: Based on the target content text, calculate the plot complexity index by counting the number of plot points, the number of plot branches, and the number of semantic transitions.

[0113] Step S1022: Based on the comparison result between the plot complexity index and the preset threshold, determine the scene splitting granularity parameter;

[0114] Step S1023: Perform sequence labeling processing on the target content text based on the scene splitting granularity parameters to obtain scene labeling results;

[0115] Step S1024: Extract scene element information based on the scene annotation results and generate a scene structure containing the multiple scene units.

[0116] Specifically, for step S1021, natural language processing technology can be used to extract core events (plot points), causal logic branching points (plot branching degree), and contextual conflicts or spatiotemporal jump points (plot semantic shift degree) from the target content text. A pre-set weighted algorithm is then used to calculate a numerical index reflecting the difficulty of text creation. This plot complexity index is a comprehensive quantitative value used to measure the narrative density and logical shift frequency of the script.

[0117] The number of plot points refers to the number of key event nodes in the target content text that can drive the plot forward. These plot points typically consist of event-triggered actions, changes in the behavior of key characters, or plot-advancing events, such as a character discovering a crucial clue, a change in character relationships, or the occurrence of a conflict. The system can identify these event nodes using an event extraction model or a semantic parsing method based on predicate-argument structure and then count the number of plot points.

[0118] The number of plot branches refers to the number of different plot development paths that exist during the development of the story. A plot branch is formed when multiple possible development paths or parallel narrative threads appear in the text. For example, the same thread may point to different suspects, or the plot may unfold around two task lines simultaneously. The system can identify different narrative paths by recognizing causal relationships or conditional statement structures (such as "if...then...", "at the same time...", etc.) and count the number of plot branches accordingly.

[0119] The number of semantic transitions refers to the number of times the semantic tendency or plot logic changes significantly during the text's narrative process. Examples include a shift from a stable narrative to a conflict-ridden plot, from a safe state to a dangerous state, or from a misjudgment to the revelation of the truth. The system can identify semantic transition points by detecting changes in distance or emotional polarity between adjacent text segments in the semantic vector space and count the number of semantic transitions.

[0120] Optionally, in some embodiments, to automatically identify the number of plot points, plot branches, and semantic transitions, the following processing flow can be adopted: For automatic identification of event nodes, the system first performs sentence segmentation on the target content text and identifies the predicate-argument structure in the sentences based on dependency parsing. Sentences containing a "subject-action-object" structure and having event-triggered features are identified as candidate event nodes. For example, when an action predicate (such as "discover," "enter," "attack," "escape," etc.) appears in a sentence and the action changes the character's state or the plot state, the sentence is marked as an event node. In practical applications, the target content text can be input into a pre-trained event extraction model as a sequence of sentences or words. This model can use a sequence labeling architecture based on BERT or Transformer to predict a BIO tag for each token: the B-EVENT tag can represent the starting boundary of an event, the I-EVENT tag can represent subsequent tokens within the event, and the O tag can represent non-event tokens. Through forward inference of the model, the event tag sequence of the entire text sequence can be obtained. Then, by counting the occurrences of all B-EVENT tags, the number of event nodes in the target content text can be obtained.

[0121] For automatic identification of plot branches, natural language processing techniques can be used to segment and perform dependency parsing on the target text to identify conditional structures (such as "if...then...") or parallel narrative structures (such as "at the same time...", "on the other hand...") and map the related sentences in these structures to plot nodes. Then, a plot relationship graph can be constructed based on these nodes, where each node represents a plot node, and the directed edges between nodes represent the logical order of plot development or causal dependencies. When a node in the plot relationship graph has two or more successor nodes, it can be determined that a plot branch structure has occurred at that position, and then the number of all branch nodes in the entire text can be counted by traversing the graph structure.

[0122] For automatic identification of semantic inflection points, the system can segment the target text into multiple continuous text segments by sentence or a fixed-length sliding window. Then, a semantic encoding model (such as Sentence-BERT) can be used to embed and encode each text segment, obtaining the corresponding semantic feature vector. Next, the cosine distance between the semantic feature vectors of two adjacent text segments can be calculated. When this cosine distance is greater than a preset semantic change threshold, a semantic inflection point can be identified. Furthermore, in some cases, a sentiment polarity detection model can be combined to classify the sentiment of each text segment. For example, when the sentiment polarity of adjacent segments undergoes a significant change (e.g., from positive to negative, or from calm to conflict), it can also be identified as a semantic inflection point. By traversing the entire text sequence, the total number of semantic inflection points can be counted, and the corresponding semantic inflection degree value can be calculated for each inflection point based on the cosine distance or the magnitude of the sentiment polarity change.

[0123] Specifically, regarding step S1022, the scene segmentation granularity parameter is a preset control variable used to control the coarseness of scene segmentation. The system can compare the calculated plot complexity index with a preset tiered threshold. In some examples, if the index exceeds a high threshold, the granularity parameter is set to "fine granularity," and vice versa, ensuring that high-conflict plots can obtain a richer visual presentation space.

[0124] In some embodiments, the scene segmentation granularity parameter can be represented numerically, for example, defined as a scene segmentation threshold parameter G, with a value range of 0 to 1. When G is close to 0, it indicates that a coarse-grained segmentation strategy is adopted, and segmentation is only performed when there is a significant scene change; when G is close to 1, it indicates that a fine-grained segmentation strategy is adopted, and scene segmentation is performed at places with minor narrative changes.

[0125] In some embodiments, the scene splitting granularity parameter G can be determined based on the comparison result between the plot complexity index C and preset thresholds C1 and C2, for example: If C < C1, then G = 0.3; If C1 ≤ C < C2, then G = 0.6; If C ≥ C2, then G = 0.9.

[0126] C1 and C2 are preset complexity thresholds.

[0127] The system can utilize sequence labeling models (such as BiLSTM-CRF or Transformer sequence labeling models) to perform word-by-word or sentence-by-sentence sequence labeling on the target content text. For example, the tagging system employed may include: B-SCENE: Scene start position; I-SCENE: Text within the scene; E-SCENE: End position of the scene.

[0128] The sequence labeling model can be trained using training data that includes script scene boundary annotations, thereby learning to automatically identify scene boundaries. When the model detects a B-SCENE tag, it can determine the starting position of a new scene unit, thus completing the scene annotation of the text.

[0129] Specifically, for step S1023, the target content text can be scanned word by word or sentence by sentence according to the determined scene splitting granularity parameters to identify the spatiotemporal transition boundaries, character appearance points or action switching points that meet the granularity requirements, and the corresponding segmentation labels can be added to obtain the scene annotation results; the scene annotation results refer to the set of indexes marked in the text sequence to indicate the scene switching positions.

[0130] Specifically, for step S1024, the text can be divided into multiple physical segments based on the scene annotation results, and key information such as location, characters, props, and core actions (i.e., scene element information) can be automatically extracted from each segment. Finally, these encapsulated scene units are reassembled in chronological order to construct a complete structured data model.

[0131] Optionally, in some embodiments, the step of calculating the plot complexity index based on the target content text by counting the number of plot points, the number of plot branches, and the number of semantic transitions, i.e., step S1031, can be specifically implemented by the following formula:

[0132] ;

[0133] in, This represents the plot complexity index; Indicates the number of the aforementioned situation nodes. Indicates the maximum number of preset text fragment plot points; Indicates the degree of plot branching; This represents the total number of semantic turning points contained in the target content text; Indicates the first The degree of semantic transition at each semantic turning point; , , These are the preset first weight, second weight, and third weight, respectively.

[0134] Specifically, natural language processing techniques (such as dependency parsing or event graph construction) can be used to logically deconstruct the target content text, identifying key events, actions, and causal logic points. Sentiment classification models or semantic shift detection algorithms can be used to monitor the fluctuations of the text stream in sentiment polarity or semantic vector space in real time. By identifying anchor points where semantic tendencies change drastically, the total number of semantic turning points contained in the target content text can be determined. Furthermore, by calculating the cosine distance or semantic variance of adjacent text segments in the feature space, a turning point value representing the intensity of the plot reversal can be obtained.

[0135] Furthermore, the aforementioned dimensions are linearly weighted using preset first, second, and third weights. This quantification method transforms emotional narrative tension into rational parameter indicators, providing objective support for matching highly dynamic or impactful camera language models in subsequent step S1032.

[0136] For example, suppose the target content text contains =6 core plot points (preset maximum value) =10), plot branching degree =2 (representing the existence of two key plot developments), and identified =3 semantic transition points, each with a transition strength of respectively =0.8、 =0.5、 =0.9; The system can first calculate the arithmetic mean of the inflection point intensity, that is, execute... =0.73, which can then be combined with preset weighting coefficients (such as...) =0.4、 =0.3、 Linear weighting is performed on the plot complexity index C = 0.4·0.6 + 0.3·2 + 0.3·0.73 = 1.059.

[0137] It should be noted that this embodiment can also be an improvement based on the second embodiment.

[0138] It is not difficult to see that in this embodiment, the plot complexity index obtained by performing multi-dimensional quantitative analysis on the target content text can provide a scientific reference benchmark for subsequent automated processing, helping to objectively assess the narrative density of the script content. By comparing the complexity index with a threshold to determine the scene splitting granularity parameter, adaptive matching between processing logic and text features can be achieved, making the splitting process more flexible. Then, by performing sequence annotation based on the granularity parameter and extracting scene element information, long texts can be transformed into scene structures containing key information such as location and subject, which helps to improve the standardization of data extraction to a certain extent. This adaptive segmentation mechanism based on complexity analysis can optimize the logical connection of scene transitions and provide more granular structural support for the accurate generation of subsequent screen solutions.

[0139] Fourth embodiment

[0140] The fourth embodiment relates to a personalized micro-drama scriptwriting and storyboard generation method based on a large model. The fourth embodiment is an improvement upon the first embodiment, specifically in that it provides a concrete implementation method for generating scene schemes based on the aforementioned scene structure.

[0141] Specifically, the step S103, which generates a screen display scheme based on the scene structure, may include:

[0142] Step S1031: Based on the scene structure, extract visual description elements;

[0143] Step S1032: Based on the visual description elements, construct a structured storyboard representation;

[0144] Step S1033: Generate a scene scheme based on the storyboard representation.

[0145] Specifically, in step S1031, entity recognition and attribute extraction can be performed on each scene unit in the scene structure to identify the visual description elements. The visual description elements refer to a set of metadata extracted from the scene text that possesses direct visual features.

[0146] In some application examples, the visual description elements may include at least one of the following: scene objects, action relationships, and environmental attributes.

[0147] For a given scene, key subjects and core props can be identified. For example, key subjects may include a person's facial features or clothing style, while core props may include specific objects in the scene that have a narrative function.

[0148] Based on action relationships, the interaction state between the key subject and the core prop or the environmental space can be identified. For example, the action relationship may include the displacement trajectory of a character walking towards a bookshelf, or the interactive gestures of a character flipping through books.

[0149] By considering environmental attributes, the category of the environmental space and the overall lighting and shadow tone can be identified. For example, the environmental attributes may include the spatial positioning of indoor or outdoor scenes, as well as the atmosphere setting of dimness, brightness, or a specific color tendency.

[0150] Regarding step S1032, the scattered visual descriptive elements can be transformed into a standardized sequence of shooting parameters according to photogrammetry and film language logic, thereby constructing a structured storyboard representation. This structured storyboard representation can be used to characterize the parameter set automatically generated by matching the corresponding shot language model based on the degree of conflict and emotional intensity of the visual descriptive elements.

[0151] In some application examples, the structured storyboard representation may include at least one of the following: lens type, perspective parameters, and composition layout.

[0152] Depending on the shot type, the spatial range covered by the shot can be determined according to the narrative needs. For example, the shot type may include a panoramic view to show a grand environment, or a close-up to depict local details of a subject or emotional tension.

[0153] The perspective parameters determine the spatial relationship between the camera and the subject. For example, the perspective parameters may include a low-angle shot to convey a sense of majesty or height, or a high-angle shot to depict a panoramic view or a sense of oppression.

[0154] Regarding composition and layout, the geometric arrangement of visual elements within the image can be determined. For example, the composition and layout may include a rule-of-thirds composition that conforms to visual aesthetic balance, or a central composition used to strengthen the central position of the subject.

[0155] For step S1033, the parameters in the storyboard representation can be used as constraints and input into a preset image or video generation model. The diffusion sampling capability of the generation model is used to generate visual content that meets the parameter constraints, and a complete visual draft is formed according to the time sequence of scene units.

[0156] Optionally, in some embodiments, generating a scene scheme based on the storyboard representation, i.e., step S1033, may include:

[0157] Step S10331: Based on the storyboard representation, construct image generation control parameters;

[0158] Step S10332: Based on the image generation control parameters, call the image generation model to generate a frame sequence;

[0159] Step S10333: Based on the frame sequence, construct the picture scheme.

[0160] Regarding step S10331, in some examples, the storyboard instructions from the director's perspective can be translated into tensor constraints or guiding weights that the generation model can recognize, thereby generating image generation control parameters. These image generation control parameters are used to precisely regulate the generation boundaries of the image at the pixel generation level.

[0161] In some examples, the image generation control parameters include at least scene layout parameters and object constraint parameters.

[0162] Specifically, for scene layout parameters, the perspective and composition information in the storyboard representation can be transformed into spatial guidance weights. For example, by generating specific plugins through numerical encoding to control the intensity (such as Control Net weights for controlling composition), the generated model can be forced to place the visual subject at the golden ratio point of the frame, or a specific upward-facing perspective effect can be achieved through depth map constraints.

[0163] Specifically, for object constraint parameters, a text feature vector containing both positive and negative prompts can be generated by combining preset prompt word templates. For example, positive prompts are used to control the materials, lighting, and specific professional attire in the scene, while negative prompts are used to avoid visual flaws or elements that do not conform to the storyboard requirements, thereby guiding the model to generate image content that meets expectations at the probability distribution level.

[0164] Regarding step S10332, the image generation control parameters can be input into a pre-trained diffusion model or generative adversarial network. Utilizing the model's denoising sampling capability in the latent space, a series of image frames conforming to the parameter constraints are generated. Seed point fixing technology ensures visual consistency between frames, resulting in the frame sequence. The frame sequence refers to multiple consecutive static images generated by the generation model according to control logic, possessing temporal or logical correlation, forming the material basis of the image.

[0165] For step S10333, the generated frame sequence can be quality evaluated, invalid frames with artifacts or logical deviations can be removed, and the retained images can be processed by frame interpolation smoothing or color correction. Finally, the images are time-bound with the corresponding scene units to form a complete visual output scheme.

[0166] Optionally, in some embodiments, the step of constructing the image scheme based on the frame sequence, i.e., step S10333, may include:

[0167] Step B1: Using the frame sequence as the initial frame sequence, calculate the semantic similarity and visual continuity index based on the frame sequence;

[0168] Step B2: Calculate the cross-frame object consistency index based on the image generation control parameters;

[0169] Step B3: When at least one of the semantic similarity, the visual continuity index, and the object consistency index does not meet the preset conditions, the order of the initial frame sequence is adjusted or the image generation model is called again to generate the image until the preset conditions are met, and the target frame sequence is obtained.

[0170] Step B4: Construct the image scheme based on the target frame sequence.

[0171] Regarding step B1, the generated original image sequence can be defined as the initial frame sequence. The cosine distance between image features and text features is calculated using a cross-modal embedding model to derive semantic similarity. In practical applications, the cross-modal embedding model can employ a pre-trained model based on the Transformer architecture (such as the CLIP model), which achieves semantic comparison by mapping text and images to the same high-dimensional feature space.

[0172] Optical flow estimation or pixel-level residual analysis techniques can also be used to compare the motion vector deviations between adjacent frames in the initial frame sequence, thereby deriving a visual continuity index reflecting whether the image flickers or has breaks. Specifically, optical flow estimation can employ dense optical flow algorithms (such as the Farneback algorithm) or deep learning optical flow networks to extract the motion trajectories of pixels. Pixel-level residual analysis can quantify the degree of image change by calculating the differences in pixel attributes at the same coordinate positions between two adjacent frames.

[0173] The semantic similarity is used to measure the degree of fit between the screen content and the script text description, while the visual continuity index is used to evaluate the smoothness of the transition between adjacent frames in terms of lighting, composition and dynamics, in order to identify whether there is flickering or discontinuity in the screen.

[0174] For step B2, feature codes related to the subject can be extracted from the image generation control parameters, and these feature codes can be used as reference features to achieve refined monitoring of the stability of the subject in the generated image.

[0175] Specifically, the baseline features can be compared with the target objects identified in each frame of the initial frame sequence using keypoint matching and texture feature comparison. A consistency value reflecting the degree of object feature drift can be obtained through quantitative calculation, serving as a cross-frame object consistency index. In this embodiment, the cross-frame object consistency index is a quantitative parameter specifically used to monitor whether the same subject (such as a character's appearance or the appearance of a specific prop) maintains morphological stability in different frame views.

[0176] The method for determining the consistency value may include: identifying multiple semantic key points (such as facial feature coordinates or prop outline anchor points) between the baseline features and the target object in the current test frame; calculating the Euclidean distance between the corresponding key points; identifying non-logical morphological distortions through key point mapping offset calculation; obtaining local texture description values ​​of the baseline features and the target object using a preset feature extraction network; performing texture feature similarity measurement by calculating the cosine similarity between the two in the feature space; and thus quantifying the degree of retention of the subject in color, material, and fine texture. Further, the key point offset and texture similarity are weighted and fused to obtain a consistency value reflecting the degree of object feature drift. A higher value generally indicates that the visual features of the subject are more stable across frames, while a lower value indicates feature drift, thereby achieving quantitative monitoring of the morphological stability of the same subject in different frames.

[0177] In some embodiments, when no preset object template exists, the same object in different frames can be automatically identified using a cross-frame object tracking algorithm. Specifically, the system can use an object detection model (e.g., YOLO) to detect potential objects in each frame and obtain the corresponding object bounding boxes. Then, an object tracking algorithm (e.g., DeepSORT) can be used to match the detected objects in adjacent frames. This algorithm can achieve object association by combining object appearance feature vectors, spatial location features, and temporal continuity constraints. When objects in different frames meet the appearance feature similarity threshold and their motion trajectories are continuous, they can be determined to be the same object. After identifying the same object, the object's keypoint positions can be extracted based on a keypoint detection model (e.g., OpenPose), and the Euclidean distance between keypoints in different frames can be calculated to obtain a distortion index reflecting the degree of change in object morphology.

[0178] Specifically, the target appearance feature vector can be extracted by a convolutional neural network to extract the visual feature vector of the object region; the spatial location feature is specifically predicted by calculating the displacement of the center point of the target bounding box between consecutive frames to predict the object's motion trajectory; and the temporal continuity constraint can be predicted by using Kalman filtering to predict the target's position in the next frame.

[0179] Regarding step B3, in some cases, a compensation mechanism can be triggered based on the specific indicators that are not met: if it is only a temporal logic deviation, then frame order replacement is performed; if it involves object distortion or semantic deviation, then error features are fed back and seed points or prompt word weights are reconfigured to drive the image generation model to perform local redrawing or full reconstruction, and the process is iterated until all indicators exceed the preset threshold.

[0180] The target frame sequence refers to a set of images that, after logical correction or resampling, meet both visual and semantic constraints.

[0181] For step B4, the target frame sequence can be timestamped according to the script rhythm, and color space unification and super-resolution enhancement processing can be performed to ensure that the visual quality meets the resolution requirements of the micro-drama, thereby outputting the picture scheme.

[0182] It should be noted that this embodiment may also be an improvement based on the second embodiment and / or the third embodiment.

[0183] It is not difficult to see that in this embodiment, by extracting visual descriptive elements such as environment, subject, and props from the scene structure, detailed and concrete input information can be provided for subsequent image generation, which helps to reduce the ambiguity of content expression. By constructing a structured storyboard representation based on the elements, literary visual elements can be transformed into professional shooting logic containing parameters such as shots, perspectives, and compositions, which can optimize the professionalism of visual storytelling to a certain extent. By generating a screen scheme based on the storyboard representation, the parameterized storyboard instructions can be concretized into an image sequence using a preset generation model, which can help improve the efficiency of visual material production and, to a certain extent, ensure the consistency between the screen presentation and the director's intentions. This progressive processing from element extraction to logical modeling and then to visual presentation can enhance the level of control over visual content in the production process of micro-dramas.

[0184] Fifth embodiment

[0185] The fifth embodiment relates to a method for personalized micro-drama scriptwriting and storyboard generation based on a large model. The fifth embodiment is an improvement upon the first embodiment, specifically in that it provides a concrete implementation method for cross-modal consistency processing of the visual scheme and the content text.

[0186] Specifically, the cross-modal consistency processing of the screen scheme and the content text, i.e., step S104, may include:

[0187] Step S1041: Calculate the semantic correspondence and consistency score between the screen scheme and the target content text;

[0188] Step S1042: Compare the consistency score with a preset threshold;

[0189] Step S1044: When the consistency score is greater than the preset threshold, the screen scheme and the target content text are time-axis aligned and content-bound.

[0190] Step S1045: When the consistency score is less than or equal to the preset threshold, adjust the target content text or the screen scheme, and re-execute the consistency processing.

[0191] For step S1041, a multimodal encoder (such as the CLIP model) can be used to extract the semantic feature vector of the target content text and the visual feature vector of the key frame in the picture scheme respectively. By calculating the cosine similarity or Euclidean distance of the two vectors in the multidimensional feature space, a consistency score reflecting the degree of audio-visual alignment can be obtained.

[0192] The semantic correspondence refers to the mapping logic between the entities described in the text and the pixel features presented in the image, and the consistency score is a numerical indicator that quantifies the degree of this mapping fit.

[0193] Regarding step S1042, the real-time consistency score generated in step S1041 can be compared with a pre-stored benchmark threshold in the database. The logical truth value (True / False) of the comparison result determines whether to proceed to the compositing stage or revert to the correction stage. The preset threshold is a system-set minimum logical pass standard for determining whether the image is acceptable; this embodiment does not specifically limit its value.

[0194] Specifically, regarding step S1044, after confirming that consistency is achieved, the system can predict the duration based on the text's speaking speed (e.g., 200 words per minute), set time anchors, stretch or cut the corresponding visual scheme to the matching length, and use the text information as input for external subtitles or text-to-speech (TTS), performing a strong metadata-level association binding with the visual stream. It can be understood that timeline alignment and content binding is the process of precisely matching and encapsulating discrete visual frames with dialogue and narration in the script according to playback duration.

[0195] Regarding step S1045, based on missing features in the consistency calculation (such as detecting the absence of the element "rain"), correction logic can be automatically triggered: either fine-tuning the target content text to adapt to the existing image (such as changing "in the rain" to "cloudy"), or re-calling the generation model to optimize the image scheme, and then re-entering S1041 for cyclical verification. In other words, this step belongs to the system's feedback correction mechanism, used to handle abnormal situations where audio and video do not match, until the generated result meets the quality requirements.

[0196] Optionally, in some embodiments, when the consistency score is less than or equal to a preset threshold, the target content text or the screen scheme can be adjusted in at least one of the following ways: text element completion adjustment, text semantic replacement adjustment, screen regeneration adjustment, and local screen repair adjustment.

[0197] Specifically, regarding text element completion adjustments, when a key object entity in the text description is missing from the visual scheme, the system can insert supplementary descriptions into the target content text, such as adding environmental elements, action descriptions, or character characteristics, thereby ensuring semantic consistency between the text and the visual. Regarding text semantic replacement adjustments, when there are discrepancies between the generated visual elements and the text description in the visual scheme, the target content text can be adjusted through semantic replacement, for example, replacing "walking in the rain" with "walking on a cloudy day" to eliminate semantic conflicts. Regarding visual regeneration adjustments, when the text content is a key narrative element and cannot be modified, the system can regenerate the corresponding visual scheme based on the target content text. For example, it can re-call the image generation model and increase the weight of keywords containing key entities to generate visual content that matches the text description. Regarding local visual repair adjustments, when only some frames have semantic deviations, the corresponding areas can be repaired using local redrawing techniques based on a diffusion model, without regenerating the entire frame sequence. Through these adjustments, the system can perform closed-loop correction of the text content and the visual scheme, ensuring that the final generated result meets preset consistency conditions.

[0198] Optionally, in some embodiments, the step S1041, which calculates the semantic correspondence and consistency score between the image scheme and the target content text, can be implemented using the following formula:

[0199] ;

[0200] in, This represents the consistency score; This represents the overall semantic feature vector of the target content text. This represents the overall visual feature vector of the aforementioned image scheme; This indicates the number of scene object entities in the aforementioned visual scheme that match the target content text; This represents the total number of scene object entities contained in the target content text; and The preset adjustment coefficient, and satisfies + =1.

[0201] Specifically, a large-scale pre-trained cross-modal model (such as the CLIP model) can be used to extract the overall semantic feature vector of the target content text and the overall visual feature vector of the visual scheme, respectively. The cosine similarity between the two can be calculated to characterize the degree of semantic fit at the macro level. Simultaneously, the system can utilize natural language processing techniques, such as dependency parsing or named entity recognition (NER), to extract the total number of scene object entities in the text, and use object detection algorithms to identify the number of successfully matched scene object entities in the generated visual scheme. The object detection algorithm can be, but is not limited to, YOLO series algorithms or Faster R-CNN.

[0202] Furthermore, through a preset adjustment coefficient and The scores from the two dimensions are weighted and summed to obtain the final consistency score. This approach combines panoramic semantic understanding with microscopic entity detection, effectively avoiding the problem of generated images that, while matching the intended mood, lack key narrative elements. This provides a precise quantitative basis for subsequent automated correction or synthesis.

[0203] For example, consistency score The calculation combines two dimensions: global semantic fit and local entity matching rate. Assuming the target content text description is "a swordsman walks alone in a bamboo forest", the system can first extract the text vector. Visual vectors generated by the image scheme The overall semantic fit value was calculated to be 0.80 using the cosine similarity algorithm. Next, the text was identified as containing the two key scene entities: "swordsman" and "bamboo forest" (i.e.,...). =2), and these two entities were successfully matched in the generated image scheme (i.e. =2), resulting in an entity matching rate of 1.0; the system can adjust according to preset adjustment coefficients (such as...). =0.6, The consistency score is calculated by performing a weighted summation of (=0.4). =0.6×0.80+0.4×(2 / 2)=0.88.

[0204] It should be noted that this embodiment may also be an improvement based on any one or more of the second to fourth embodiments.

[0205] It is easy to see that in this embodiment, by calculating the semantic correspondence and consistency score between the visual scheme and the content text, a quantitative logical reference can be provided for evaluating the audio-visual fit. By comparing the score with a preset threshold, the system can automatically identify potential narrative deviations or visual conflicts. By aligning the timeline and binding the content when the score meets the standard, the synchronicity of the final output micro-drama in terms of audiovisual rhythm is ensured. Furthermore, by triggering adjustments and re-verification when the score does not meet the standard, a closed-loop correction mechanism is constructed, which can reduce the risk of audio-visual disconnect caused by the randomness of generation to a certain extent. This consistency processing based on quantitative scoring and dynamic feedback helps to improve the yield rate of micro-drama content and optimizes the user's audiovisual experience to a certain extent.

[0206] Sixth Embodiment

[0207] The sixth embodiment relates to a method for personalized micro-drama scriptwriting and storyboard generation based on a large model. The sixth embodiment is an improvement on the first embodiment, specifically in that the method is executed in a hybrid architecture that combines desktop and cloud server collaboration.

[0208] Specifically, the desktop client performs user interaction responses, local data caching, and lightweight inference tasks based on the user's interest categories; the cloud server performs large-scale text generation, image generation model inference, and cross-modal consistency processing computationally intensive tasks; the desktop client and the cloud server synchronize data and collaborate on tasks through a secure communication protocol.

[0209] It is understandable that existing technical solutions mainly focus on single architectures of pure cloud deployment or pure local deployment, lacking a holistic consideration of collaborative computing between the desktop and cloud servers. In a pure cloud architecture, user-created data needs to be continuously uploaded to a remote server, posing risks of data privacy leaks and network transmission latency. In a pure local architecture, the computing power of the terminal device is limited, making it difficult to support the inference needs of large language models, resulting in low content generation efficiency. Furthermore, existing solutions mostly adopt end-to-end automated generation models, lacking effective human-computer collaboration mechanisms, making it difficult for users to intervene and make fine-tuning adjustments in real time during the creation process. In this embodiment, by adopting a hybrid architecture that combines desktop and cloud server collaboration, sensitive information is processed locally to reduce the risk of privacy leaks, while high computing power demands are offloaded to the cloud to overcome terminal performance bottlenecks, thus balancing data security, response efficiency, and model performance.

[0210] In some examples, to achieve a human-computer collaborative creation mechanism, the method may further include: receiving real-time intervention commands from users regarding the scriptwriting process through a Kanban-style interactive interface. The Kanban interface presents each scene unit and its status information in card format, with each card corresponding to a scene unit. The card displays a summary of the scene unit, its generation status, and executable operation options. Users can adjust the scene order by dragging and dropping cards, click on cards to enter detailed editing mode, or perform operations such as deleting, copying, and merging cards.

[0211] The dashboard interface can adopt a two-way real-time synchronization mechanism: when the cloud server generates a new scene unit or updates an existing scene, the dashboard interface will automatically refresh and display; when the user performs an intervention operation through the dashboard interface, the system will immediately synchronize the operation command to the cloud server, triggering the corresponding model regeneration or structural adjustment process.

[0212] Furthermore, the real-time intervention commands may include at least one of the following: plot development adjustment commands, used to modify the plot development direction of a scene unit; character attribute modification commands, used to adjust the personality traits or behavioral patterns of characters in the scene; dialogue content editing commands, used to directly modify or rewrite the dialogue text in the scene; and scene addition / deletion commands, used to add scene cards or delete existing scene cards in the dashboard interface. After receiving the real-time intervention commands, the system can determine the scope of impact based on the command type. In some examples, if the command only affects a single scene unit, a partial regeneration is performed on that single scene unit; if the command involves the logical connection of multiple scene units, consistency checks and coordinated adjustments of the associated scenes are triggered. After the adjustment is completed, the system can render the updated scene structure to the dashboard interface in real time for users to continue reviewing and intervening.

[0213] At the video editing level, existing technologies typically abstract editing operations into black-box parameter configurations, making it impossible for users to intuitively preview and adjust the editing effects. Therefore, optionally, in some embodiments, before performing cross-modal consistency processing on the visual scheme and the target content text, the process may further include: receiving user editing operation instructions for the visual scheme through a visual editing editor. The visual editing editor adopts an interactive mode that links the timeline and preview window, and references the interactive design concepts of LTX-Desktop editing software, presenting professional editing functions to the user in an intuitive and visual manner.

[0214] For example, the visual editing editor generally includes the following core components: a timeline component, used to display the time arrangement of the frame sequence in the form of tracks, supporting multi-track overlay and hierarchical management; a preview window component, used to render the editing effect at the current timeline position in real time, supporting frame-by-frame preview and playback control; a media library component, used to manage available video materials, transition effects, and audio resources; and a properties panel component, used to display and edit detailed parameters of the selected material, including in point, out point, transparency, transformation attributes, etc. Optionally, the editing operation commands executed by the user through the visual editing editor include at least one of the following: cut commands, splicing commands, transition commands, speed adjustment commands, and special effects commands. The preview window uses real-time rendering technology, combined with graphics processor acceleration technology to ensure the smoothness of preview rendering, and adopts a progressive rendering strategy for complex special effects to achieve WYSIWYG editing feedback.

[0215] Regarding the underlying model support, existing large language model fine-tuning schemes suffer from slow training speed, high memory consumption, and unstable model quality after fine-tuning. In this embodiment, generating the target content text may include: generating initial content text based on the user's interest category using a Qwen3.5-27B or Qwen3.5-35B-A3B large language model fine-tuned with Unsloth. The Unsloth fine-tuning employs low-rank adaptation technology to generate editing-specific LoRA weights; these editing-specific LoRA weights are then loaded into the large language model, enabling it to generate script content that conforms to editing constraints.

[0216] Specifically, in practical applications, the dedicated LoRA weights for editing can be obtained as follows: A multimodal annotation dataset containing numerous script segments and corresponding editing parameters is constructed; the base model is fine-tuned using the Unsloth framework with low-rank adaptation. During fine-tuning, 4-bit quantization is used to compress the weights of the base model, gradient checkpointing is used to dynamically recalculate forward activation values ​​during backpropagation, the Flash Attention-2 algorithm is used to optimize attention calculation, and a dynamic batch processing strategy is employed. These techniques reduce memory usage during fine-tuning by more than 50% and increase fine-tuning speed by 1.5 to 12 times, thereby significantly reducing training costs while obtaining a high-quality dedicated model for editing.

[0217] Optionally, in some embodiments, IC-LoRA technology can also be introduced. During inference, the cloud server dynamically loads different LoRA weight combinations based on the current creation context, enabling the large language model to adaptively switch generation styles. For example, when the current scene is detected as a suspenseful reasoning type, the system can automatically load the corresponding suspenseful style LoRA weights; when switching to a romantic love scene, it can load the corresponding romantic style LoRA weights. By combining efficient fine-tuning with IC-LoRA technology, dynamic switching of the model's generation style and context adaptation can be achieved.

[0218] It should be noted that this embodiment may also be an improvement based on any one or more of the second to fifth embodiments.

[0219] It is not difficult to see that in the embodiments of this application, by combining the end-to-cloud collaboration, the refined human-computer interaction mechanism of the dashboard interface, and the efficient model fine-tuning and inference technology, the generation efficiency and large model processing capabilities are guaranteed, while giving users the ability to intervene and finely adjust in real time, which can greatly improve the flexibility and controllability of the micro-drama creation process.

[0220] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this application.

[0221] Seventh Embodiment

[0222] The seventh embodiment of this application relates to a personalized micro-drama scriptwriting and storyboard generation system based on a large model. The system includes: one or more processors; and a memory storing computer program instructions, which, when executed, cause the processors to perform the steps of the methods provided in any one or more of the above embodiments.

[0223] Figure 2An exemplary structural diagram of the system is disclosed. The system includes one or more processors 1101, a memory 1102, an input device 1103, and an output device 1104. The components are interconnected via a bus or other means (the diagram illustrates a bus connection). The processor 1101 can be used to execute instructions stored in the memory 1102 to control the overall operation of the system. The memory 1102 may include a program storage area and a data storage area, wherein the program storage area stores the operating system and applications required for at least one function; the data storage area stores data created according to system usage, etc. The memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as disk storage devices, flash memory devices, or other non-transitory solid-state storage devices. In some embodiments, the memory 1102 may also include storage resources located remotely to the processor and accessible via a network.

[0224] Input device 1103 can be used to receive input numerical or character information or user operation signals, such as a touch screen, keypad, mouse, trackpad, touchpad, indicator, one or more mouse buttons, trackball, joystick, etc. Output device 1104 may include display devices (such as liquid crystal displays, light-emitting diode displays, plasma displays, and optional touch screens), auxiliary lighting devices (such as LEDs), and haptic feedback devices (such as vibration motors), etc.

[0225] To facilitate user interaction, the system may be configured to include a display device (such as an LCD or CRT monitor) and input devices such as a keyboard and pointing devices (e.g., a mouse or touchpad). Feedback can be any form of sensory feedback (e.g., visual feedback, auditory feedback); input may also be received via voice, touch, or other means.

[0226] This application also relates to a computer-readable medium having a computer program / instructions stored thereon, which, when executed by a processor, implement the steps of the methods provided in any one or more of the above embodiments. The computer-readable medium may be a memory included in a system, or it may be a standalone storage medium not assembled into the device.

[0227] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium, a computer-readable storage medium, or a combination of both. Examples include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. Specific examples of storage media may include, but are not limited to, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory, optical fibers, portable CD-ROMs, optical storage devices, magnetic storage devices, etc., or any suitable combination thereof.

[0228] Computer-readable media may store one or more programs that can be used by or in conjunction with an instruction execution system. The media may be permanent or non-permanent, removable or non-removable, and may store information by any method or technology, including computer-readable instructions, data structures, program modules, or other data.

[0229] The computer program code used to implement the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Java, Smalltalk, and C++) and conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. The remote computer can be connected to the user's computer via any network (including a local area network or a wide area network) or can be connected to an external computer.

[0230] In the above embodiments, the functions can be implemented in whole or in part by software, hardware, firmware, or any combination thereof, for example, by using application-specific integrated circuits, general-purpose computers, or other similar hardware devices. In some embodiments, the software program of this application can be executed by a processor to implement the steps or functions; it can also be implemented by hardware, for example, as a circuit that works in conjunction with the processor to execute the steps or functions.

[0231] This application also provides a computer program product, including one or more computer programs / instructions, which, when executed by a processor, generate all or part of the processes or functions described in this application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one storage medium to another via wired (e.g., DSL) or wireless (e.g., wireless, microwave) means. The computer-readable storage medium may be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive).

[0232] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0233] The scope of this application is defined by the appended claims rather than the foregoing description, and is therefore intended to encompass all variations falling within the meaning and scope of equivalents of the claims. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device in software or hardware. Terms such as "first," "second," etc., are used only for distinguishing descriptions and do not indicate any particular order, nor should they be construed as indicating or implying relative importance.

[0234] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily made by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.

Claims

1. A method for personalized micro-drama scriptwriting and storyboard generation based on a large model, characterized in that, The method includes: Generate target content text based on user interest categories determined from historical user behavior data; The target content text is structured to obtain a scene structure containing multiple scene units, including: calculating a plot complexity index based on the target content text by counting the number of plot points, plot branches, and semantic transitions; determining scene splitting granularity parameters based on the comparison result of the plot complexity index and a preset threshold; performing sequence labeling processing on the target content text based on the scene splitting granularity parameters to obtain scene labeling results; and extracting scene element information based on the scene labeling results to generate a scene structure containing the multiple scene units; wherein, the number of plot points is the number of key event nodes in the target content text that can drive the plot development, the number of plot branches is the number of different plot paths that exist during the plot development, and the number of semantic transitions is the number of times the semantic tendency or plot logic changes significantly during the text narrative; Based on the scene structure, a visual scheme is generated; Cross-modal consistency processing is performed on the aforementioned visual scheme and the target content text, and the short drama content result is generated based on the processing result.

2. The method according to claim 1, characterized in that, The step of generating target content text based on user interest categories determined from user historical behavior data includes: Based on the user's interest categories, determine the target content theme; Based on the target content theme, sentiment feature vectors are extracted as sentiment elements using sentiment classification models or semantic embedding models. Based on the aforementioned emotional elements, a plot structure is constructed using preset plot templates or plot state transition rules; Initial content text is generated based on the aforementioned plot structure; The target content text is generated based on the initial content text.

3. The method according to claim 2, characterized in that, The process of generating the target content text based on the initial content text includes: Based on the semantic features of the initial content text, an initial storyboard prediction representation is generated; the initial storyboard prediction representation is used to characterize the preliminary storyboard description data deduced from the semantic information in the initial content text. Based on the initial storyboard prediction representation, storyboard constraint parameters are extracted; the storyboard constraint parameters are used to characterize the set of storyboard control parameters extracted from the initial storyboard prediction representation. Multiple plot branches are generated from the initial content text based on the large model; The storyboard constraint parameters are introduced during the plot branch generation process to constrain and control the plot branches, resulting in modulated plot branches. The modulated plot branches are weighted and then subjected to sequence concatenation or linear combination processing to generate the target content text.

4. The method according to claim 3, characterized in that, The step of performing sequence concatenation or linear combination processing on the weighted plot branches to generate the target content text includes: Based on the weighted plot branches, perform sequence splicing or linear combination processing to generate candidate content text; A corresponding posterior segmentation representation is generated based on the candidate content text, and verification constraint parameters are extracted based on the posterior segmentation representation; the posterior segmentation representation is used to characterize the visual presentation structure corresponding to the candidate content text; the verification constraint parameters are used to characterize a set of quantitative indicators for the visual rationality of the candidate content text. The candidate content text is subjected to consistency verification based on the verification constraint parameters. When the candidate content text does not meet the consistency check, the large model is triggered to correct and adjust the candidate content text until the consistency condition is met, and then the target content text is generated.

5. The method according to claim 1, characterized in that, The image generation scheme based on the scene structure includes: Based on the scene structure, extract visual description elements; Based on the aforementioned visual description elements, a structured storyboard representation is constructed; Based on the storyboard representation, a scene scheme is generated.

6. The method according to claim 5, characterized in that, The scheme for generating images based on the storyboard representation includes: Based on the storyboard representation, image generation control parameters are constructed; Based on the image generation control parameters, the image generation model is invoked to generate a frame sequence; Based on the frame sequence, the picture scheme is constructed.

7. The method according to claim 6, characterized in that, Constructing the image scheme based on the frame sequence includes: The frame sequence is used as the initial frame sequence, and semantic similarity and visual continuity indices are calculated based on the frame sequence. Calculate the cross-frame object consistency index based on the image generation control parameters; When at least one of the semantic similarity, the visual continuity index, and the object consistency index fails to meet the preset conditions, the initial frame sequence is adjusted in order or the image generation model is called again to generate the image until the preset conditions are met, and the target frame sequence is obtained. The image scheme is constructed based on the target frame sequence.

8. The method according to claim 1, characterized in that, The cross-modal consistency processing of the visual scheme and the content text includes: Calculate the semantic correspondence and consistency score between the image scheme and the target content text; The consistency score is compared with a preset threshold. When the consistency score is greater than the preset threshold, the screen scheme and the target content text are time-axis aligned and content-bound. When the consistency score is less than or equal to the preset threshold, the target content text or the screen scheme is adjusted, and the consistency processing is re-executed.

9. The method according to any one of claims 1 to 8, characterized in that, The method is executed in a hybrid architecture that combines desktop and cloud server collaboration; The desktop client performs user interaction responses, local data caching, and lightweight inference tasks based on the user's interest categories; The cloud server performs computationally intensive tasks such as large-scale text generation, image generation model inference, and cross-modal consistency processing for the large model. The desktop client and the cloud server synchronize data and collaborate on tasks through a secure communication protocol.

10. A personalized micro-drama scriptwriting and storyboard generation system based on a large model, characterized in that, include: Memory and processor; The memory is used to store computer programs; The processor is used to call and execute the computer program to implement a personalized micro-drama scriptwriting and storyboard generation method based on a large model as described in any one of claims 1 to 9.