Movie script generation method and system based on generative large model

By performing multi-dimensional analysis and evaluation on the film and television scripts generated by the generative large model, revision instructions are generated and the drafts are optimized. This solves the problem of the lack of shot-oriented thinking and cost awareness in existing film and television scripts, and ensures that the generated scripts meet the standards of the film and television industry.

CN121859858BActive Publication Date: 2026-05-29HANGZHOU WANXIANG TIANYING FILM & TELEVISION TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU WANXIANG TIANYING FILM & TELEVISION TECHNOLOGY CO LTD
Filing Date
2026-03-17
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing generative large models lack shot-based thinking and cost awareness when generating film and television scripts, making it difficult for the generated scripts to be implemented in the film and television industry and failing to effectively balance textual creativity with film and television applicability assessment.

Method used

By obtaining user prompts, the system generates initial scene drafts using a large creator model based on a large language model, performs multi-dimensional analysis and evaluation for filmability, generates filmability evaluation vectors and production warning reports, generates revision instructions and optimizes based on these evaluation results, and finally outputs a filmable script.

Benefits of technology

This ensures that the generated scripts meet the production standards of the film and television industry, solves the problem of difficulty in implementing generated content, and achieves a balance between textual creativity and filmability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121859858B_ABST
    Figure CN121859858B_ABST
Patent Text Reader

Abstract

The application discloses a film and television script generation method and system based on a generative large model, relates to the field of film and television script generation, and introduces a filmizable multidimensional analysis mechanism by taking the method as an intermediate product. By deconstructing an initial scene draft, quantifying visual potential and production risks of the initial scene draft, generating a filmizable evaluation vector and a production warning report containing objective evaluation indexes, and enabling the system to perceive shooting logic behind the text, specific revision instructions are generated based on the evaluation data and user constraints, and the model is guided to perform targeted constraint rewriting and optimization on the draft. The method forces the model to consider both narrativity and filmability during the creation process, ensures that the final output script meets the production standards of the film and television industry, and effectively solves the problem that generated content is difficult to land.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of film and television script generation, and more specifically, to a method and system for generating film and television scripts based on a generative large model. Background Technology

[0002] With the rapid development of artificial intelligence technology, generative large-scale models have demonstrated astonishing creativity and understanding in the field of natural language processing, especially showing great potential in assisting text creation. In the film and television industry, scriptwriting, as the source of the industry chain, directly determines the success or failure of film and television works through its efficiency and quality. Traditional scriptwriting often relies on the inspiration and lengthy polishing of human screenwriters, which is time-consuming and highly uncertain. This paper proposes a film and television script generation solution based on generative large-scale models. The aim is to leverage the language generation capabilities of large-scale models to assist screenwriters in quickly constructing story frameworks, generating scene descriptions and dialogues, thereby significantly improving creative efficiency, stimulating creative inspiration, and meeting the high-efficiency content output demands of the film and television industrial production.

[0003] However, directly applying existing general generative models to film and television script generation faces a core problem: a severe lack of cinematic perception and integration capabilities. Existing text generation technologies, whether based on the Transformer architecture or other sequence models, primarily rely on plain text corpora for training data. The models learn more about the script's textual format and narrative logic than the underlying visual elements and production logic. This results in scripts that, while grammatically and plot-wise coherent, often exhibit a strong literary orientation rather than a cinematic one. Specifically, the models lack awareness of physical simulation, scene mise-en-scène complexity, and filming costs. For example, a model might generate descriptions of a protagonist leaping from a high-rise to a speedboat—physically impractical or with extremely high budgets—but it cannot anticipate the feasibility, danger, and special effects requirements of such actions like a professional screenwriter or director. This lack of cinematic thinking and cost awareness results in outputs that are closer to novels than executable filming blueprints, making it extremely difficult to implement and costly to rewrite when entering the actual filming pipeline, hindering their integration into the industrialized film and television production process.

[0004] Therefore, there is an urgent need for a generation method that can take into account both textual creativity and film and television applicability to solve the above problems. Summary of the Invention

[0005] To address the aforementioned problems in the existing technology, according to one aspect of this application, a method for generating film and television scripts based on a generative large model is provided, comprising:

[0006] Get user prompts;

[0007] User prompts are used to input the creator's big model based on a big language model to obtain an initial scene draft;

[0008] Perform multi-dimensional analysis and evaluation of the initial scene drafts to obtain a filmability evaluation vector and a production early warning report;

[0009] Based on the filmability assessment vector, production early warning report, and user constraints, revision instructions are generated.

[0010] Based on the revision instructions, the initial scene draft is rewritten and optimized under constraints to obtain the revised scene draft;

[0011] The revised scene drafts are iteratively reviewed and finalized to obtain the final script suitable for film and television adaptation.

[0012] According to another aspect of this application, a film and television script generation system based on a generative large model is provided, comprising:

[0013] The user prompt acquisition module is used to acquire user prompts;

[0014] The initial scene draft generation module is used to generate an initial scene draft by taking user prompts and inputs based on a large language model of the creator's big model.

[0015] The filmability multi-dimensional analysis module is used to perform filmability multi-dimensional analysis and evaluation on the initial scene draft to obtain filmability evaluation vectors and production early warning reports.

[0016] The revision instruction generation module is used to generate revision instructions based on the filmability evaluation vector, production warning report, and user constraints;

[0017] The draft revision module is used to perform constrained rewriting and optimization of the initial scene draft based on revision instructions to obtain the revised scene draft;

[0018] The filmable script generation module is used to iteratively review and finally output the revised scene drafts to obtain the final filmable script.

[0019] Compared to existing technologies, this application provides a method and system for generating film and television scripts based on a generative large model, addressing the technical problems of existing generative models lacking shot-oriented thinking and neglecting production costs and physical feasibility. This solution does not directly output the initial text generated by the large model, but instead treats it as an intermediate product, introducing a multi-dimensional analysis mechanism for film and television applicability. By deconstructing the initial scene draft, its visual potential and production risks are quantified, generating a film and television applicability evaluation vector containing objective evaluation indicators and a production warning report, allowing the system to perceive the shooting logic behind the text. Furthermore, based on this evaluation data and user constraints, specific revision instructions are generated, guiding the model to perform targeted, constrained rewriting and optimization of the draft. This method forces the model to consider both narrative and filmability during the creation process, ensuring that the final output script meets the production standards of the film and television industry, effectively solving the problem of difficulty in implementing generated content. Attached Figure Description

[0020] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain the application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0021] Figure 1 This is a flowchart of a video script generation method based on a generative large model according to an embodiment of this application.

[0022] Figure 2 This is a schematic diagram of the data flow of a film and television script generation method based on a generative large model according to an embodiment of this application.

[0023] Figure 3 This is a flowchart of step S2 in the video script generation method based on a generative large model according to an embodiment of this application.

[0024] Figure 4 This is a block diagram of a film and television script generation system based on a generative large model according to an embodiment of this application. Detailed Implementation

[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0026] To address the technical problems of existing technologies, this application proposes a film and television script generation method based on a generative large model. Figure 1 This is a flowchart of a video script generation method based on a generative large model according to an embodiment of this application. Figure 2 This is a schematic diagram of the data flow in a film and television script generation method based on a generative large model according to an embodiment of this application. Figure 1 and Figure 2 As shown, the film and television script generation method based on a generative large model according to an embodiment of this application includes: S1, obtaining user prompts; S2, inputting the user prompts into a creator large model based on a large language model to obtain an initial scene draft; S3, performing multi-dimensional analysis and evaluation of the initial scene draft for film and television adaptability to obtain a film and television adaptability evaluation vector and a production warning report; S4, generating revision instructions based on the film and television adaptability evaluation vector, the production warning report, and user constraints; S5, performing constrained rewriting and optimization of the initial scene draft based on the revision instructions to obtain a revised scene draft; S6, performing iterative review and final output on the revised scene draft to obtain the final film and television adaptable script.

[0027] In step S1, user prompts are obtained. It should be understood that in the industrialized process of film and television content generation, scriptwriting is not a random process arising from nothing, but rather begins with a clear and specific creative intention. While large language models possess massive knowledge reserves and text generation capabilities, their essence is a general model based on probability prediction. Without initial constraints and directional guidance specific to a particular domain, the generated text is prone to divergence, leading to output content that deviates from the expected genre style or narrative core. To ensure that the subsequently generated film and television scripts accurately match the human creator's ideas in terms of theme expression, character setting, and scene atmosphere, an interface needs to be established at the beginning of the generation chain to capture and anchor these unstructured creative ideas. Therefore, step S1 aims to transform the abstract creative ideas in the user's mind into initial text data that can be processed by a computer, serving as the startup command and core context for the entire generative large model system. This provides a unique semantic benchmark and logical starting point for subsequent scene draft generation, multi-dimensional evaluation, and constrained rewriting.

[0028] In one possible implementation, step S1 is carried out as follows: The user prompt is an unstructured natural language text containing creative intent, the content of which covers the core elements of scriptwriting, such as the story's theme, genre (e.g., science fiction, suspense, romance), introductions of the main characters, specific objectives of the scene, and the expected emotional tone. Data acquisition is accomplished through the text input area of ​​the graphical user interface, which allows users to input long text strings.

[0029] The system captures data by listening for submission events at the input interface. Once the user completes and confirms the input, the system reads the text paragraph into memory and stores it as a user prompt variable. For example, a user might input: "Generate a cyberpunk-style action movie scene. The protagonist is an injured veteran detective named Kyle, hiding behind a neon billboard in the rain, holding a pistol with only one bullet left, trying to evade a villainous company's drone search. The scene needs to convey a tense and oppressive atmosphere." This text is the original user prompt. At this point, the data is unstructured, containing the ambiguity and polysemy of natural language. To ensure the prompt can be effectively processed by the subsequent large-scale creator model based on a large language model, basic format validation is first performed on the text, including removing invisible special characters at the beginning and end and unifying the character encoding to UTF-8 format to prevent garbled characters from affecting subsequent word vector conversion.

[0030] The core concept involved in this stage is the creator big model based on a large language model, which serves as the downstream processing object. Although inference has not yet been performed at this step, the specifications of the input data must conform to the architectural requirements of this model. The creator big model in this application adopts a Transformer-based deep neural network architecture, whose core components include a multi-head self-attention mechanism and a feedforward neural network. The model contains hundreds of millions to trillions of parameters, which are mainly composed of weight matrices and bias terms. The values ​​of the weights are obtained through self-supervised learning on massive text datasets (such as books, web pages, and script corpora) during a large-scale pre-training phase. Specifically, the model continuously updates the weights and biases by minimizing the cross-entropy loss function for predicting the next token, using the backpropagation algorithm and gradient descent optimizer (such as Adam), and finally converges to obtain the current parameter state. These parameters solidify the grammatical rules, world knowledge, and narrative logic of the language. The user prompts obtained in step S1 will subsequently be converted into a token sequence, which is then mapped into a high-dimensional vector as the input context of the model, stimulating specific neuronal activation patterns within the model.

[0031] To prevent excessively long input from exceeding the model's context window limit, the system presets a maximum length threshold when receiving user prompts. This threshold is determined based on the maximum number of tokens supported by the selected base model, such as 4096 or 8192 tokens. For example, if the preset threshold is 2000 characters, and the user's input exceeds this length, the system will prompt the user to shorten or segment the input. Continuing with the cyberpunk scenario example, the system will perform a preliminary length check on the text to confirm that its character count is within the valid range and does not contain illegal injection commands. The validated text stream is then marked as a valid user prompt.

[0032] In step S2, user prompts are input into a large-scale creator model based on a large language model to obtain an initial scene draft. Correspondingly, while large language models demonstrate powerful general capabilities in text generation, they often struggle to accurately capture the complex elements unique to film and television scripts, such as scene staging, character subtext, and visual atmosphere, when directly confronted with unstructured user prompts. Raw user input is often filled with the ambiguity and vagueness of natural language; if directly used as generation instructions, the model is prone to misinterpretation, resulting in scripts that are formatted incorrectly, deviate from the expected genre and style, and even ignore crucial dramatic conflicts. To transform a general generative large-scale model into a professional scriptwriting aid, the input prompts must undergo in-depth preprocessing and enhancement. Step S2 aims to build a bridge from vague user intent to precise model instructions. By performing structured parsing, creative expansion, and standardized encapsulation of the original prompts, the user's natural language is transformed into enhanced prompts containing clear semantic tags, rich related concepts, and strict format constraints. This guides the creator's large model to generate initial scene drafts that are rich in content, logically rigorous, and conform to industry standards, laying a solid content foundation for subsequent film and television adaptation evaluation.

[0033] In one possible implementation, Figure 3 This is a flowchart of step S2 in the film and television script generation method based on a generative large model according to an embodiment of this application. Figure 3 As shown, step S2, which inputs user prompts into the creator big model based on the big language model to obtain an initial scene draft, includes: S21, performing prompt word structuring and creative enhancement on the user prompts to obtain structured enhanced prompt words; S22, inputting the structured enhanced prompt words into the creator big model based on the big language model to obtain the initial scene draft.

[0034] Under the above implementation method, the implementation process of step S2 is as follows: In one possible implementation method, step S21, performing structuring and creative enhancement on the user prompt to obtain structured enhanced prompt words, includes: S211, performing structured parsing and entity linking on the user prompt to obtain a parsed element mapping table; S212, performing vector retrieval-based association concept enhancement on the parsed element mapping table to obtain an expanded keyword set; S213, performing template instantiation on the parsed element mapping table and the expanded keyword set to obtain structured enhanced prompt words.

[0035] S211, the first step in the entire processing flow, involves structured parsing and entity linking of the user prompt to obtain a parsed element mapping table. This step receives the user prompt output from the previous stage S1 as input. The text is first fed into a Named Entity Recognition (NER) model specifically fine-tuned for the film and television creation field. This NER model is based on the BERT architecture, which consists of a multi-layer bidirectional Transformer encoder. It is pre-trained on massive amounts of general text through a masked language model and a next-sentence prediction task, thus gaining profound language understanding capabilities. For this solution, the model is fine-tuned on a dedicated dataset with tens of thousands of labeled scripts, enabling it to recognize specific entity categories such as type, theme, character introduction, scene setting, scene objectives, and key props. The model's parameters (weights and biases) are continuously optimized during fine-tuning using a backpropagation algorithm to minimize the loss function between predicted and true labels. When the user prompt is input into the NER model, the model outputs a labeled sequence of terms, i.e., the annotated term sequence. For example, cyberpunk is tagged as [genre], action film as [genre], the wounded veteran detective named Kyle is tagged as [character description], the neon billboard behind a rainy night is tagged as [setting], the pistol with only one bullet left is tagged as [key prop], evading the villain's company's drone search is tagged as [scene target], and oppression and tension are tagged as [theme / atmosphere]. Next, the processing flow enters the entity grouping and aggregation stage, traversing the sequence and merging consecutive lexical units belonging to the same category. Then, the coreference resolution and attribute linking module intervenes. This module uses an attention-based relation extraction network to analyze the semantic relationships between entities. For example, it identifies that the protagonist refers to Kyle and links "wounded" and "veteran" as Kyle's attributes, forming a structured entity description. Finally, all identified and linked entities and their attributes are populated into a standard key-value pair data structure, generating a parsed element mapping table. This mapping table may contain the following key-value pairs: Key="Type", Value="Cyberpunk, Action"; Key="Character Introduction", Value="Kyle, a wounded veteran detective"; Key="Setting", Value="Rainy night, behind a neon billboard"; Key="Key Prop", Value="A pistol with only one bullet left"; Key="Scene Objective", Value="Evading the villainous company's drone search"; Key="Theme", Value="Oppressive, Tensive". Trivial words that were not recognized were filtered out.

[0036] Step S212 aims to address the issue of insufficient user input information by enriching scene details through the introduction of an external knowledge base. First, key creative seed fields are extracted from the parsed element mapping table obtained in the previous step, primarily selecting values ​​corresponding to type and theme, namely cyberpunk, action film, repression, and tension. These text values ​​are input into a pre-trained text encoder. This encoder uses the Sentence-BERT model, a variant of BERT that encodes sentences through a Siamese network structure, generating semantically meaningful sentence embeddings. Sentence-BERT is also based on the Transformer architecture, and after fine-tuning on the NLI (Natural Language Inference) dataset, its generated vectors better reflect the semantic similarity between texts. The input text is converted into a high-dimensional floating-point vector, called the query vector. Next, using query vectors Similarity searches are performed on a pre-built vector knowledge base of film and television themes and styles. This knowledge base is a pre-built vector database that stores a massive amount of film and television concepts and their corresponding vector representations. The knowledge base is constructed by collecting a large number of classic film scripts, reviews, and visual descriptions, extracting high-frequency visual images and style terms, such as neon rainy nights, holographic projection advertisements, cyborgs, steam and smoke, and reflective metallic pavements. These terms are then encoded into vectors using the same Sentence-BERT model. This data is then stored in vector retrieval engines such as Faiss. During retrieval, the query vector is calculated. With each concept vector in the knowledge base Cosine similarity measures the directional similarity between two vectors; a value closer to 1 indicates greater semantic similarity. To filter out high-quality related concepts, a preset relevance threshold is set. This threshold was determined experimentally on the validation set, aiming to balance recall and precision. For example, setting... =0.85. Filter out the text descriptions corresponding to all concept vectors with a cosine similarity > 0.85. If there are too many search results, sort them from highest to lowest similarity score and select the top K concepts. The value of K is determined based on experience or the user's needs for creative divergence, usually set between K=5 and 10. In this example, K=5 is set. The top 5 retrieved and filtered concepts are: holographic geisha advertisement, acid rain-corroded metal, flashing police lights, prosthetic repair shop, and dystopian slum. These text descriptions, after automatic deduplication, are added to a set data structure, forming an extended keyword set. This set enriches the original cyberpunk concept, providing specific visual materials for subsequent generation.

[0037] Step S213 is a crucial step in integrating the structured information and creative materials obtained in the previous two steps into the final executable instructions. The first step is template selection. Based on the core intent identified in the parsed element mapping table, the most suitable template is matched from the advanced cue word template library. Since the mapping table contains clear scene objectives and character introductions, and the intent is to generate a specific scene, the standard scene generation template is selected. This template library contains various predefined templates for different generation tasks, such as outline generation, character biographies, and dialogue optimization. Each template is carefully designed and includes instructions to guide the large model in generating output in a specific format. The selected template might look like this: "You are a professional screenwriter. Please create a scene for a film script based on the following information: [Genre and Style]: {GENRE_STYLE}; [Character Information]: {CHARACTER_PROFILE}; [Scene Environment]: {SCENE_SETTING}; [Key Props]: {KEY_PROPS}; [Creative Extension Elements]: {EXPANSION_KEYWORDS}; [Plot Objective]: {SCENE_GOAL}; [Emotional Atmosphere]: {MOOD}. Please use standard script format (scene title, action description, character dialogue), focusing on visual details and action scenes, integrating the above creative extension elements, and ensuring logical coherence. Next, data serialization will be performed. The Values ​​in the parsed element mapping table and all keywords in the extended keyword set will be converted into string format. For example, the extended keyword set {"Holographic Geisha Advertisement", "Acid Rain Corroded Metal", "Flashing Police Lights", "Cyborg Repair Shop", "Dystopia"} The string "}" is serialized into the string "Holographic Geisha Advertisement, Acid Rain Corroded Metal, Flashing Police Lights, Prosthetic Repair Shop, Dystopian Slum". Other fields in the mapping table are also converted to their corresponding strings. Placeholder replacement and instantiation are then performed. All predefined placeholders in the selected template (such as {Type and Style}, {Creative Extension Elements}, etc.) are iterated through and replaced with their corresponding serialized strings. For example, {GENRE_STYLE} is replaced with "Cyberpunk, Action Movie"; {CHARACTER_PROFILE} is replaced with "Kyle, Wounded Veteran Detective"; {EXPANSION_KEYWORDS} is replaced with "Holographic Geisha Advertisement, Acid Rain Corroded Metal, Flashing Police Lights, Prosthetic Repair Shop, Dystopian Slum". The final synthesized string is the structured enhanced prompt. This prompt not only includes the user's original requirements but also incorporates specific visual imagery and strict formatting instructions, such as: "You are a professional screenwriter... [Creative Extension Element]: Holographic Geisha Advertisement, Acid rain corrodes metals...

[0038] Step S22 is a complex automated process involving high-dimensional vector computation, probability space sampling, and rule constraint generation. The core execution entity of this step is the Creator Big Model, based on a large language model. This model employs a Decoder-only deep neural network structure based on the Transformer architecture, with a parameter scale typically in the hundreds of billions (e.g., 175B parameters). The specific architecture of the model consists of dozens of stacked decoder layers, each containing a multi-head self-attention mechanism sublayer and a feedforward neural network sublayer, integrated through residual connections and layer normalization. The weight matrix and bias vector in the model are obtained during the pre-training phase through a self-supervised learning task on a dataset containing massive amounts of books, scripts, and internet text—that is, maximizing the likelihood probability of predicting the next lexical unit given the context—and iteratively updated via backpropagation and the AdamW optimizer. These parameters implicitly encode the grammatical structure of human language, the formatting conventions of film and television scripts, and world knowledge about cyberpunk styles. In the specific implementation, context initialization is performed first. The structured enhanced prompts output from S21 are input into the tokenizer of the Creator Big Model. The token segmenter uses a byte-pair encoding algorithm to convert a string into a corresponding sequence of token IDs. For example, "cyberpunk" might be decomposed into [34521,223]. The entire prompt word is converted into an integer tensor, which serves as the initial input context sequence for the model. After receiving the sequence, the model maps each Token ID to a high-dimensional vector through an embedding layer, and overlays positional encoding to preserve sequence order information. Subsequently, the data flows through dozens of Transformer blocks in the model, undergoing layer-by-layer feature extraction and attention calculations. Finally, at the output layer, a linear transformation maps the hidden states to a vector the size of the vocabulary, i.e., the original logistic vector `logits`. At this point, the model enters the crucial stage of autoregressive loop generation. At each time step `t`, the model calculates the unnormalized log probability of the next lexical term on each word in the vocabulary, based on the current context (the initial prompt word and all previously generated content). Where V is the vocabulary size, for example, 50257. To ensure the generated script is both stunningly cyberpunk and logically sound, a specific creative decoding strategy is configured, starting with temperature sampling. Temperature sampling aims to adjust the smoothness of the probability distribution. A preset temperature coefficient T is set. This coefficient is set to a value greater than 1.0, for example, T=1.2. The rationale is that a higher temperature value can reduce the weight of high-probability words and increase the chance of selecting low-probability (but potentially more creative) words, thereby increasing the diversity of generated content. For each element in logits... The initial probability distribution is obtained by scaling and normalizing using the formula. : When T=1.2, the advantage of previously highly probable common words, such as "speak" and "see," is weakened, while the probability of previously less probable but more descriptive words, such as "scream" and "gaze," relatively increases. This makes the generated cyberpunk scene more likely to show holographic projection flickering rather than simply lights on. Next, to prevent excessive temperature from causing the model to generate completely incoherent gibberish or logically broken text, core sampling is applied. A cumulative probability threshold is set. This value is set to 0.95. A probability distribution of 0.95 means the system only trusts the set of head candidate words whose combined probability reaches 95%, while cutting off the remaining 5% of long-tail low-probability words, which are either grammatically incorrect or irrelevant. The specific implementation logic is: the temperature-adjusted probability distribution... Sort by probability value from highest to lowest. Accumulate the probability value of each word starting from the first word, until the cumulative probability is reached. At this point, the lexical units included in the cumulative range constitute the candidate core set. All lexical terms not in this set have their probabilities set to 0. For example, when describing Kyle's actions, the candidate set might include drawing a gun, dodging, and panting, while illogical terms like dancing and sleeping are excluded due to their extremely low probabilities. Subsequently, for The probabilities of the words in the sequence are renormalized so that their sum is 1 again, resulting in the final word probability distribution. Based on this final word probability distribution adjusted by the dual strategy, a word is randomly sampled. However, before it can be officially added to the generation sequence, it needs to pass a format compliance check. This is a lightweight state machine module. This module maintains the context state of the script format. For example, if the previously generated text is "Scene Title: Interior Abandoned Factory - Night", the state machine enters the action description state; if Kyle was generated, the state machine enters the expect dialogue or subtext state. As currently sampled It is a punctuation mark or action verb, but the state machine expects modifiers such as (VO) (voice-over) or (OS) (voice-over), or even dialogue content itself, and the state machine will check its reasonableness. If the generated... For tokens that severely violate script indentation or tagging rules (e.g., directly following a scene title after a character name), the checker will reject the token and trigger model resampling, or apply an extremely large negative bias (e.g., -10000) to the invalid token in the next round of generated logits, forcing the model to correct. Tokens that pass the compliance check... The tokens are appended to the end of the current text sequence and used as the input context for the next generation. This autoregressive process repeats continuously: calculate logits -> temperature scaling -> core sampling -> format checking -> appending. During this process, the model gradually constructs every detail of the scene: rainwater flowing down Kyle's prosthetic eye (action), Kyle: Damn, the scanner overheated (dialogue). The loop continues until the model generates a predefined special termination token, such as [END_SCENE], or the generated text length reaches a set maximum length limit, such as 2048 tokens. Finally, the entire generated token ID sequence is fed into a desegmenter, restoring the numerical sequence to human-readable natural language text. This text is the initial scene draft. For example, the output might be a standard script-formatted text: "Scene: Exterior, New Tokyo Lower District - Rainy Night\n Action: Acid rain corrodes the rusty tin roof... Kyle... clinging to...\n Kyle (whispering...) How much longer?...." This draft not only includes all the elements required by the structured cue words (cyberpunk, Kyle, tense atmosphere), but also fills in a large number of specific visual details and coherent action logic through creative decoding of the model, providing a solid material foundation for subsequent filmability evaluation and optimization.

[0039] In step S3, the initial scene draft undergoes a multi-dimensional analysis and evaluation for filmability to obtain a filmability evaluation vector and a production warning report. It is understandable that while the initial scene draft generated by the creator's large model achieves narrative coherence and format standardization at the textual level, from the perspective of actual film and television production, its essence remains a literary description untested by the physical world and economic costs. Generative large models lack an intrinsic perception of real-world shooting conditions, often producing fantastical content such as blowing up an entire city or large-scale scenes involving tens of thousands of people—exquisite in text but unrealistic due to budget constraints or technical limitations in actual filming. Directly incorporating such drafts into the production process will lead to repeated revisions, budget overruns, and even project shelving. To bridge the significant gap between textual creativity and practical execution, this application constructs a multi-dimensional analysis and evaluation system for filmability to deeply deconstruct the elements of the draft and quantify cost risks, outputting an objective filmability evaluation vector and a production warning report.

[0040] In one possible implementation, step S3, performing a multi-dimensional analysis and evaluation of the initial scene draft for filmability to obtain a filmability evaluation vector and a production warning report, includes: S31, deconstructing the production elements of the scene text in the initial scene draft to obtain an extracted element list; S32, quantifying the multi-dimensional cost of the production elements in the extracted element list to obtain an element score mapping table; S33, evaluating the visual potential of the text description in the initial scene draft to obtain a visual potential score; S34, aggregating the element score mapping table and the visual potential score to obtain a filmability evaluation vector; and S35, performing element-by-element analysis on the filmability evaluation vector based on a preset threshold to obtain a production warning report.

[0041] Under the above implementation method, the specific process of step S3 is as follows: Step S31 is a key step in converting natural language text into structured data that can be analyzed by a computer. This step receives the initial scene draft output from S2 as input. First, the text is fed into a script format parser. This is a rule engine built on regular expressions and finite state machines. It utilizes the strict format specifications of film and television scripts, such as scene titles usually being uppercase and occupying a single line, action descriptions being left-aligned, character names being centered, and dialogue being below character names with indentation, to accurately segment the continuous text stream into blocks with specific semantic functions. For the example draft, the parser will identify: "Scene: Exterior - Lower Tokyo - Rainy Night" as the scene title block; "Acid rain corrodes the rusty tin roof... Kyle... clinging to..." as the action description block; "Kyle" as the character name block; "(whispering...)" as the secondary language prompt block; and "How much longer?..." as the dialogue block. Subsequently, the core multi-task entity and event extraction stage begins. Each segmented action description block is fed into a multi-task joint extraction model fine-tuned using a film and television production dataset. This model employs a pre-trained language model based on BERT or RoBERTa as a shared encoder, with multiple task-specific classification heads connected to the backend. This architecture allows the model to complete multiple related tasks simultaneously in a single forward propagation, sharing underlying feature representations and improving processing efficiency and accuracy. The model's parameters are trained on a dataset containing a large amount of script annotation data (annotated with props, special effects, special skills, etc.) by jointly optimizing the loss functions of multiple tasks (such as the sum of cross-entropy losses). In the specific processing, the model executes three sub-tasks in parallel. The first is named entity recognition. The model scans the action description text and identifies noun entities that significantly impact production costs. For example, in describing a giant holographic geisha advertisement flashing in mid-air, the model identifies the holographic geisha advertisement as a prop, i.e., a special prop / set; in an armed drone bearing the 'Arasaka Corporation' logo, it identifies the armed drone as a prop. The second is event extraction. The model focuses on identifying verb phrases and their associated objects to capture events requiring special filming techniques. For example, from the image of a blue scanning laser beam cutting through the rain like a sword, the model identifies the laser beam cutting through the rain as a visual effect; from the image of acid rain corroding a rusty tin roof, it identifies acid rain corrosion as a physical / environmental effect. Finally, there's attribute classification. The model performs attribute judgments on specific entities. For example, analyzing the exterior scene in the scene title, the classifier labels it as an exterior scene, meaning it requires location shooting or the construction of a large outdoor set, rather than a relatively inexpensive interior set. Simultaneously, the model counts unique character names, such as Kyle, calculating an actor size of 1. All the results identified and classified by the model are compiled and output as a list—the list of extracted elements.This list is a structured list containing all quantifiable production elements in the scene, where each entry is a tuple (element type, element text). For the aforementioned cyberpunk scene, the output list of extracted elements might contain the following entries: [('Scene Type', 'Exterior'), ('Scene Name', 'New Tokyo Lower District'), ('Weather / Environment', 'Rainy Night'), ('Physical / Environmental Effects', 'Acid Rain Corrosion Effect'), ('Settings / Furnishings', 'Rusty Tin Roof'), ('Visual Effects Props', 'Holographic Geisha Advertisement'), ('Lighting Effects', 'Flickering Pink Light'), ('Special Props', 'Mechanical Eye'), ('Regular Props', 'Communicator'), ('Mechanical Props', 'Armed Drone'), ('Visual Effects', 'Blue Scanning Laser Beam'), ('Regular Props', 'Kinematic Pistol')]. This list clearly outlines the material conditions and technical means required to realize the scene, stripping away literary embellishments and restoring the production skeleton.

[0042] Step S32 is a high-precision calculation step that transforms qualitative semantic tags into quantitative multidimensional values. First, each element tuple (element type, element text) in the list is traversed. For each element, an exact match is first attempted in a pre-built film production cost database. This database is a large-scale relational database (e.g., built on PostgreSQL) specifically designed for the film industry. Its data comes from the cleaning and structuring of production budget tables for tens of thousands of films over the past decade, as well as the regular collection of standard price lists from major film equipment rental companies and special effects studios. The core table structure of the database includes fields such as standard name, thesaurus, budget score, time score, and complexity / risk score. The budget score represents the level of financial investment, the time score represents the preparation and filming time, and the complexity / risk score represents the technical difficulty or safety risk. Taking the entry ('common props', 'communicator') in the list as an example, the query engine uses 'communicator' as the primary key to retrieve data from the database. Since 'communicator' is a very common standard prop, an exact match exists in the database. The associated data is retrieved: Since the communicator only requires renting or purchasing off-the-shelf items, the cost is low, hence the budget score is 0.05; no special preparation is needed, hence the time score is 0.02; there is no usage risk, hence the complexity / risk score is 0.01. These pre-stored values ​​are standard reference values ​​derived from the average of historical big data. However, for highly customized or uniquely described non-standard elements in the list, such as ('visual effects props', 'holographic geisha advertisements') or ('physical / environmental effects', 'acid rain corrosion effects'), the database often struggles to directly match them. In this case, the processing flow automatically switches to the regression model estimation path. A set of specially trained cost prediction regression models is used here. Different sub-models are pre-trained for different element types (such as props, effects, and sets). Taking the SFX_Cost_Model for effects elements as an example, its specific architecture is based on a deep bidirectional Transformer encoder. For example, if RoBERTa-Base is used as the backbone network for feature extraction, its backend is connected to a multilayer perceptron (MLP) regression head to output continuous numerical predictions. The model's input processing first involves segmenting the text descriptions of special effects elements, such as acid rain corrosion effects, into words and adding special start tags. <s> and end marker< / s>Simultaneously, the type label corresponding to the element is assigned. For example, physical effects are mapped to type embedding vectors through a separate embedding layer. These type embedding vectors are concatenated or added to the initial label representation of the text sequence to incorporate prior type information. After the input sequence is processed by the 12-layer self-attention mechanism and feedforward network of the RoBERTa model, the hidden state vectors at the CLS positions, which are rich in contextual semantics, are extracted. These hidden state vectors are then fed into the MLP regression head, which contains two fully connected layers. The interlayer uses the Tanh activation function instead of ReLU to smooth the gradient and uses a Dropout layer to prevent overfitting. The final output layer has 3 neurons, corresponding to the prediction scores in the three dimensions of budget, time, and complexity. The processing of this regression model is as follows: First, the text description T is fed into the Transformer layer for deep feature extraction, outputting a context vector representation containing semantic information. The vector has 768 dimensions. Next, The data is fed into the MLP layer. The MLP layer consists of two fully connected layers and an activation function. The first layer compresses the high-dimensional features, and the second layer outputs three scalar values, corresponding to the original predicted scores in the three dimensions of budget score, time score, and complexity / risk score, respectively. Mathematically, this is expressed as: ,in, , It is a weight matrix. , It is a bias vector. It is the activation function of the linear rectifier unit. That is, the predicted score vector The weights and bias parameters in the model are obtained during the training phase by iteratively updating the model using the backpropagation algorithm by minimizing the mean squared error (MSE) loss function between the predicted values ​​and historical real project data (labeled cost, time, and risk level). For the holographic geisha advertisement, the model analyzes its semantic features, including holography (high-tech terminology), geisha (complex art design), and advertisement (requiring dynamic content production), and predicts a high value. For example, the original score output by the model is: high budget (requiring CG production), long time (requiring modeling and rendering), and medium complexity. Similarly, for the acid rain corrosion effect, the model identifies that acid rain and corrosion involve on-site chemical reactions or physical damage, with extremely high irreversible risks and set repair costs, thus predicting an extremely high complexity / risk score. After all elements in the list have obtained their original scores through database queries or model predictions, the model enters the score normalization and mapping table construction stage. Since the dimensions of the pre-stored values ​​in the database and the predicted values ​​in the model may be inconsistent, they must be uniformly mapped to the standard interval [0,1]. Using a maximum-minimum normalization method, all normalized data is ultimately assembled into a nested dictionary structure called an element score mapping table. For the aforementioned cyberpunk scenario, the specific content of this parsed element map table might be as follows: {'Communicator':{'Budget Score':0.05,'Time Score':0.02,'Complexity / Risk Score':0.01}; 'New Tokyo Lower District':{'Budget Score':0.90,'Time Score':0.85,'Complexity / Risk Score':0.60}, #Large-scale exterior construction required; 'Rainy Night':{'Budget Score':0.75,'Time Score':0.60,'Complexity / Risk Score':0.70}, #Involves water trucks Lighting equipment and night shooting risks; 'Acid rain corrosion effect': {'Budget score': 0.65, 'Time score': 0.50, 'Complexity / Risk score': 0.95}, #High-risk physical effects; 'Holographic geisha advertisement': {'Budget score': 0.80, 'Time score': 0.75, 'Complexity / Risk score': 0.40}, #High post-production costs; 'Armed drone': {'Budget score': 0.50, 'Time score': 0.30, 'Complexity / Risk score': 0.45} #Custom props and models required}. This mapping table precisely quantifies the production cost of each sub-element in the scene. For example, it clearly shows that although the budget score of the acid rain corrosion effect (0.65) is lower than the 0.90 of the "New Tokyo Lower District" construction, its risk score (0.95) is the highest.

[0043] In one possible implementation, step S33, which evaluates the visual potential of the initial scene draft by text description to obtain a visual potential score, includes: S331, text encoding the initial scene draft to obtain a scene text embedding; S332, calculating the concept similarity of the scene text embedding based on the target aesthetic concept vector set to obtain an initial visual potential score; and S333, performing score calibration based on a sigmoid function on the initial visual potential score to obtain a visual potential score.

[0044] S331, text encoding the initial scene draft to obtain the scene text embedding, is the first step in mapping natural language text to a semantic-visual joint space. This initial scene draft text is input into a pre-trained multimodal model's text encoder. Here, the CLIP model's Text Encoder component is chosen. The CLIP model is based on a Transformer architecture, containing a Text Encoder and an Image Encoder. The Text Encoder's specific architecture is a 12-layer or deeper Transformer, incorporating multi-head self-attention mechanisms and layer normalization. During processing, the initial scene draft is first tokenized, with start tokens [SOS] and end tokens [EOS] added, and then truncated or padded to a fixed length, such as 77 tokens. Subsequently, the ID sequences of these tokens are converted into vectors through the embedding layer and forward propagated in the Transformer layer. The model's parameters are trained on a large-scale image-text pair dataset (such as 400 million pairs of internet data) through contrastive learning. The training objective is to maximize the cosine similarity of matched image-text pairs in the feature space while minimizing the similarity of unmatched pairs. Therefore, the Text Encoder learns not only word meanings but also the visual features corresponding to the text. After encoding, the output vector at the end of the sequence (or global average pooling) is taken and mapped to the joint embedding space through a linear projection layer, resulting in a high-dimensional vector of dimension d, for example, 512 dimensions. This refers to scene text embedding. Mathematically, this vector represents the coordinates of the text in the multimodal semantic space, capturing the comprehensive features of visual concepts such as cyberpunk, neon lights, and oppressive atmosphere.

[0045] S332 is a process that uses a preset benchmark for measurement. First, a predefined set of target aesthetic concepts is loaded. This set is pre-constructed and obtained by selecting a group of keywords representing high visual value in the film and television industry, such as epic feel, visual spectacle, cinematic feel, cyberpunk aesthetics, film noir atmosphere, and action impact. These keywords are then input into the same CLIP Text Encoder to obtain the corresponding feature vectors. These vectors constitute the target aesthetic concept vector set. Next, the scene text embedding is calculated. With each concept vector in this set cosine similarity For example, in cyberpunk settings... Its similarity to cyberpunk aesthetics vectors might be as high as 0.85, to visual spectacle 0.78, while its similarity to pastoral scenery might be only 0.1. The maximum value among all calculated similarity scores is selected as the initial visual potential score. .Right now In this example, the maximum value is 0.85, which indicates that the scene has extremely high expressiveness in at least one important aesthetic dimension.

[0046] S333 aims to transform the raw similarity scores into standardized values ​​that are more intuitive for scoring. Since cosine similarity ranges from -1 to 1, and in CLIP's high-dimensional space, a similarity difference of 0.2 to 0.3 can represent a significant semantic gap, directly using the raw scores is not intuitive enough. Therefore, a sigmoid function is used for non-linear mapping and calibration. The calibration formula is as follows: .in, It is the final output visual potential score. This is the initial score obtained in the previous step, such as 0.85; It is a scaling factor that controls the steepness of the curve, for example, set to 10; This is the center offset, for example, set to 0.25, representing a baseline for average similarity. Specifically, the scaling factor and center offset are determined by performing logistic regression fitting or grid search optimization on a validation dataset containing human-annotated aesthetic scores to minimize the error between the model's output score and the human expert score. (As set...) =10, =0.25, if =0.85, then the exponent part ≈0.0025, ≈1 / (1.0025)≈0.997. This means the scene has extremely high visual potential, close to a perfect score of 1.0. If A score of only 0.25 (mediocre text) results in a score of 0.5. This transformation stretches and smoothly maps the original similarity to the [0,1] interval, enhancing the distinction between high and low scores. The final score of 0.997 is the visual potential score, a single scalar metric that quantifies the likelihood that the description of acid rain corroding a rusty tin roof could be transformed into a compelling cinematic image.

[0047] Step S34 is the core step in integrating fragmented analytical data into a global evaluation metric. First, scene-level cost scores are aggregated. In film production, the overall difficulty of a scene often depends on its most challenging weakest link (the "weakest link" effect). For example, even if other props are inexpensive, the budget for the entire scene can skyrocket if there's a large outdoor set that needs to be built. Therefore, instead of using the average, a maximum value pooling strategy is employed. All entries in the element score mapping table are iterated through, and the maximum value for each dimension is extracted. The final budget cost score is then calculated. =max(0.05,0.90,0.75,0.65,0.80,0.50)=0.90; Final time score =max(0.02,0.85,0.60,0.50,0.75,0.30)=0.85; Final complexity / risk score =max(0.01,0.60,0.70,0.95,0.40,0.45)=0.95. Next, the total score for filmability is calculated. User weight configuration is introduced. This is a parameter object set by users such as producers or directors during the project initialization phase to define the current project's emphasis on cost, time, risk, and visual effects. This configuration contains four floating-point weights: Visual weight, Budget weighting, Time weighting Risk weighting. If a user is setting up a project that prioritizes high quality but has a relatively generous budget, the configuration would be as follows: =1.0 Extreme emphasis on visual perception, =0.4 indicates a certain tolerance for budget constraints. =0.3 There is ample time. =0.5 indicates a need to control risk. Substituting the above values ​​into the weighting formula: The formula's design logic is as follows: visual score contributes positively, while cost, time, and risk all contribute negatively (deductions). Substituting specific values ​​into the calculation: =1.0×0.997-(0.4×0.90+0.3×0.85+0.5×0.95)=-0.093. This indicates that despite the excellent visual effects, the overall score becomes negative due to the high risk of acid rain corrosion and the high budget of the New Tokyo Lower District, suggesting that the scene is not feasible under the current configuration and requires significant optimization. Finally, a filmable evaluation vector is constructed, and the order of each dimension is predefined, i.e., its structure is [ , , , , The specific values ​​are: [-0.093, 0.997, 0.90, 0.85, 0.95]. This vector encapsulates all the key quantitative indicators of the scene and will serve as the direct basis for generating subsequent automatic revision instructions.

[0048] Step S35 is the process of translating the data into human-readable suggestions. First, preset warning thresholds are loaded. These thresholds are constants set according to risk control standards in the film and television industry, such as setting a budget warning threshold. =0.8, time warning threshold =0.8, risk warning threshold =0.9. Iterate through each element in the element score mapping table, checking if its scores in each dimension exceed the corresponding threshold. 1. Check 'New Tokyo Lower Ward': its final budget cost score > That is, 0.90 > 0.8. This triggers a high budget warning. Its final time score > That is, 0.85 > 0.8. This triggers a long-term periodic warning. 2. Check 'Acid Rain Corrosion Effect': its final risk score > That is, 0.95 > 0.9. Triggering a high-risk warning. 3. Check other elements (such as communicators): all scores are below the threshold, ignore. Based on the above checks, generate and format the output string "Production Warning Report". This report will highlight the culprit that caused the scene-level cost score to reach the maximum value. The generated report content is as follows: Production Warning Report: [Severe Warning - Extremely High Risk] Element 'Acid Rain Corrosion Effect' (Complexity / Risk Score: 0.95) exceeds the threshold of 0.9. This physical effect involves on-site chemical reaction control, which has uncontrollable safety hazards and the risk of reshoots. It is the project with the highest risk in the current scene. [High Budget Warning] Element 'New Tokyo Lower District' (Budget Score: 0.90) exceeds the threshold of 0.8. This outdoor set construction is huge and involves the 'Rainy Night' environment (Time Score: 0.85), which will lead to a serious extension of the shooting cycle. Recommendation: Consider changing 'Acid Rain Corrosion' to post-production VFX, or changing 'New Tokyo Lower District' to partial construction with green screen extension. This report not only points out the problem ( (The reason for the negative result) also pinpointed the specific word, such as "acid rain," that caused the problem, thus providing a precise target for generating specific revision strategies in subsequent steps, such as changing "acid rain" to "ordinary rainwater."

[0049] In step S4, revision instructions are generated based on the filmability assessment vector, production warning report, and user constraints. It should be understood that after the in-depth analysis in the preceding steps, while the initial scene draft may be excellent in content, quantitative assessment has shown that it may exceed the project's capacity in terms of production costs, risks, or timelines. While a simple assessment report points out problems, such as budget overruns or excessive risks, it is merely a diagnosis and does not directly provide specific action guidelines for modification. Without clear, data-driven modification instructions, the large model may again fall into blind creation during rewriting, either making excessive changes that lose the original ideas or making insufficient changes that fail to solve the core problems. To transform cold assessment data into meaningful and actionable natural language feedback for the large model, a bridge from diagnosis to prescription needs to be built. Step S4 is to deeply integrate the numerical deviations in the filmability assessment vector, the specific problems in the production warning report, and the hard constraints set by the user to generate a set of precise and targeted revision instructions. This guides the model to accurately revise the draft while retaining the core of the story, ensuring that the script generated in the next round can converge to the user's expected feasibility range.

[0050] In one possible implementation, step S4, based on the filmability evaluation vector, production warning report, and user constraints, generates revision instructions, including: S41, performing constraint mapping and difference quantization on the user constraints and filmability evaluation vector to obtain a difference vector; S42, performing root cause diagnosis and correction strategy generation on the difference vector and production warning report to obtain a correction strategy set; S43, instantiating the correction strategy set and the initial scene draft using natural language to obtain revision instructions.

[0051] Under the above implementation method, the specific process of step S4 is as follows: Step S41 is a calculation step that transforms the user's subjective natural language requirements into objective mathematical constraints and quantifies the degree of deviation of the current draft. First, obtain the user constraints. This is a JSON-formatted configuration object containing qualitative descriptions, generated by the producer or screenwriter by selecting preset options in the system interface. For example, for the current cyberpunk short film project, the user may have set the following constraints: {'Budget': 'Medium', 'Time': 'Urgent', 'Risk Tolerance': 'Low'}. These descriptive terms, such as medium and urgent, are intuitive but cannot be directly used for numerical calculations. Therefore, constraint threshold conversion needs to be performed. First, access a preset qualitative-quantitative constraint mapping table. This mapping table is a key-value database that stores the correspondence between common qualitative descriptive terms and specific normalized numerical thresholds. This table is built based on industry experience data. For example: for budget: 'Low' -> 0.6, 'Medium' -> 0.75, 'High' -> 0.9; for time: 'Sufficient' -> 0.8, 'Normal' -> 0.6, 'Urgent' -> 0.5; for risk: 'High' -> 0.8, 'Medium' -> 0.6, 'Low' -> 0.4. Each item in the user constraints is parsed and converted into a specific numerical threshold by looking up a mapping table. `map('Medium Budget')` returns the budget threshold. =0.75. map('Time is tight') returns the time threshold. =0.50. map('Low risk tolerance') returns the risk threshold. =0.40. This provides a visually verifiable evaluation vector for each key dimension ( , , This generated the corresponding set of hard constraint thresholds. Next, difference calculation and positive filtering are performed. This involves iterating through each actual score in the video-capable evaluation vector. Compare it with the corresponding threshold A comparison is then made. To accurately capture the degree of exceeding the limit, the formula for the modified linear unit (ReLU) logic is used for calculation: This formula ensures that a positive difference is only generated when the actual score exceeds the threshold (i.e., violates the constraint); if the score is below or equal to the threshold (i.e., meets the constraint), the difference is 0, indicating that no correction is needed. Substituting the aforementioned specific values ​​into the calculation: Budget Difference Actual budget breakdown The threshold is 0.90. It is 0.75. =max(0,0.90-0.75)=0.15. This means the budget exceeded the limit by 0.15 units and needs to be reduced accordingly. Time difference. Actual time The threshold is 0.85. It is 0.50. =max(0,0.85-0.50)=0.35. This indicates that the shooting schedule was severely exceeded because the user had tight deadlines, and the scenes involved time-consuming outdoor shots and rainy nights. The difference value is as high as 0.35, which is a bottleneck that urgently needs to be addressed. Risk Difference Actual risk score The threshold is 0.95. It is 0.40. =max(0,0.95-0.40)=0.55. This indicates that the risk is extremely high, the user requires low risk, but physical effects such as acid rain corrosion bring extremely high uncontrollability, resulting in the largest difference value. Finally, the output is vectorized. All calculated difference values These are combined into a structured variance vector. This vector corresponds to the evaluation vector dimension and is specifically used to represent the degree to which each score exceeds the constraints. The output variance vector is: {0.15, 0.35, 0.55}. This variance vector precisely quantifies the problems in the current draft: it shows that risk is the primary concern (with the largest variance), followed by time constraints, and finally budget constraints.

[0052] Step S42 is a decision-making process combining logical reasoning and knowledge retrieval. First, key issues are identified. The input difference vector is analyzed, comparing the magnitude of differences across various dimensions to determine the primary contradiction facing the current scenario. In the example data, the risk difference value is 0.55, significantly higher than the time difference (0.35) and budget difference (0.15). Therefore, exceeding the risk limit is identified as the most serious problem and must be addressed first. This means that subsequent strategy generation will focus on reducing production difficulty and safety hazards. Next, problem attribution is performed. To find the specific culprit causing the risk exceeding the limit, Natural Language Processing (NLP) technology is used to analyze the production warning report. Since the report is structured text, it is extracted using simple keyword matching or regular expressions. For the identified risk exceeding the limit problem, warning entries containing the keyword "risk" or "Risk" are searched in the report. The element 'Acid Rain Corrosion Effect' (complexity / risk score: 0.95) is located in: [Severe Warning - Extremely High Risk]... From this, the root cause entity text causing the problem is successfully extracted as "Acid Rain Corrosion Effect," and its corresponding entity type is identified as a physical / environmental effect. This step completes the process of locking in data from abstract to concrete objects. Next, a core strategy knowledge base query is performed. A pre-built problem-strategy mapping knowledge base is loaded. This knowledge base is an expert system storing a large number of correction rules in the film and television production field, existing in the form of lookup tables or graphs. Its construction method is based on summarizing the common handling methods used by senior producers and screenwriters when facing budget cuts or technical limitations. For example, the knowledge base contains the following mapping rules: Rule 1: (High budget, props) -> ['Replace with cheaper alternatives', 'Delete entity']; Rule 2: (High risk, physical / environmental effects) -> ['Switch to post-production effects', 'Simplify effects description', 'Use sound cues']; Rule 3: (High time cost, scene / exterior) -> ['Change to interior scenes', 'Reduce scene length']. The currently identified problem type and root entity type are used as the joint query key (high risk, physical / environmental effects) to search the knowledge base. The query matches rule 2, returning the strategy list ['Switch to post-production effects', 'Simplify effects description', 'Use sound cues']. Finally, the strategy is instantiated and output. Based on the magnitude of the difference value or the preset strategy priority, the most suitable strategy is selected from the returned list. Since the risk difference is as high as 0.55, a significant difference, simple simplification may not be sufficient to solve the problem, and sound cues may weaken the visual impact. Therefore, a more moderate and effective compromise is chosen: converting it to a post-production effect. This strategy is bound to the root entity text acid rain corrosion effect, forming a structured correction strategy object: {'Strategy Code': 'Convert to Post-Production Effects', 'Strategy Description': 'Convert physical effects to post-production visual effects to reduce on-site risks', 'Target Entity': 'Acid Rain Corrosion Effect'}.In addition, a quick diagnostic round may be performed for minor issues (time exceeding the limit, time difference = 0.35). The root cause entity, New Tokyo Lower District (exterior view), is located. The knowledge base is queried (high time cost, scene / exterior view), and the strategy is changed to an interior view. The generated strategy object is: {'Strategy Code':'Change to Interior View','Strategy Description':'Change exterior view to interior view to reduce the impact of construction and weather','Target Entity':'New Tokyo Lower District'}. All generated strategy objects are added to a set, forming the final set of corrective strategies.

[0053] Step S43 is a Natural Language Generation (NLG) step that utilizes a large language model for role-playing and instruction generation. First, a prompt is constructed. A complex and information-rich prompt needs to be dynamically assembled. This prompt aims to activate specific capabilities of the large model, enabling it to simulate an experienced director or production consultant. The prompt is composed of the following predefined template modules: 1. Role Setting Module: Instills a specific role identity and task objective into the model. Insert text: "You are an experienced film director, skilled at achieving optimal artistic effects within budget constraints. Your task is to provide the screenwriter with clear and actionable revisions." This sets the tone (authoritative, constructive) and perspective (director's perspective) of the output. 2. Contextual Information Module: To make the generated instructions more targeted, a summary text of the "initial scene draft" is embedded within. For example: "Current script scene summary: A cyberpunk-style rainy night. The protagonist, Kyle, is in an exterior scene of 'New Tokyo Lower District,' dodging environmental damage with 'acid rain corrosion effects,' with holographic advertisements and drones in the background."

[0054] 3. Task Module: This is the core part. It parses each strategy object in the "Correction Strategy Set" and converts it into an instruction-style natural language description. For the strategy {'Strategy Code':'Convert to Post-Production Visual Effects', 'Target Entity':'Acid Rain Corrosion Effect'}, it translates to: "Task 1: For 'Acid Rain Corrosion Effect', execute the strategy 'Convert physical effects to post-production visual effects'. Please suggest how to achieve this through simple on-site shooting combined with post-production compositing, avoiding dangerous real-life chemical corrosion filming." For the strategy {'Strategy Code':'Change to Indoor Scene', 'Target Entity':'New Tokyo Lower District'}, it translates to: "Task 2: For the exterior setting of 'New Tokyo Lower District', execute the strategy 'Change to Indoor Scene'. Please suggest moving the scene to a specific indoor space (such as an abandoned warehouse or subway station) to avoid the cost of building large exterior scenes and controlling the weather." 4. Output Format Requirements Module: Clearly defines the constraints on the output format. The inserted text reads: "Based on the above tasks, generate a specific revision instruction. The language style should be inspiring; don't just list the operations, but explain how this modification will maintain the original tension." After concatenation, a complete long text prompt is formed, ready to be input into the downstream model. Next, large-scale model inference generation is performed. This step calls a large language model similar to S2, but with fine-tuned instructions, referred to here as the director's large model. This model (such as GPT-4 or Claude 3 architectures) receives the constructed prompt. Internally, the model is based on a Transformer architecture, its core consisting of stacked Transformer decoder blocks, each containing a multi-head self-attention layer and a positional feedforward network layer, using residual connections and layer normalization for deep signal transmission. This architecture leverages its powerful self-attention mechanism to understand the logical relationships between different parts of the prompt, i.e., how to maintain the director's tone required by the character setting while satisfying the hard constraints of Task 1 and Task 2, combined with the contextual information of the plot. The model begins to generate text word by word. It's not just a simple translation strategy, but involves inference and creative filling. Regarding the acid rain issue, the model might deduce: since we can't actually throw acid, let's use light and sound to represent it. Regarding the transition from exterior to interior scenes, the model might associate cyberpunk interiors with narrow, dirty spaces, which perfectly matches Kyle's oppressive feeling of being hunted. Finally, the model outputs a coherent natural language text, i.e., a revision instruction. In this example, the generated content might be as follows: Revision Instruction: Dear Screenwriter, to optimize production feasibility and maintain the film's tension, we need to make the following adjustments to the scene: 1. Scene Replacement: Move the scene from the grand exterior of 'New Tokyo Lower District' to the maintenance room (interior) of an abandoned subway station. This not only avoids the high costs of building exterior scenes and artificial rain, but the enclosed space further enhances the oppressive atmosphere of Kyle's 'nowhere to escape'. 2. Special Effects Adjustment: Regarding the 'acid rain corrosion effect,' please remove the physical corrosion depiction on site.The revised script should focus on sound design (the sizzling sound of corrosion) and post-production VFX hints (the warning light on Kyle's prosthetic body turns red, and the UI displays 'Armor Damaged'). Instead of depicting rain directly corroding the roof, use Kyle's reaction to indirectly highlight the lethality of the acid rain. Please rewrite the script based on these two points. This output not only clearly conveys the revised intentions (interior scenes, post-production), but also provides specific creative solutions (subway station, indirect imagery).

[0055] In step S5, based on the revision instructions, the initial scene draft is rewritten and optimized under constraints to obtain the revised scene draft. Correspondingly, through the evaluation and strategy generation in the preceding steps, the problems existing in the script have been identified and specific modification suggestions have been given, such as changing the high-risk physical acid rain to post-production special effects and converting high-cost exterior scenes to interior scenes. However, these revision instructions are merely guiding principles; the real challenge lies in how to implement these abstract logical instructions into the specific script text without compromising the coherence and artistic appeal of the original story. Simply deleting or replacing certain words may lead to contextual breaks, illogical character behavior, or failure to create atmosphere. For example, after deleting the rainy night exterior scene, if the characters' actions (such as wiping away rain) and dialogue (such as complaining about the weather) are not adjusted accordingly, the script will become fragmented. To achieve a smooth transition from instructions to final text, this application guides the large model to creatively execute the modification requirements based on a deep understanding of the original text's intent through a dedicated rewriting generation process.

[0056] In one possible implementation, step S5, based on the revision instructions, involves constrained rewriting and optimization of the initial scene draft to obtain a revised scene draft, including: S51, structurally assembling rewriting prompts for the initial scene draft and revision instructions to obtain rewriting prompts; S52, inputting the rewriting prompts into a rewriting generator based on a large language model to obtain candidate revisions; and S53, verifying the dramatic core retention of the candidate revisions and the initial scene draft to obtain the revised scene draft.

[0057] Under the above implementation method, the specific process of step S5 is as follows: Step S51 is a key preprocessing step that logically integrates and formats heterogeneous text information from multiple sources. First, instruction intent parsing and keyword extraction are performed. Using Natural Language Processing (NLP) technology, a shallow semantic analysis is performed on the revision instruction generated in S4. The instruction text is scanned to identify and extract key action verbs, such as substitution, cancellation, and change, as well as their target entities, such as Neo-Tokyo Lower District, acid rain corrosion effect, and abandoned subway station. These extracted action-entity pairs are temporarily stored. Although they do not directly constitute the main body of the prompt, they will be used to generate emphasis markers in the prompt or for subsequent verification steps. For example, ('action': 'substitution', 'source entity': 'Neo-Tokyo Lower District', 'target entity': 'abandoned subway station') are extracted. Next, the core multi-segment template filling is performed. A guided rewriting template designed specifically for text editing tasks is enabled. This template is designed as a long text structure with clear partitions to minimize the illusion of a large model and focus on the editing task. The template contains four strictly defined sections, into which the system fills in the corresponding content in sequence: 1. [Role and Task] Section: This is the header of the prompt, used to define the identity and behavioral guidelines of the large model. Insert the preset text: "You are a senior screenwriter. Your task is to rewrite the provided original scene draft according to the following director's modification instructions. The core objective is to maintain or enhance the original dramatic tension while reducing production costs. Please ensure that the modified scene is logically consistent and does not omit key plot information." This text sets the screenwriter's role for the model and clarifies the optimization goal of cost reduction and quality assurance. 2. [Original Text] Section: The complete initial scene draft output by S2 is seamlessly inserted into this area. To distinguish them, separators such as "" or [ORIGINAL_SCRIPT_START] will be added before and after the text. The inserted content is: "Scene: Exterior - New Tokyo Lower District - Rainy Night\n Action: Acid rain corrodes the rusty tin roof... (complete text omitted here)". This provides the model with the story foundation and narrative starting point that it must follow. 3. [Modification Instructions] section: Insert the complete revision instructions output by S4 into this area. Also enclosed in separators. The inserted content is: "Revision Instructions: 1. Scene Replacement: Move the scene from the grand exterior of 'New Tokyo Lower District' to 'the maintenance room (interior) of an abandoned subway station'... 2. Effects Adjustment: Remove the on-site physical corrosion description and replace it with side rendering...". This part is the core constraint that drives the model to change the generation trajectory. 4. [Output Requirements] section: This is the end of the prompt words, used to standardize the output format. Inserted text: "Please output the rewritten complete scene script. The output should strictly follow the standard script format (scene title, action description, character name, dialogue). Please output the script content directly, without including explanatory text." Finally, the prompt words are synthesized.The four parts filled in above are then pieced together in sequence to form the final rewrite prompt. This is a structured, multi-part composite text object. Its structure is roughly as follows: [Characters and Tasks]...[Original Text]...[Modification Instructions]...[Output Requirements]... This rewrite prompt integrates the context of the original text, the logical instructions for modification, and the formatting specifications. This meticulous assembly process ensures that the model does not write instructions as dialogue into the script, nor does it forget the character's name when modifying the scene, thus guaranteeing the accuracy and usability of the rewrite.

[0058] Step S52 is an advanced generation control step combining deep learning inference and real-time probabilistic intervention. The execution entity is a rewriting generator based on a large language model. This generator is also based on the Transformer Decoder architecture, sharing the same parameter weights (such as GPT-4 or PaLM 2, models with hundreds of billions of parameters) as the creator's large model in previous steps, but in this step, its operating mode is switched to controlled generation mode. First, context loading is performed. The long text rewriting prompts assembled in S51 are input into the embedding layer of the rewriting generator. The model converts them into a token sequence and performs positional encoding, serving as the initial context for generation. Next, the crucial dynamic logistic bias step begins. The model starts autoregressive generation per token. At each generation time step t, the model's last linear layer calculates the original log probability vectors (i.e., the original logits) for all V tokens in the vocabulary (e.g., 50257 tokens). At this point, before these logits are fed into Softmax normalization, a pre-defined constraint application module forcibly intervenes. This module modifies logits in real time based on the action-entity pair information parsed from S51. Specifically, it first constructs a muted set based on the action-entity pairs. For example, the instruction requires "Scene replacement: change 'New Tokyo Lower District' (exterior) to 'Abandoned Subway Station' (interior)", and remove physical acid rain. Add tokens and their synonyms related to concepts such as New Tokyo Lower District, exterior, acid rain, and corrosion to the muted list. For example, Token ID 3452 corresponds to the exterior, and Token ID 8912 corresponds to acid rain. Apply the following bias formula to adjust the original logits:

[0059]

[0060] in, It is the i-th word element in the large model vocabulary. It is the original logistic value. It is the adjusted logistic value. This is a preset negative bias value, empirically set to a sufficiently small number (such as -100.0 or negative infinity) to ensure that the generation probability of the corresponding word unit approaches zero after Softmax normalization. To ensure that these words will absolutely not be selected, it is set... =-100.0. This means that even if the model is extremely inclined to generate the word "acid rain" based on the original context (the original logit may be very high), adding -100 will make its probability negligible. Conversely, for concepts that the instruction requires to be enhanced or replaced, such as interior, subway station, and sound, an incentive set is constructed. For tokens belonging to the incentive set, such as the token ID of the subway station, a positive bias is applied, for example, +5.0, which significantly increases its probability of being selected. Finally, decoding and generation under constraints are performed. The vector after the above bias adjustment, i.e., the adjusted logits, is sent to the decoding strategy module, such as Temperature Sampling or Nucleus Sampling. This module uses random sampling algorithms such as core sampling or temperature sampling. Taking core sampling as an example, a cumulative probability threshold is set, such as 0.9. The adjusted logits are transformed into a probability distribution through the Softmax function, and then sorted from high to low probability. The top k tokens whose cumulative probability reaches the cumulative probability threshold are selected to form a candidate set. Because the logit of terms like "acid rain" has been significantly reduced, while the logit of terms like "subway station" has been increased, the sampling algorithm naturally avoids the former and chooses the latter. The model continuously generates the next token in this controlled probability space and appends it to the context. This process is repeated until an end token is generated. Finally, the model outputs a complete text, i.e., the candidate revision. For example, the generated text no longer includes an exterior view of a rainy night, but becomes: "Scene: Interior of an abandoned subway station maintenance room - Night\n Action: The hissing sound of acid rain corrosion coming from the ventilation ducts overhead..." This text strictly follows the modification instructions at the underlying computational level, achieving a physically constrained rewrite.

[0061] In a possible preferred implementation, step S52, inputting the rewrite prompt words into a rewrite generator based on a large language model to obtain candidate revisions, includes: generating candidate revisions by applying a temperature-adaptive dynamic logistic bias strategy to the rewrite prompt words.

[0062] In detail, firstly, the rewrite generator loads the rewritten prompts and performs autoregressive calculations. At the current time step, the model outputs the original logits vectors for all words in the vocabulary. To prevent the model from getting bogged down in mediocre repetitive generation, a temperature scaling mechanism is first introduced. A preset temperature coefficient greater than 1.0 is set. This value is obtained by the system from a preset parameter configuration table based on the creative requirement level of the currently generated task (such as "high creativity divergence" mode). =1.2, used to smooth the original probability distribution. The basic probability calculation formula is as follows:

[0063]

[0064] in, This indicates the lexical units before any constraint bias is applied. Sampling probability after temperature adjustment; The original logarithmic value is given; the denominator is the normalized term. When When the value is greater than 1, the difference between high-probability and low-probability words is reduced, increasing the diversity of generation.

[0065] In particular, when inputting rewriting prompts into a large language model-based rewriting generator to generate candidate revisions, the model faces an inherent contradiction between creative generation and constraint control. To make the generated script dialogue vivid and the plot dramatic, a high temperature coefficient is usually required to flatten the probability distribution, thereby encouraging the model to select non-high-frequency but highly expressive lexical units. However, the core requirement of the rewriting task is to strictly adhere to hard instructions such as removing physical acid rain and avoiding outdoor shooting. If a traditional fixed negative bias strategy is used to penalize prohibited entity lexical units, in the high-temperature, high-creativity mode, the fixed penalty value is often diluted due to the flattened overall probability distribution, leading to suppression failure, and the model may still generate prohibited words. In the low-temperature deterministic mode, the fixed penalty value may seem too harsh, pushing the probability of related lexical units to the edge of arithmetic underflow, making it difficult for the model to find reasonable alternatives, such as replacing acid rain with corrosive sounds, thus disrupting the fluency of the text. Therefore, this step adopts a temperature-adaptive dynamic logistic bias strategy, which aims to establish a mechanism that links the penalty intensity to the current generation temperature in real time. This ensures that the suppression ratio of a specific entity remains constant regardless of the creative level of the model, thereby achieving precise execution of production constraints without sacrificing the artistic appeal of the script.

[0066] Next, based on the revision instructions, the specific set of terms (the muted set) that need to be suppressed is identified, such as tokens corresponding to acid rain and corrosion. To achieve precise suppression, instead of using a fixed subtraction bias, a temperature-adaptive logistic negative bias strategy is introduced. The goal of this strategy is to set a constant suppression scaling factor. (in Based on the strictness requirements for script modification, the posterior probability of the penalized word is obtained through pre-set empirical values ​​(such as 0.01 representing strong suppression) or by mapping according to the priority weight of instructions. It is its original probability of This objective relationship can be expressed as:

[0067]

[0068] in, This is the target logarithmic value after applying dynamic bias, which needs to be solved. By taking the natural logarithm of both sides of the above formula, the final execution formula is derived: In this formula, That is, the negative bias term calculated dynamically. .because It is a constant between 0 and 1 (e.g., set). =0.01, meaning the probability of occurrence is reduced to 1% of the original value. It must be a negative number. This formula reveals the relationship between bias strength and temperature. The direct proportional relationship: 1. In high-temperature and high-creativity scenarios, such as setting =1.5: Dynamic bias is used to effectively penetrate specific lexical units in a flat probability distribution. The absolute value will automatically increase. 2. In low-temperature, high-determinism scenarios, such as setting... =0.7: To avoid text breakage due to excessive pressure, dynamic offset is used. The absolute value will automatically decrease. That is, at each generation time step, after the model calculates the original logits, for each token involved in the rewritten prompt word... Calculate its temperature-adaptive negative bias And the calculated Add the original logistic value to the lexicon involved in the rewrite prompt word. Above, we obtain the adjusted logic. ,Should This will then be used for temperature scaling (if temperature scaling has not yet been applied, temperature scaling T will be performed first, followed by softmax, but generally...). It will be added directly to Up, then the whole Divide by T again to enter softmax. Here... The effect of T has already been taken into account, so it is added directly to... Then, when T finally takes effect in the softmax layer, the desired probability suppression effect is achieved. Specific implementation example: For instance, when generating "Scene: [MASK]", the model originally tends to generate an exterior scene (Token ID 3452), its original... =10.0. The revision directive requires changing the exterior scene to an interior scene and setting a suppression factor. =0.01, meaning a reduction of 100 times. ≈-4.605. Scenario A (High Creativity, =1.5): Dynamically calculate the bias =1.5×(-4.605)≈-6.9. (Adjusted) =10.0-6.91=3.09. When entering Softmax, the exponent term is exp(3.09 / 1.5)=exp(2.06), which, compared to the original exp(10 / 1.5)=exp(6.66), has its probability precisely scaled by a factor of s, effectively suppressing the generation of exterior scenes while retaining their potential as semantic substitutes. Case B (low creativity, =0.7): Dynamically calculate the bias =0.7×(-4.605)≈-3.22. (Adjusted) =10.0-3.22=6.78. At this point, the bias value is relatively small, avoiding the direct zeroing of the probability of this word in a sharp distribution.

[0069] Finally, the dynamic strategy adjusted as described above will be applied to the product. The vector input is normalized by the Softmax layer, and the next token is sampled and appended to the end of the sequence. This autoregressive process is repeated until an end token is generated. The final generated complete token sequence is de-segmented and restored to natural language text, which is the candidate revised draft that strictly follows the modification instructions while maintaining narrative coherence. This process ensures that regardless of the randomness of the generation process, the model's suppression effect on high-cost elements such as acid rain and exterior scenery remains consistent, ultimately generating candidate revised drafts that meet cost reduction requirements and have rich details.

[0070] Step S53 is a quality control step based on deep learning semantic matching technology. First, core semantic embeddings are extracted. The two long texts are input into a pre-trained sentence embedding model. The model chosen here is SimCSE. SimCSE is based on the BERT architecture and fine-tuned using a contrastive learning framework. Specifically, a pre-trained BERT is used as the backbone network, which consists of 12 stacked Transformer encoder layers, each containing a multi-head self-attention mechanism and a feedforward neural network. During unsupervised training, SimCSE cleverly utilizes the standard Dropout operation enabled by default within BERT (applied to attention weights and fully connected layer outputs with a probability of 0.1) as a data augmentation technique. It treats the different embeddings obtained from two inputs of the same sentence as positive sample pairs. Due to the randomness of Dropout, even with the same input, the embedding vectors generated by the two forward propagations will differ slightly in space, thus forming a pair of mutually reinforcing positive examples. Meanwhile, the embeddings of all other sentences within the batch are considered negative examples. Embeddings of different sentences are used as negative sample pairs. By optimizing the InfoNCE loss function, the model can generate high-quality semantic vectors with good alignment and uniform distribution. During processing, the initial scene draft is input into the SimCSE model. After multiple layers of self-attention computation, the model outputs the vector representation corresponding to its [CLS] label, i.e. Similarly, candidate revisions are also input into the model to obtain vectors. These two vectors are 768-dimensional high-dimensional floating-point vectors, condensing the deep semantic information of their respective texts, such as plot development, emotional tone, and core events. Next, similarity calculation and evaluation are performed. The cosine similarity between these two vectors is calculated to quantify their proximity in the semantic space. The result of this formula is the retention score. Since the SimCSE model performs well on semantic similarity tasks, this score effectively reflects the consistency between the two texts in telling the same story. For example, if the original text describes Kyle dodging a drone in the rain, and the revised text describes Kyle dodging a drone indoors, the core event (evading pursuit) is the same, resulting in a score of 0.85; however, if the revised text changes to Kyle drinking tea in a coffee shop, the core event changes, and the score may drop to 0.4. Finally, a threshold is determined and output. A preset minimum retention threshold is loaded, for example, set to 0.8. This threshold is set based on test results from a large number of script revision samples, aiming to balance the extent of revisions with plot fidelity. The calculated retention score is compared with the minimum retention threshold: Case 1: If the retention score ≥ the minimum retention threshold, this indicates that although the scene has been replaced (exterior to interior), the core plot (the tension of Kyle being chased) has been well preserved. The verification is successful, and the current candidate revision is directly renamed to the revised scene draft. Case 2: If the retention score < the minimum retention threshold, this indicates that the modification was too drastic and may have resulted in the loss of the core plot. In this case, the system triggers a rollback mechanism. The process automatically jumps back to step S52 with an enhanced instruction, such as "Warning: The previous modification caused the plot to deviate. Please more strictly maintain the original plot development and character motivations when performing modifications." The model will be regenerated based on the new instruction. If the score still does not meet the standard after a preset number of retries (e.g., 3 times), the problem will be marked and submitted for manual review to avoid outputting poor-quality scripts.

[0071] In step S6, the revised scene draft undergoes iterative review and final output to obtain the final film-ready script. That is, although the revised scene draft, after the restrictive rewriting in step S5, should theoretically have followed the cost-reduction and efficiency-enhancing modification instructions, due to the complexity and uncertainty of the natural language generation process, a single modification may not solve all problems in one step. For example, the model might unexpectedly introduce new high-cost props while changing exterior scenes to interior scenes, or even after avoiding the physical risk of acid rain, the scene's time-cycle calculation might still exceed the limit. To ensure that the final script delivered to the user is not only fluent in text but also truly converges to a safe range in various production metrics, a closed-loop verification mechanism is introduced. Step S6, to build an automated quality feedback loop, sends the modified results back to the evaluation system for secondary verification, ensuring that only scripts that fully meet the user's constraints can be output as the final product. This achieves a qualitative leap from linear generation to iterative optimization, maximizing the script's industrial applicability.

[0072] In one possible implementation, step S6 is carried out as follows: First, the revised scene draft is sent back to step S3 (for multi-dimensional analysis and evaluation of film and television adaptation). This process fully reuses the various sub-modules in S3: S31 deconstructs the modified text again to identify a new list of elements; S32 queries the database or calls the regression model again to calculate the cost scores of the new elements; S33 calls the CLIP model again to calculate the visual potential score of the new text; finally, in S34, a new film and television adaptation evaluation vector is obtained. For example, the new vector may be [0.55, 0.90, 0.40, 0.30, 0.20]. It can be seen that compared with the negative score of the original draft, the new vector has significantly reduced risk and budget. Next, a user constraint satisfaction check is performed. The user constraint thresholds determined in S41 (such as budget threshold 0.75, time threshold 0.50, risk threshold 0.40) are called. The scores in the new vector are compared item by item: the new budget score 0.40 < 0.75, which satisfies the requirement. The new time score is 0.30 < 0.50, which meets the requirements. The new risk score is 0.20 < 0.40, which also meets the requirements. If the check results are satisfactory, meaning all key indicators are below the warning threshold, the loop terminates. Finally, metadata tags are added to the current revised scene draft, it is officially renamed the final filmable script, and output to the user through the user interface as a shooting blueprint that can be directly put into production.

[0073] If the requirements are not met (for example, although the interior scenes were changed, the model depicted hundreds of rats to increase tension, causing the cost of the biological props to exceed the budget), the process enters an iterative branch. The current iteration count counter is checked. If the preset maximum number of iterations, such as 3, has not been reached, S4 and S5 will be automatically repeated: generating a new difference vector, diagnosing that the rats are a new problem, generating a strategy to reduce their number, and rewriting the code again. If the maximum number of iterations has been reached, automatic attempts will stop to prevent infinite loops, and the version with the best current evaluation score will be submitted as a draft, along with the final production warning report and modification suggestions, to a human screenwriter or producer for final human decision-making and intervention. This mechanism ensures a smooth transition to human assistance when automation is hindered.

[0074] In summary, the film script generation method based on a generative large model, as described in this application, addresses the technical problem of existing generative models lacking shot-oriented thinking and neglecting production costs and physical feasibility in generated scripts. This solution does not directly output the initial text generated by the large model, but instead treats it as an intermediate product, introducing a multi-dimensional analysis mechanism for film and television applicability. By deconstructing the initial scene draft, its visual potential and production risks are quantified, generating a film and television applicability evaluation vector containing objective evaluation indicators and a production warning report, thereby allowing the system to perceive the shooting logic behind the text. Furthermore, based on this evaluation data and user constraints, specific revision instructions are generated, guiding the model to perform targeted, constrained rewriting and optimization of the draft. This method forces the model to consider both narrative and filmability during the creation process, ensuring that the final output script meets the production standards of the film and television industry, effectively solving the problem of difficulty in implementing generated content.

[0075] Figure 4 This is a block diagram of a film and television script generation system based on a generative large model according to an embodiment of this application. Figure 4 As shown, the film and television script generation system 100 based on a generative large model according to an embodiment of this application includes: a user prompt acquisition module 110, used to acquire user prompts; an initial scene draft generation module 120, used to input user prompts into a creator large model based on a large language model to obtain an initial scene draft; a filmability multi-dimensional analysis module 130, used to perform filmability multi-dimensional analysis and evaluation on the initial scene draft to obtain a filmability evaluation vector and a production warning report; a revision instruction generation module 140, used to generate revision instructions based on the filmability evaluation vector, the production warning report, and user constraints; a draft revision module 150, used to perform constrained rewriting and optimization of the initial scene draft based on the revision instructions to obtain a revised scene draft; and a filmability script generation module 160, used to iteratively review and finally output the revised scene draft to obtain a final filmability script.

[0076] Here, those skilled in the art will understand that the specific operations of each step in the above-described generative large model-based film and television script generation system have been referenced above. Figures 1 to 3 The description of the film and television script generation method based on generative large models has been detailed, and therefore, its repeated description will be omitted.

Claims

1. A method for generating film and television scripts based on a generative large model, characterized in that, include: Get user prompts; User prompts are used to input the creator's big model based on a big language model to obtain an initial scene draft; The initial scene draft undergoes multi-dimensional analysis and evaluation for filmability to obtain a filmability evaluation vector and a production warning report. This includes: deconstructing the production elements of the scene text in the initial scene draft to obtain an extracted element list; quantifying the multi-dimensional cost of the extracted element list to obtain an element score mapping table; evaluating the visual potential of the initial scene draft based on its text description to obtain a visual potential score; aggregating the element score mapping table and the visual potential score to obtain a filmability evaluation vector; and performing element-by-element analysis of the filmability evaluation vector based on a preset threshold to obtain a production warning report. Based on the filmability assessment vector, production early warning report, and user constraints, revision instructions are generated. Based on the revision instructions, the initial scene draft is rewritten and optimized under constraints to obtain the revised scene draft; The revised scene drafts are iteratively reviewed and finalized to obtain the final script suitable for film and television adaptation.

2. The film and television script generation method based on a generative large model according to claim 1, characterized in that, The user-provided input is used to create an initial scene draft based on a large language model of the creator, including: The prompts to users are structured and creatively enhanced to obtain structured enhanced prompts; The structured enhanced prompts are input into the creator big model based on a large language model to obtain the initial scene draft.

3. The film and television script generation method based on a generative large model according to claim 2, characterized in that, Structuring and creatively enhancing user prompts to obtain structured enhanced prompts, including: The user prompts are structured and linked to entities to obtain a parsed element mapping table; Vector-based retrieval-based concept enhancement is performed on the parsed element mapping table to obtain an expanded keyword set; Template instantiation is performed on the parsed element mapping table and the expanded keyword set to obtain structured enhanced prompt words.

4. The film and television script generation method based on a generative large model according to claim 1, characterized in that, A visual potential assessment of the initial scene draft is performed using a textual description to obtain a visual potential score, including: Text encoding is performed on the initial scene draft to obtain the scene text embedding; Based on the target aesthetic concept vector set, conceptual similarity is calculated for scene text embedding to obtain an initial visual potential score; The initial visual potential score is calibrated using a sigmoid function to obtain the final visual potential score.

5. The film and television script generation method based on a generative large model according to claim 1, characterized in that, Based on the filmability assessment vector, production warning report, and user constraints, revision instructions are generated, including: Constraint mapping and difference quantization are performed on user constraints and filmability evaluation vectors to obtain difference vectors; Root cause diagnosis and correction strategy generation are performed on the difference vector and production early warning report to obtain a set of correction strategies; The revision strategy set and the initial scene draft are instantiated using natural language to obtain revision instructions.

6. The film and television script generation method based on a generative large model according to claim 1, characterized in that, Based on the revision instructions, the initial scene draft is rewritten and optimized under constraints to obtain the revised scene draft, including: The initial scene draft and revision instructions are structurally assembled to obtain rewrite prompts; Input the rewrite prompts into a large language model-based rewrite generator to obtain candidate revisions; The dramatic core retention of the candidate revised draft and the initial scene draft was verified to obtain the revised scene draft.

7. A film and television script generation system based on a generative large model, characterized in that, include: The user prompt acquisition module is used to acquire user prompts; The initial scene draft generation module is used to generate an initial scene draft by taking user prompts and inputs based on a large language model of the creator's big model. The filmability-to-filmability multi-dimensional analysis module is used to perform multi-dimensional analysis and evaluation of the initial scene draft to obtain a filmability-to-filmability evaluation vector and a production warning report. This includes: deconstructing the production elements of the scene text in the initial scene draft to obtain an extracted element list; quantifying the multi-dimensional cost of the extracted element list to obtain an element score mapping table; evaluating the visual potential of the initial scene draft based on its text description to obtain a visual potential score; aggregating the element score mapping table and the visual potential score to obtain a filmability-to-filmability evaluation vector; and performing element-by-element analysis of the filmability-to-filmability evaluation vector based on preset thresholds to obtain a production warning report. The revision instruction generation module is used to generate revision instructions based on the filmable evaluation vector, production warning report, and user constraints; The draft revision module is used to perform a constrained rewrite and optimization of the initial scene draft based on revision instructions to obtain the revised scene draft; The filmable script generation module is used to iteratively review and finally output the revised scene drafts to obtain the final filmable script.