AI movie long film prompt word generation and asset multi-dimensional consistent management method and system

CN122513632APending Publication Date: 2026-08-04QUANZHOU MAIFENG FILM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QUANZHOU MAIFENG FILM CO LTD
Filing Date
2026-05-11
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

[0003]现有AIGC在生成时长不低于90分钟的电影长片时,由于现有AI生成模型缺乏跨时序单元的角色外观特征持久化约束机制,导致同一角色在跨幕、跨场、跨镜头的连续生成过程中出现外貌、服饰、体型等视觉特征的不可控漂移,以及角色在剧情时间线中发生年龄变化、受伤等外观状态变更时,前后镜头缺乏特征继承性与逻辑一致性,尤其是在复杂多角色情景下,人物站位也会出现对应的混乱和时序错误;

Benefits of technology

本发明通过构建顶层锚定提示词与全链路分层提示词集群,结合人物多视特征锚定资产包、场景360°环绕视频资产包、复杂度分级站位控制与终态帧继承机制、以及基于心理-情绪映射的表演增强方法,实现了AIGC制作长篇电影过程中全流程的视觉一致性、逻辑连贯性与情感细腻度管控;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122513632A_ABST
    Figure CN122513632A_ABST
Patent Text Reader

Abstract

This invention relates to the field of AIGC (AI Generic Content Generation) digital film and television technology, and particularly to a method and system for AI-generated feature film prompts and multi-dimensional consistency management of assets. The method includes: S1, structured decomposition of the long script and anchoring of top-level prompts; S2, generation of hierarchical prompts across the entire process; S3, anchoring and consistency management of character multi-view feature assets; S4, construction and consistency management of 360° surround video assets for scenes; S5, hierarchical control of multi-person positioning complexity and inheritance of final state frames; and S6, enhancement of character emotional performance. This invention achieves visual consistency, logical coherence, and emotional nuance management throughout the entire process of AIGC feature film production by constructing a top-level anchored prompts and a hierarchical prompts cluster across the entire process, combined with character multi-view feature anchored asset packages, 360° surround video asset packages for scenes, a complexity-hierarchical positioning control and final state frame inheritance mechanism, and a performance enhancement method based on psychological-emotional mapping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of AIGC (AI Generic Content Generation) digital film and television technology, and in particular to a method and system for generating AI-generated cinematic prompts and managing multi-dimensional consistent assets. Background Technology

[0002] AIGC (AI Generic Content Generation) digital film and television production is a process that uses artificial intelligence algorithms to automatically generate video content. It breaks away from the traditional model of video production, which mainly relies on manual shooting and editing, and realizes an intelligent transformation from creative conception to content presentation. AIGC digital film and television production technology integrates the achievements of multiple disciplines such as deep learning, computer vision, and natural language processing. Its generation process includes key steps such as data preparation and preprocessing, feature extraction and encoding, generation and optimization.

[0003] When existing AIGC generates feature films with a runtime of at least 90 minutes, the lack of a persistent constraint mechanism for character appearance features across time units leads to uncontrollable drift of visual features such as appearance, clothing, and body shape during the continuous generation of the same character across scenes, scenes, and shots. Furthermore, when a character undergoes changes in appearance such as age or injury in the plot timeline, there is a lack of feature inheritance and logical consistency between shots. Especially in complex multi-character scenarios, the positioning of characters also becomes chaotic and has temporal errors. Meanwhile, because the script contains a large number of recurring scenes, and the same scene needs to be presented from different shooting angles in different shots, the existing AI generation technology can only output images of the scene from a single perspective. It lacks the support of 360° spatial continuity assets for the scene, which leads to visual inconsistencies such as misaligned background elements and drifting prop positions in the same scene across shots and scenes. Moreover, when the scene changes naturally or is damaged by external forces as the plot progresses, the existing technology cannot maintain the consistency of the scene after the change of scene state across shots, which seriously restricts the industrial usability of AI in multi-character ensemble scenes. Therefore, it is urgent to develop an AI-generated feature film prompt generation and multi-dimensional consistency management method and system to solve the above problems. Summary of the Invention

[0004] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the description and other accompanying drawings.

[0005] The purpose of this invention is to overcome the above-mentioned shortcomings and provide a method and system for AI-generated feature film prompts and multi-dimensional consistency management of assets. This method constructs a top-level anchored prompt and a full-link hierarchical prompt cluster, combined with character multi-view feature anchored asset packages, scene 360° surround video asset packages, complexity-level position control and final state frame inheritance mechanism, and a performance enhancement method based on psychological-emotion mapping, to achieve full-process visual consistency, logical coherence and emotional subtlety management in the production of AIGC feature films. The multi-view character feature anchoring method establishes a multi-view character feature asset package bound to the character ID, serving as the character anchor point throughout the entire film. Driven by the timestamp of the feature film, the character's appearance can be modified and automatically matched to the scene. The 360° surround video asset method establishes a 360° surround video asset package bound to the spatial ID. Through video sequence retrieval and matching, it automatically obtains background images that strictly correspond to the shooting parameters. Moreover, the new asset package is an expansion and iteration based on the old asset package, thereby ensuring the feature inheritance of character appearance and scene changes, achieving consistency throughout the entire film. To address the issue of complex multi-character positioning drift, the system is tiered based on the number of characters appearing. When the number of characters is small, positioning is directly controlled through text prompts. When the number of characters is large, AI-generated or manually formatted positioning reference images are used as compositional constraints to ensure the continuity of positioning across different characters. To address the issue of AI characters lacking micro-expressions and emotional layers, a psychological-emotional mapping knowledge base is established to generate emotional representations that are causally related to the plot context. The model is then set up for automatic verification and iterative correction, forming a closed-loop optimization. This invention provides a method for generating AI-powered feature film prompts and managing multi-dimensional asset consistency, including: S1. Structured decomposition and top-level keyword anchoring of long-form scripts: For feature films with a runtime of no less than 90 minutes, the complete script is decomposed into narrative main line units, plot sub-line units, and three-level temporal units of scene-scene-shot, generating top-level anchoring keywords covering the entire film. S2. Full-link hierarchical prompt generation: Based on the top-level anchored prompts, hierarchical sub-prompt clusters are generated according to the full-link dimension of script-storyboard-camera-character-lighting-sound effects-rhythm. Each level of sub-prompt inherits the corresponding weight parameters of the top-level anchored prompts, and each level of sub-prompt is set with independent adjustable weights and association constraint rules. S3. Character Multi-View Feature Asset Anchoring and Consistency Control: Construct a multi-view feature asset package for anchoring each character in the script, specifically including: S31. Based on the character biography of each character, obtain the corresponding set of appearance feature description words, generate a single-view white background baseline image of each character, and store it as the initial feature anchoring asset after manual confirmation. S32. Using the initial feature anchoring assets as graph constraint input, generate at least four view graph sets for the role, and bind them with the role's unique identity ID to form a multi-view feature anchoring asset package; S33. When generating an AI screen containing any temporal unit of the character, automatically call the multi-view feature anchoring asset package of the character and input it as the character appearance constraint into the latent space control layer of the AI ​​generation model. S34. For events in the script where the character's status changes with timeline, such as age changes, injury, or appearance changes, generate new multi-view features to anchor the asset package after the status change based on the current valid asset package, and mark the valid timestamp interval. S35. Based on the timestamp interval of the current time sequence unit, automatically match and call the corresponding version of the character asset package; S4. Scene 360° Surround Video Asset Construction and Consistency Control: This involves constructing a 360° surround video asset package for each independent scene in the script, specifically including: S41. Extract the spatial description text of all independent scenes in the script and generate a single-view scene baseline map for each scene. S42. Using the single-view scene reference map as the initial frame constraint input, generate a 360° surround video sequence with the center point of the scene as the axis, and bind it with the scene's unique spatial identifier ID to form a 360° surround video asset package. S43. When generating AI character images for any shot in the scene, obtain the preset shooting parameters of the current shot, retrieve and extract frames with matching perspectives from the surrounding video asset package as the scene background layer. S44. Establish a mapping between the surround video asset package and all scenes in the script where the scene appears. When generating any shot in any scene, retrieve and automatically call the surround video asset package of the corresponding scene based on the current scene to ensure visual consistency of the same scene throughout the film. S45. For scenes in the script that change over time or are destroyed by external forces, generate a new surround video asset package after the state change based on the base frame of the current valid asset package, and mark the valid timestamp interval. S5. Multi-player positioning complexity hierarchical control and final state frame inheritance: Perform positioning control for scenes with multiple roles, specifically including: S51. Obtain the number N of characters appearing simultaneously in the current session and the interaction data between characters. Based on the value of N and the interaction complexity, automatically divide the positioning control strategy into at least two processing levels: Pre-build cue text containing descriptions of the relative space and position of each character. When N≤3, enable direct position control mode and adjust the positions of multiple characters through the cue text. Pre-construct static positioning reference images generated by AI, video keyframe positioning reference images generated by AI, manually laid-out positioning reference images, or real-life shooting positioning reference images. When N≥4, enable the auxiliary positioning mode and input one of the following as composition constraints into the AI ​​generation model to constrain the spatial distribution of multiple characters in the generated image. S52. When multiple characters undergo plot-driven position changes in a scene, the final frame of the previous shot is obtained as the anchor asset of the position state. This anchor asset of the position state is used as the initial frame reference input or composition constraint input, and the AI ​​generates the next shot to achieve strict inheritance of the position state across shots. S6. Enhanced Character Emotional Performance: Enhances the emotional nuances of character performances within a single shot, specifically including: S61. Extract action sequence data, dialogue text data, behavioral antecedent data, and behavioral consequence data from single-shot storyboard text; S62. Query the preset psychological-emotion mapping knowledge base to obtain a set of emotional representation descriptions that match the psychological changes before and after the behavior; S63. Integrate the emotional representation description set with action sequence data and dialogue text data to generate performance-enhanced storyboard prompts; S64. Input the performance-enhanced storyboard prompts generated in step S63 into the AI ​​generation model to obtain preliminary performance footage. Call the performance emotion recognition model to classify the preliminary performance footage by emotion label and score the emotion intensity. If the recognition result deviates from the emotion representation description set set in step S62 by more than the preset threshold, the weight parameters of the micro-expression or body language description in the prompts will be automatically adjusted and regenerated until the verification is passed. S65. For the performance output of the same character in consecutive shots, extract the performance emotion recognition results of each shot, construct the emotional intensity change curve of the character in the scene or act, compare the emotional intensity change curve with the pre-set emotional rhythm spectrum of the script, and perform the supplementary enhancement operation of steps S63-S64 for shots that deviate from the pre-set curve to ensure the emotional progression logic of the character's performance in cross-shot narrative. S7. Cross-unit consistency verification and correction for feature films: For the sub-prompt word clusters and generated images of all time units in the whole film, perform multi-dimensional consistency verification across scenes, scenes and shots. Specifically, it includes verification of character setting consistency, visual style consistency, narrative sequence consistency, emotional curve coherence, and lighting logic coherence. For sub-prompt words that do not meet the verification threshold, weight correction and content iteration are automatically triggered. S8, Industrial-grade rendering adaptation and dynamic optimization: The verified layered prompt word cluster is mapped to a standardized parameter set that can be recognized by the AI ​​film and television generation engine, and a mapping relationship between prompt word parameters and rendering engine interface is established. During the production of feature films, the weight parameters and constraint rules of subsequent time-series unit sub-prompt words are dynamically adjusted based on the effect data of the generated segments.

[0006] In some embodiments, in step S34, the specific steps for handling the character's status change events—such as age change, injury, or appearance change—over time in the script are as follows: S341. Read the currently valid multi-view feature anchored asset package for this role; S342. Obtain the set of target feature descriptive words that describe the state transition event; S343. Input the single-view baseline map in the current valid asset package and the target feature descriptor set into the AI ​​image generation model to generate a new single-view baseline map after the state change. S344. Perform the multi-view expansion operation of steps S31-S32 on the new single-view baseline map to generate a new multi-view feature anchored asset package after the state change. S345. Bind the new asset package to the character ID and mark its valid timestamp range in the corresponding scene or act of the script.

[0007] In some embodiments, in step S35, based on the current time sequence unit's position in the script timeline, the valid multi-view feature anchor asset package version of the character within the timestamp interval is automatically matched and invoked to ensure that the character's appearance is strictly synchronized with the plot development.

[0008] In some embodiments, in step S45, the specific operational steps for the scene in the script to undergo natural changes or external destruction events over time are as follows: S451. Read the base frame from the currently valid 360° surround video asset package for this scene; S452. Obtain the set of target feature descriptive words that describe the natural change or external destruction event; S453. Input the baseline frame and the target feature descriptor set into the AI ​​image generation model to generate a new single-view scene baseline map after the state change. S454. Perform the surround video extension operation of steps S41-S42 on the new single-view scene baseline map to generate a new 360° surround video asset package after the state change. S455. Bind the new asset package to the scene ID and mark its valid timestamp range in the corresponding scene or act of the script.

[0009] In some embodiments, when an external force damage event occurs in the scene, the following steps are further performed: The system invokes a pre-defined AI video generation model to generate video clips depicting the dynamic process of the destructive event. The stable final state frame after the destruction is completed is extracted from the video clip of the destruction process and used as the base image for generating the new single-view scene reference map in step S45; The video clip of the destruction process is stored as a dynamic destruction event asset for the scene. When generating subsequent continuous shots of the scene after destruction, the intermediate state frames are extracted as the scene background layer to maintain the consistency of the distribution of debris and fragments across shots.

[0010] In some embodiments, for the auxiliary control mode in step S51: Call the preset AI image generation model, input the position description text containing the number of characters and their rough positional relationships, and generate a static position reference image; The preset AI video generation model is invoked. With the functions of dialogue generation and complex action generation turned off, only a panoramic video of the spatial distribution of multiple characters is generated, and key frames are extracted from it as a reference map for the video key frame positions. Use image processing tools to manually arrange the preset character materials to generate a layout and positioning reference diagram; We obtained real-life footage of the pre-rehearsed positioning as a reference image for the actual shooting.

[0011] In some embodiments, the cross-cell consistency check in step S7 specifically includes: Character consistency check: Based on the character's entire life cycle, check the deviation values ​​of appearance, clothing, age, personality and behavioral logic of characters across shots, with the deviation threshold set to no more than 5%; Visual style consistency check: Based on the overall visual tone of the film, check the deviation values ​​of color system, art style and picture quality of cross-scene shots, with the deviation threshold set to no more than 8%; Narrative timeline consistency verification: Based on the overall timeline of the film, verify the timeline logic of cross-scene plots, the continuity of props, and the compliance of causal relationships between events; Emotional curve coherence verification: Based on the overall narrative rhythm spectrum of the entire film, verify the matching degree between the fluctuation of the emotional curve across scenes and the preset overall spectrum, with the matching degree threshold set to no less than 90%; Lighting and shadow logic consistency verification: Based on the overall lighting and shadow rules of the whole film, verify the logical consistency of light source direction, day and night sequence, and ambient lighting across shots.

[0012] In some embodiments, in step S7, the automatic correction iteration specifically involves: when the verification result of a sub-prompt word exceeds a preset deviation threshold, the system automatically locks the parent-level related prompt word corresponding to the sub-prompt word, and automatically adjusts the weight parameters and content description of the sub-prompt word based on the top-level anchor prompt word until the verification result meets the threshold requirement, and generates a correction log archive.

[0013] In some embodiments, in step S8, the standardized parameter set includes a shot parameter set, a character model parameter set, a rendering environment parameter set, an audio parameter set, and a timing control parameter set. Each type of parameter set forms a unique mapping relationship with the corresponding level of prompt content, supporting direct access to the native rendering interface of mainstream AI film and television generation engines.

[0014] In some embodiments, in step S8, dynamic adjustment specifically involves: during the production of a feature film, after each scene-level unit is generated and its effect is evaluated, automatically extracting quantitative data on the unit's visual consistency, narrative matching degree, and rendering quality, and dynamically fine-tuning the sub-cue word weight parameters of subsequent scene-level units within a range of ±10% based on this data, without changing the fixed weight parameters of the top-level anchor cue words during the fine-tuning process.

[0015] An AI-powered feature film prompt generation and multi-dimensional consistent asset management system, which includes: Top-level anchoring module: used to structurally decompose the script, generate and store top-level anchoring prompts and their fixed weight parameters; Hierarchical prompt word generation module: Connected to the top-level anchoring module, it is used to generate a cluster of sub-prompt words with full-link hierarchical binding, and configure the independent weights and association constraint rules for each level; Character Asset Anchoring Module: Used to generate and manage multi-view feature anchored asset packages for characters, and supports story-driven asset version iteration and timestamp matching; Scene Asset Anchoring Module: Used to generate and manage 360° surround video asset packages for scenes, and supports multi-angle background screenshot calls and plot-driven asset version iteration; Multi-person positioning control module: Used to execute positioning control strategies with varying levels of complexity and manage the inheritance of positioning status anchor point assets across shots; Emotional Performance Enhancement Module: This module generates performance-enhanced storyboard prompts based on a psychological-emotional mapping knowledge base, and calls the performance emotion recognition model for verification and iterative correction. Consistency verification module: used to perform multi-dimensional consistency verification across time units of the entire chip, and trigger correction iterations for unqualified sub-prompt words; Render Adaptation and Dynamic Optimization Module: This module maps the hierarchical prompt word clusters to a standardized set of rendering parameters and dynamically adjusts the prompt word parameters based on the generated effect data.

[0016] In some embodiments, the system also includes an access control module connected to the top-level anchoring module, which is used to set the modification permissions for the top-level anchoring prompts, authorize only users to modify the fixed weight parameters, and retain log records of all modification operations.

[0017] In some embodiments, the system also includes an effect quantification evaluation module, which is connected to the consistency verification module and the dynamic optimization module, respectively. This module is used to perform a full-dimensional effect quantification score on the generated feature film clips, generate a traceable effect evaluation report, and provide data support for dynamic optimization.

[0018] By adopting the above technical solution, the beneficial effects of the present invention are: This invention achieves full-process control over visual consistency, logical coherence, and emotional subtlety in the production of AIGC feature films by constructing top-level anchored prompts and a full-link hierarchical prompt cluster, combined with character multi-view feature anchored asset packages, scene 360° surround video asset packages, complexity-level position control and final state frame inheritance mechanism, and performance enhancement method based on psychological-emotion mapping. The multi-view character feature anchoring method establishes a multi-view character feature asset package bound to the character ID, serving as the character anchor point throughout the entire film. Driven by the timestamp of the feature film, the character's appearance can be modified and automatically matched to the scene. The 360° surround video asset method establishes a 360° surround video asset package bound to the spatial ID. Through video sequence retrieval and matching, it automatically obtains background images that strictly correspond to the shooting parameters. Moreover, the new asset package is an expansion and iteration based on the old asset package, thereby ensuring the feature inheritance of character appearance and scene changes, achieving consistency throughout the entire film. To address the issue of complex multi-character positioning drift, the system is tiered based on the number of characters appearing. When the number of characters is small, positioning is directly controlled through text prompts. When the number of characters is large, AI-generated or manually formatted positioning reference images are used as compositional constraints to ensure the continuity of positioning across different characters. To address the issue of AI characters lacking micro-expressions and emotional layers, a psychological-emotional mapping knowledge base is established to generate emotional representations that are causally related to the plot context. The model is then set up for automatic verification and iterative correction, forming a closed-loop optimization.

[0019] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure.

[0020] Undoubtedly, such and other objects of the present invention will become more apparent after the following detailed description of the preferred embodiments, which are illustrated in various accompanying drawings and figures.

[0021] To make the above and other objects, features and advantages of the present invention more apparent and understandable, one or more preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0022] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0023] In the accompanying drawings, the same parts use the same reference numerals, and the drawings are schematic and not necessarily drawn to actual scale.

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only one or more embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on such drawings without creative effort.

[0025] Figure 1 This is a schematic diagram of the overall process of the control method in some embodiments of the present invention; Figure 2 This is a schematic diagram of the process for constructing a multi-view feature anchored asset package for people in some embodiments of the present invention; Figure 3 This is a schematic diagram of the scene 360° surround video asset construction process in some embodiments of the present invention; Figure 4 This is a schematic diagram of the Wei Junlong 18-year-old multi-view asset anchoring asset package in some embodiments of the present invention; Figure 5 This is a schematic diagram of the 360° environmental asset package of the Wei Family Mansion in its early intact state in some embodiments of the present invention; Figure 6 This is a schematic diagram of the arrangement of multiple people standing in some embodiments of the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0027] Furthermore, in the description of this invention, it should be understood that the terms "center," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "axial," "radial," and "circumferential" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0028] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral unit; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. However, specifying a direct connection indicates that the two main bodies are not connected through a transitional structure, but rather formed as a whole through a connecting structure. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0029] In this invention, unless otherwise expressly specified and limited, the first feature "on" or "below" the second feature may be in direct contact with the first and second features, or indirect contact through an intermediate medium. In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0030] Reference Figure 1 , Figure 1 This is a schematic diagram of the overall process of the control method in some embodiments of the present invention.

[0031] According to some embodiments of the present invention, the present invention provides a method for AI-generated feature film prompts and multi-dimensional consistent asset management, including: S1. Structured Decomposition and Top-Level Anchoring of Long-Form Scripts: For feature films with a runtime of 90 minutes or more, the complete script is decomposed into narrative main line units, plot sub-line units, and three-level temporal units of scene-scene-shot, generating top-level anchoring keywords covering the entire film. These top-level anchoring keywords include fixed weight parameters for the film's world-building, the life-cycle settings of core characters, the overall visual style and tone, the narrative rhythm spectrum, and the general rules of lighting and sound effects.

[0032] S2. Full-link hierarchical prompt generation: Based on the top-level anchored prompts, hierarchical sub-prompt clusters are generated according to the full-link dimensions of script-storyboard-camera-character-lighting-sound effects-rhythm. Each level of sub-prompt inherits the corresponding weight parameters of the top-level anchored prompts, and each level of sub-prompt is set with independent adjustable weights and association constraint rules.

[0033] Reference Figure 2 , Figure 2 This is a schematic diagram illustrating the process of constructing a multi-view feature anchored asset package for a person in some embodiments of the present invention.

[0034] S3. Character Multi-View Feature Asset Anchoring and Consistency Control: Construct a multi-view feature asset package for anchoring each character in the script, specifically including: S31. Based on the character biography of each character, obtain the corresponding set of appearance feature description words, generate a single-view white background baseline image of each character, and store it as the initial feature anchoring asset after manual confirmation. The text of character biographies is obtained after the script of a feature film has been read by a human. The set of descriptive words for the appearance features of each character is extracted. The set of descriptive words for the appearance features is input into a preset AI image generation model to generate a single-view white background baseline image for each character. The single-view white background baseline image is submitted to the human confirmation terminal. After receiving the human confirmation instruction, the single-view white background baseline image is stored as the initial feature anchoring asset of the corresponding character. S32. Using the initial feature anchoring assets as graph constraint input, generate at least four view graph sets for the role, and bind them with the role's unique identity ID to form a multi-view feature anchoring asset package; The character's four-view atlas is generated by calling a preset multi-view generation model. The four-view atlas specifically includes the front, back, left, and right sides. For some characters, five view atlases can also be selected to add 3 / 4 side or top / bottom auxiliary perspectives. S33. When generating an AI screen containing any temporal unit of the character, the multi-view feature anchor asset package of the character is automatically called and used as a character appearance constraint condition input into the latent space control layer of the AI ​​generation model. The latent space control layer can be a ControlNet reference graph input channel or an IP-Adapter feature embedding layer to constrain the visual features of the character in the generated screen to maintain consistency with the anchor assets. S34. For events in the script where the character's status changes with timeline, such as age changes, injury, or appearance changes, generate new multi-view features to anchor the asset package after the status change based on the current valid asset package, and mark the valid timestamp interval. In step S34, the specific steps for handling events related to the character's changing status over time, such as aging, injury, or appearance, are as follows: S341. Read the currently valid multi-view feature anchored asset package for this role; S342. Obtain the set of target feature descriptive words that describe the state transition event, such as "same character, age increased by 20 years", "new scar on left cheek", etc. S343. Input the single-view baseline map in the current valid asset package and the target feature descriptor set into the AI ​​image generation model to generate a new single-view baseline map after the state change. S344. Perform the multi-view expansion operation of steps S31-S32 on the new single-view baseline map to generate a new multi-view feature anchored asset package after the state change. S345. Bind the new asset package to the character ID and mark its valid timestamp range in the corresponding scene or act of the script; S35. Based on the timestamp interval of the current time sequence unit, automatically match and call the corresponding version of the character asset package; In step S35, based on the current position of the time sequence unit in the script timeline, the system automatically matches and calls the valid multi-view feature anchor asset package version of the character within the timestamp interval to ensure that the character's appearance is strictly synchronized with the plot development. The precision of the valid timestamp interval is at the scene level, corresponding to the time range of a single scene in the script. That is, the same character uses the same version asset package in the same scene, and the system automatically detects the status change event and switches the version when crossing scenes. When generating a new single-view baseline map in steps S341-343, the single-view baseline map of the current effective asset package is used as the initial latent space encoding input for the graph-generated map. The noise reduction intensity parameter is set in the range of 0.4-0.7 to balance feature preservation and state change amplitude.

[0035] Reference Figure 3 , Figure 3 This is a schematic diagram of the 360° surround video asset construction process in some embodiments of the present invention.

[0036] S4. Scene 360° Surround Video Asset Construction and Consistency Control: This involves constructing a 360° surround video asset package for each independent scene in the script, specifically including: S41. Extract the spatial description text of all independent scenes in the script and generate a single-view scene baseline map for each scene. The script of a feature film is obtained, and the spatial description text of all independent scenes in the script is extracted through a preset AI semantic analysis model. The spatial description text of each scene is input into a preset AI image generation model to generate a single-view scene baseline map of each scene. The single-view scene baseline map corresponds to the typical shooting angle when the scene first appears in the script. S42. Using the single-view scene reference map as the initial frame constraint input, generate a 360° surround video sequence with the center point of the scene as the axis, and bind it with the scene's unique spatial identifier ID to form a 360° surround video asset package. The 360° surround video sequence is generated by calling a preset AI video generation model and includes continuous frames of the scene within any horizontal azimuth angle and a certain pitch angle range; the video sequence covers at least the range of horizontal azimuth angle 0°-360° and pitch angle -30° to +60°, with a frame rate of no less than 24fps and a total surround time of no less than 4 seconds to ensure that there are corresponding background frames available for any camera angle; S43. When generating AI character images for any shot in the scene, obtain the preset shooting parameters of the current shot, retrieve and extract frames with matching perspectives from the surrounding video asset package as the scene background layer; the preset shooting parameters include but are not limited to the azimuth angle, pitch angle and focal length of the camera position. When retrieving matching frames based on the shooting parameters, the azimuth angle matching tolerance does not exceed ±5° and the pitch angle matching tolerance does not exceed ±3° to ensure the visual coordination between the background perspective relationship and the character synthesis. S44. Establish a mapping between the surround video asset package and all scenes in the script where the scene appears. When generating any shot in any scene, retrieve and automatically call the surround video asset package of the corresponding scene based on the current scene to ensure visual consistency of the same scene throughout the film. S45. For scenes in the script that change over time or are destroyed by external forces, generate a new surround video asset package after the state change based on the base frame of the current valid asset package, and mark the valid timestamp interval; scenes that change over time can be day and night alternation, seasonal changes, era changes, etc., and external destruction events can be explosions, collapses, fires, etc. In step S45, the specific operational steps for the scene in the script that undergoes natural changes or external destruction over time are as follows: S451. Read the reference frame in the currently valid 360° surround video asset package of the scene; that is, the single-view scene reference map in step S41 or the reference frame of the previous state version. S452. Obtain the set of target feature descriptive words that describe the natural change or external destruction event; S453. Input the baseline frame and the target feature descriptor set into the AI ​​image generation model to generate a new single-view scene baseline map after the state change. S454. Perform the surround video extension operation of steps S41-S42 on the new single-view scene baseline map to generate a new 360° surround video asset package after the state change. S455. Bind the new asset package to the scene ID and mark its valid timestamp range in the corresponding scene or act of the script; The precision of this effective timestamp interval is at the field level. The same version of the asset package is used in the same scene and the same session. When crossing sessions, the system detects scene state change events and automatically switches the asset version. When an external force damages the scene, the following steps are performed: The system invokes a pre-defined AI video generation model to generate video clips depicting the dynamic process of the destructive event, such as wall collapse and flame spread. The stable final state frame after the destruction is completed is extracted from the video clip of the destruction process and used as the base image for generating the new single-view scene reference map in step S45; The video clip of the destruction process is stored as a dynamic destruction event asset for the scene. When generating subsequent continuous shots of the scene after destruction, the intermediate state frames are extracted as the scene background layer to maintain the consistency of the distribution of debris and fragments across shots. When generating a new single-view scene baseline map, the noise reduction intensity parameter of the image is set in the range of 0.3-0.6 to ensure that the spatial structure (wall position, door and window layout, large props) is preserved, and only the visual details (tone, vegetation, degree of damage) are changed in accordance with the script requirements.

[0037] S5. Multi-player positioning complexity hierarchical control and final state frame inheritance: Perform positioning control for scenes with multiple roles, specifically including: S51. Obtain the number N of characters appearing simultaneously in the current session and the interaction data between characters. Based on the value of N and the interaction complexity, automatically divide the positioning control strategy into at least two processing levels: Pre-construct cue text containing descriptions of the relative space and position of each character. When N≤3, enable direct position control mode and adjust the positions of multiple characters through cue text. For example, the specific content of the cue text can be: "A is to the left of B", "C is about one meter in front of A", etc. You can also add action status descriptors and dialogue descriptions of each character. Input the cue text into the AI ​​generation model to generate a multi-character screen that conforms to the position constraints. Pre-construct static positioning reference images generated by AI, AI-generated video keyframe positioning reference images, manually laid-out positioning reference images, or real-life shooting positioning reference images. When N≥4, enable the auxiliary positioning mode. Input one of the following as a composition constraint into the AI ​​generation model to constrain the spatial distribution of multiple characters in the generated image. The composition constraint input can be input into the model's ControlNet or the raw image reference channel. The boundary threshold between the first processing level and the second processing level is set to N=3 (i.e., N≤3 is the first level, N≥4 is the second level). According to the actual needs of the scenario, preferably, when N≥6, the third processing level - the storyboard grouping and compositing mode can be further enabled, which splits multiple characters into multiple subgroups and generates them separately before compositing. For the auxiliary control mode: Call the preset AI image generation model, input the position description text containing the number of characters and their rough positional relationships, and generate a static position reference image; The preset AI video generation model is invoked. With the functions of dialogue generation and complex action generation turned off, only a panoramic video of the spatial distribution of multiple characters is generated, and key frames are extracted from it as a reference map of video key frame positions. The frame rate of the panoramic video of the position preview is not less than 8fps and the duration is not less than 2 seconds, so as to provide sufficient position status information for manual confirmation. Use image processing tools to manually arrange the preset character materials to generate a layout and positioning reference diagram; Obtain real-life footage of the pre-rehearsal positioning as a reference image for the actual positioning; S52. When multiple characters undergo plot-driven position changes (such as character movement, character addition or removal, formation change), the final state frame of the previous shot is obtained as the anchor asset of the position state. This final state frame records the position state and action freeze state of all characters at the moment the position change is completed. The anchor asset of the position state is used as the initial frame reference input or composition constraint input. The AI ​​generates the next shot, realizing strict inheritance of the position state across shots. The final frame inheritance operation in step S52 is triggered when any of the following conditions are met: 1. Any character undergoes a spatial displacement greater than 0.5 meters; 2. The number of characters in the scene increases or decreases; 3. The formation structure undergoes an unpredictable reorganization. The anchor point assets of this position are stored at the shot level, that is, the final frame of each shot serves as the position anchor point of the next shot, forming a chain inheritance relationship. If the next shot involves a new character not included in the position state anchor point asset, the new character is superimposed onto the corresponding position of the position state anchor point asset using the reference position map generation method in step S51, and then used as a constraint input; if the next shot adds a new character, repeat steps S51-S52. When generating multi-character screens in step S51 or S52, the multi-view feature anchoring asset package corresponding to each character can be called simultaneously. The appearance feature constraints and spatial positioning constraints of each character are jointly encoded and then input into the AI ​​generation model to ensure the consistency of appearance and the accuracy of positioning of each character.

[0038] S6. Enhanced Character Emotional Performance: Enhances the emotional nuances of character performances within a single shot, specifically including: S61. Extract action sequence data, dialogue text data, behavioral antecedent data, and behavioral consequence data from single-shot storyboard text; Action sequence data: The specific physical actions performed by the character in this shot and their temporal sequence; Dialogue text data: The character's lines and tone in this shot; Antecedent data: The character's immediate psychological state and triggering events in the story before performing the action or uttering the line; Behavioral consequence data: the expected emotional feedback and plot progression after a character performs this action or utters this line; S62. Query the preset psychological-emotion mapping knowledge base to obtain a set of emotional representation descriptions that match the psychological changes before and after the behavior; The psychology-emotion mapping knowledge base is stored in a four-level association structure of "triggering event - psychological state - emotion label - explicit feature", and pre-stores no less than 50 basic emotion labels and their corresponding micro-expression or body language description templates. This set of emotion representation descriptions includes, but is not limited to: Micro-expression features include: eyebrow shape, eyelid opening and closing, corner of the mouth curvature, and changes in the nostrils; Supplementary description of body language: shoulder posture, hand micro-movements, trunk tilt, center of gravity distribution, etc.; Description of breathing and rhythm: breathing frequency cues, duration of pauses in movement, eye movement patterns, etc.; S63. The emotional representation description set is fused with the action sequence data and dialogue text data to generate performance-enhanced storyboard prompts; the performance-enhanced storyboard prompts supplement the original action or dialogue descriptions with descriptions of the character's emotional outward characteristics when performing the action or saying the line. S64. Input the performance-enhanced storyboard prompts generated in step S63 into the AI ​​generation model to obtain preliminary performance footage. Call the performance emotion recognition model to classify the preliminary performance footage by emotion label and score the emotion intensity. If the recognition result deviates from the emotion representation description set set in step S62 by more than the preset threshold, the weight parameters of the micro-expression or body language description in the prompts will be automatically adjusted and regenerated until the verification is passed. The performance emotion recognition model outputs a continuous-dimensional emotion intensity score (Valence-Arousal two-dimensional model, range 0-1), with the deviation threshold set at ±0.15 for the Valence dimension and ±0.15 for the Arousal dimension. When automatically adjusting the weight parameters of micro-expression / body language description, the single adjustment step size is ±5%, and the maximum adjustment range does not exceed ±30% of the initial weight.

[0039] S65. For the performance output of the same character in consecutive shots, extract the performance emotion recognition results of each shot, construct the emotion intensity change curve of the character in that scene or act, compare the emotion intensity change curve with the preset emotion rhythm spectrum of the script, and perform the supplementary enhancement operation of steps S63-S64 for shots that deviate from the preset curve to ensure the emotional progression logic of the character's performance in cross-shot narrative; the matching degree threshold between the emotion intensity change curve and the preset emotion rhythm spectrum is set to no less than 85%, and when it is lower than this threshold, the performance layer supplementary enhancement operation of the corresponding shot is triggered.

[0040] S7. Cross-unit consistency verification and correction for feature films: For the sub-prompt word clusters and generated images of all time units in the whole film, perform multi-dimensional consistency verification across scenes, scenes and shots. Specifically, it includes verification of character setting consistency, visual style consistency, narrative sequence consistency, emotional curve coherence, and lighting logic coherence. For sub-prompt words that do not meet the verification threshold, weight correction and content iteration are automatically triggered. In step S7, cross-cell consistency verification specifically includes: Character consistency check: Based on the character's entire life cycle, check the deviation values ​​of appearance, clothing, age, personality and behavioral logic of characters across shots, with the deviation threshold set to no more than 5%; Visual style consistency check: Based on the overall visual tone of the film, check the deviation values ​​of color system, art style and picture quality of cross-scene shots, with the deviation threshold set to no more than 8%; Narrative timeline consistency verification: Based on the overall timeline of the film, verify the timeline logic of cross-scene plots, the continuity of props, and the compliance of causal relationships between events; Emotional curve coherence verification: Based on the overall narrative rhythm spectrum of the entire film, verify the matching degree between the fluctuation of the emotional curve across scenes and the preset overall spectrum, with the matching degree threshold set to no less than 90%; Light and shadow logic consistency verification: Based on the overall lighting and shadow rules of the whole film, verify the logical consistency of light source direction, day and night sequence, and ambient lighting across shots; The automatic correction iteration is as follows: when the verification result of a sub-prompt word exceeds the preset deviation threshold, the system automatically locks the corresponding parent-level related prompt word of the sub-prompt word, and automatically adjusts the weight parameters and content description of the sub-prompt word based on the top-level anchor prompt word until the verification result meets the threshold requirements, and generates a correction log for archiving.

[0041] S8, Industrial-grade rendering adaptation and dynamic optimization: The verified layered prompt word cluster is mapped to a standardized parameter set that can be recognized by the AI ​​film and television generation engine, and a mapping relationship between prompt word parameters and rendering engine interface is established. During the production of feature films, the weight parameters and constraint rules of subsequent time-series unit sub-prompt words are dynamically adjusted based on the effect data of the generated segments. The standardized parameter sets include shot parameter sets, character model parameter sets, rendering environment parameter sets, audio parameter sets, and timing control parameter sets. Each type of parameter set has a unique mapping relationship with the corresponding level of prompt words, and supports direct access to the native rendering interface of mainstream AI film and television generation engines. The dynamic adjustment is as follows: During the production of the feature film, after each segment-level unit is generated and its effect is evaluated, the quantitative data of the unit's visual consistency, narrative matching degree, and rendering quality are automatically extracted. Based on this data, the weight parameters of the sub-cue words of the subsequent segment-level units are dynamically fine-tuned within a range of ±10%. The fine-tuning process does not change the fixed weight parameters of the top-level anchor cue words. This invention also provides an AI-powered feature film prompt generation and multi-dimensional consistent asset management system, which includes: Top-level anchoring module: used to structurally decompose the script, generate and store top-level anchoring prompts and their fixed weight parameters; Hierarchical prompt word generation module: Connected to the top-level anchoring module, it is used to generate a cluster of sub-prompt words with full-link hierarchical binding, and configure the independent weights and association constraint rules for each level; Character Asset Anchoring Module: Used to generate and manage multi-view feature anchored asset packages for characters, and supports story-driven asset version iteration and timestamp matching; Scene Asset Anchoring Module: Used to generate and manage 360° surround video asset packages for scenes, and supports multi-angle background screenshot calls and plot-driven asset version iteration; Multi-person positioning control module: Used to execute positioning control strategies with varying levels of complexity and manage the inheritance of positioning status anchor point assets across shots; Emotional Performance Enhancement Module: This module generates performance-enhanced storyboard prompts based on a psychological-emotional mapping knowledge base, and calls the performance emotion recognition model for verification and iterative correction. Consistency verification module: used to perform multi-dimensional consistency verification across time units of the entire chip, and trigger correction iterations for unqualified sub-prompt words; Rendering adaptation and dynamic optimization module: used to map the layered prompt word cluster to a standardized set of rendering parameters, and dynamically adjust the prompt word parameters based on the generated effect data; The system also includes an access control module, which is connected to the top-level anchoring module. It is used to set the modification permissions for the top-level anchoring prompts, authorizing only users to modify the fixed weight parameters, and retaining log records of all modification operations. The system also includes an effect quantification evaluation module, which is connected to the consistency verification module and the dynamic optimization module, respectively. It is used to perform full-dimensional effect quantification scoring on the generated feature film clips, generate traceable effect evaluation reports, and provide data support for dynamic optimization.

[0042] Example Reference Figures 4-6 , Figure 4 This is a schematic diagram of the Wei Junlong 18-year-old multi-view asset anchoring asset package in some embodiments of the present invention; Figure 5 This is a schematic diagram of the 360° environmental asset package of the Wei Family Mansion in its early intact state in some embodiments of the present invention; Figure 6 This is a schematic diagram of the arrangement of multiple people standing in some embodiments of the present invention.

[0043] Taking the three-act Minnan Kung Fu feature film "Never Surrender" as an example, the film has a standard theatrical length of 100 minutes, involves 15 main characters, 26 independent scenes, and is broken down into 1920 shots. The entire production process adopts the AI ​​feature film prompt generation and multi-dimensional consistent asset management method of this invention. The complete production process and specific operations are as follows: S1. Structured decomposition and top-level keyword anchoring of long-form scripts: For feature films with a runtime of no less than 90 minutes, the complete script is decomposed into narrative main line units, plot sub-line units, and three-level temporal units of scene-scene-shot, generating top-level anchoring keywords covering the entire film. Narrative Unit Decomposition: Extract 3 core narrative main lines (Wei Junlong's growth and awakening, the fight against human trafficking, and the confrontation with the corrupt officials) and 4 plot sub-lines (Wei Junlong and You Keying's emotional sub-line, the master-disciple martial arts inheritance sub-line, the change of stance of the Quanzhou martial arts community sub-line, and the fate of the Wei family sub-line), and strictly decompose them into a three-level chronological unit of 3 acts, 52 scenes, and 1920 shots. The first act (establishment) has 18 scenes and 620 shots, the second act (confrontation) has 24 scenes and 980 shots, and the third act (ending) has 10 scenes and 320 shots. Top-level anchoring keyword generation: Based on the breakdown results, generate top-level anchoring keywords covering the entire video and set fixed weight parameters, which cannot be modified throughout the entire process. The film's world-building is set in Quanzhou, Fujian, during the late Qing Dynasty and early Republic of China period. It features the martial arts world of five major schools of martial arts, the dark reality of human trafficking through collusion between officials and merchants, and the core themes of family, country, and chivalry, with a fixed weight of 25%. Core Character Lifecycle Settings: The background, appearance, personality, martial arts characteristics, and character arc of 15 main characters are set throughout their entire lifecycle, with a fixed weight of 30%; The film's overall visual style is characterized by the regional aesthetics of traditional red-brick houses in southern Fujian, a realistic and hardcore kung fu film feel, warm yellow sunlight contrasting with cool dark scenes, and 2K theatrical-quality image with film grain, all with a fixed weight of 20%. Narrative rhythm: Act 1: Lighthearted youthful narration → Act 2: Tense and suspenseful confrontation → Act 3: High-octane final battle and climax. The average shot length is 2.8 seconds, kung fu fight shots are 1.2 seconds, and emotional scenes are 4.5 seconds, with a fixed weight of 15%. General rules for lighting and sound effects: Daytime exterior shots with top lighting and hard shadow textures in southern Fujian style; nighttime interior shots with side lighting and contour lighting; ambient sounds of the city as the base; realistic martial arts sound effects; core segments paired with Nanyin music; fixed weight 10%.

[0044] S2. Full-link hierarchical prompt generation: Based on the top-level anchored prompts, hierarchical sub-prompt clusters are generated according to the full-link dimension of script-storyboard-camera-character-lighting-sound effects-rhythm. Each level of sub-prompt inherits the corresponding weight parameters of the top-level anchored prompts, and each level of sub-prompt is set with independent adjustable weights and association constraint rules. Script-level sub-cognitive keywords: Fully inherit the weight of the top-level anchored keywords, set scene-level narrative constraint rules, and allow ±10% independent adjustable weight; Storyboard level sub-keywords: Inherit the weight parameters of the parent level, set the scene-level plot and composition constraint rules, and open ±15% independent adjustable weights; Lens level sub-keyword: Inherit the weight parameters of the parent level, set the lens level image and camera movement constraint rules, and enable ±20% independent adjustable weight; Character, lighting, sound effects, and rhythm sub-keywords: Each is bound to its corresponding asset package, inherits the corresponding weight from the top-level anchor, and sets exclusive association constraint rules for each dimension to ensure that sub-keywords at each level are strongly bound to the top-level anchor, while also supporting fine-tuning of single shots.

[0045] S3. Character Multi-View Feature Asset Anchoring and Consistency Control: Build a multi-view feature asset package for anchoring each character in the script. Asset construction for the core protagonist, Wei Junlong: A set of descriptive terms for physical features was extracted from his biography, generating single-view white-background baseline images for both childhood (6-7 years old) and adulthood (18 years old). After manual verification, initial features were used to anchor the asset. These baseline images served as the input for graph-based constraints, such as... Figure 4 As shown, five viewpoints (front, back, left, right, and 3 / 4 side views) are generated and bound to the character's unique ID to generate two sets of basic multi-view feature anchoring asset packages. For the changes in the character's appearance and state of mind at three key nodes in the plot—family destruction, death at the execution ground, and underwater enlightenment—three state iteration version asset packages are generated based on the corresponding stage baseline map, with a noise reduction intensity of 0.55, and each is marked with the corresponding plot timestamp interval to achieve accurate matching between the character's state and the plot sequence. Full character asset coverage: For the remaining 14 main characters, including Dong Fei, A Jun, Wei Daxun, Liu Shi, Chen Yueshan, You Zhenkun, and You Keying, we have completed the extraction of facial features, generation of single-view baseline maps, and production of five-view atlases. We have also bound unique IDs to generate exclusive multi-view feature anchoring asset packages. Among them, female characters have generated corresponding version assets based on changes in Xunpu flower wreath headdress and clothing. We have also added pitch and tilt auxiliary view assets for martial arts characters to ensure consistency between cross-camera movements and appearance. Full-process asset retrieval: When generating any shot containing the corresponding character, the system automatically retrieves the asset package corresponding to the timestamp range of that character and injects it as an appearance constraint into the latent space control layer of the AI ​​generation model, thus eliminating the problem of character appearance drift from the root.

[0046] S4. Construction and Consistency Management of 360° Surround Video Assets for Each Scene: 360° surround video asset packages are constructed for each of the 26 independent scenes in the entire film to achieve spatial logic consistency across shots within the same scene. The specific operations are as follows: Basic asset construction: Extract the spatial description text of each independent scene in the script, generate a single-view scene baseline map, use the baseline map as the initial frame constraint input, generate a 360° surround video sequence with the scene center point as the axis, covering the horizontal azimuth angle 0°-360° and the pitch angle -30° to +60°, with a frame rate of 24fps and a surround time of 6 seconds, bind it with the scene's unique spatial identifier ID to form an asset package, and establish an association mapping with all scene indexes that appear in the script for that scene.

[0047] Multi-version state iteration: Generate corresponding version asset packages for scenarios where state changes occur in the storyline. Scene in the Wei Family Mansion: such as Figure 5 As shown, three asset packages are generated: the early stage of the story in an intact state, the middle stage of the story in a damaged state due to fighting, and the late stage of the story in a scorched earth state, with corresponding timestamp intervals marked. Among them, the scorched state version generates a dynamic video clip of the fire burning process, and the final frame is extracted as the base map to store the dynamic destruction event asset. "Riverside Fields" Scene: Generate two asset packages: a daytime scene of rice planting and martial arts practice in the early morning and a nighttime scene of a baby falling into the water in the rain, matching the day and night sequence and the atmosphere of the story. "Dock Warehouse" scenario: Generate two versions of asset packages: normal storage state and combat damage state, lock core spatial elements such as cages, stacks of boxes, and pig troughs to ensure spatial logic consistency.

[0048] Lens perspective matching call: When generating a single-lens scene, based on the azimuth and pitch parameters of the current lens, the matching frame is retrieved from the asset package. The azimuth matching tolerance is ±5° and the pitch matching tolerance is ±3°. The corresponding frame is extracted as the scene background layer to achieve spatial consistency between the foreground character and the background scene.

[0049] S5. Multi-character positioning complexity hierarchical control and final state frame inheritance: For scenes with multiple characters in the same frame throughout the film, hierarchical positioning control and cross-shot state inheritance are implemented. The specific operations are as follows: Complexity-based hierarchical control strategy: For scenes with N≤3 characters (such as Wei Junlong and A Jun escaping from the alley, Wei Junlong and You Keying facing each other in the ring, etc.), the text prompt word direct position control mode is enabled. Descriptor prompt words containing the characters' relative spatial positions, foreground and background relationships, and height differences are constructed and directly injected into the generation model. For scenes where 4≤N<6 (such as the scene of the five people confronting each other in the main hall of the Wei family), the reference position map auxiliary control mode is enabled, and the AI ​​generates a position preview map as the input for composition constraints. Scenes with N≥6 use a storyboard grouping and compositing mode: the core scene "Execution Ground" (N=12 characters in the same frame) is split into four sub-groups: the high platform officials group, the kneeling family members group, the executioners group, and the outer soldiers group, and the scenes are then composited after being generated separately; the scene "Warehouse Showdown" (N=18 characters in the same frame) uses a combination of keyframe constraints from the position preview video and grouping and compositing mode to ensure the spatial rationality of complex position scenes; Final state frame chain inheritance mechanism: With the shot level as the storage granularity, the final state frame of each shot is used as the anchor point asset of the standing state of the next shot, and as the initial frame reference and composition constraint input for the AI ​​generation of the next shot; for example, in the three consecutive fighting shots of "Jubao Street kicking the groin", the final state frame of the standing position of Wei Junlong, Shapi, and the martial arts group in the previous shot is used as the initial anchor point of the next shot, realizing strict inheritance of the standing state across shots and preventing the characters' spatial position from jumping.

[0050] S6. Enhanced Character Emotional Performance: For core emotional shots throughout the film, a performance enhancement process based on psychological-emotional mapping is executed. An example of the operation for core shots is as follows: The core scene of Wei Junlong witnessing his father's suicide at the execution ground: Action sequences, dialogue texts, and antecedent and consequential data of the behavior were extracted from the storyboard text. A psychological-emotion mapping knowledge base was queried to match the psychological change sequence of "shock-grief-despair-anger," obtaining the corresponding emotional representation description set: micro-expressions (pupil constriction, red eyes, tense jaw, trembling lips), body language (knees weakening, body leaning forward, clenched fists with white knuckles, uncontrollable shoulder trembling), and breathing rhythm (rapid breathing, sudden cessation of breathing followed by heavy breathing). The emotional representation description set was then fused with the action and dialogue data to generate performance-enhanced storyboard prompts. The performance emotion recognition model was then used for verification; the Valence dimension deviation was 0.12, and the Arousal dimension deviation was 0.10, meeting the threshold requirements. Wei Junlong's key scene of "underwater enlightenment and overcoming inner demons": Extract core data from the storyboard, match the psychological change sequence of "fear-hesitation-determination-relief", generate enhanced cue words containing micro-expressions and body language descriptions of the change of eyes from avoidance to determination and the body from curling up to stretching. The emotional verification results show that the matching degree meets the standard. Cross-camera emotional continuity control: For Wei Junlong's performance output in continuous shots throughout the film, an emotional intensity change curve was constructed from "mischievous youth" to "grief and despair" and then to "firm awakening". The match rate reached 92% when compared with the overall narrative rhythm spectrum of the film, which met the preset threshold requirements.

[0051] S7. Cross-unit consistency verification and correction for feature films: For the sub-prompt word clusters and generated images of all time units in the whole film, perform multi-dimensional consistency verification across scenes, scenes and shots. Specifically, it includes verification of character setting consistency, visual style consistency, narrative sequence consistency, emotional curve coherence, and lighting logic coherence. For sub-prompt words that do not meet the verification threshold, weight correction and content iteration are automatically triggered. For the sub-cue word clusters and generated images of all 1920 shots in the film, a multi-dimensional consistency check was performed across shots, scenes, and screens. The specific operations are as follows: Verification rules and threshold settings: Character consistency deviation threshold ≤ 5%, visual style consistency deviation threshold ≤ 8%, emotion curve matching threshold ≥ 90%, and simultaneously perform narrative sequence and lighting logic compliance checks; Automatic correction iteration: During the verification process, it was found that the facial feature deviation of Wei Junlong in 3 shots reached 7.2%. The system automatically locked the parent related prompt words of this sub-prompt word, and adjusted the weight of the character feature description in a step size of 5% per iteration based on the top-level anchor prompt word. After 2 iterations, the deviation value was reduced to 3.8%, which met the threshold requirement, and a correction log was generated and archived simultaneously. Only 2 shots in the entire film had logical deviations in the direction of the light source. After the system automatically adjusted the weight of the light and shadow prompt words, the correction was completed. Overall verification results: The average deviation of character consistency was 3.6%, the average deviation of visual style was 6.1%, the emotional curve matching degree was 91.5%, there were no errors in narrative sequence and lighting logic, and all met the verification threshold requirements.

[0052] S8, Industrial-grade rendering adaptation and dynamic optimization: The verified layered prompt word cluster is mapped to a standardized parameter set that can be recognized by the AI ​​film and television generation engine, and a mapping relationship between prompt word parameters and rendering engine interface is established. During the production of feature films, the weight parameters and constraint rules of subsequent time-series unit sub-prompt words are dynamically adjusted based on the effect data of the generated segments.

[0053] Rendering interface adaptation: The verified full-film layered prompt word cluster is mapped to a standardized parameter set that can be recognized by the AI ​​film and television generation engine. It is divided into 5 categories: shot parameter set, character model parameter set, rendering environment parameter set, audio parameter set, and timing control parameter set. Each parameter set has a unique mapping relationship with the corresponding layer prompt word, and can be directly connected to the native rendering interface of mainstream AI film and television generation engines. Dynamic optimization mechanism: After each scene-level unit is generated and its effect is evaluated, the unit's visual consistency, narrative matching degree, and rendering quality quantification data are automatically extracted. The weight of the sub-cue words in subsequent scene-level units is dynamically fine-tuned within a range of ±10%, without changing the fixed weight of the top-level anchor cue words. After the first act was generated, the anchoring weight of character features in the second act was fine-tuned by +8%, increasing the character consistency rate from 92% to 95%. Before the third act, the final battle unit, was generated, the weight of action sequence cue words was fine-tuned by +6% based on the martial arts rendering data from the first two acts, significantly improving the continuity of action shots.

[0054] It should be understood that the embodiments disclosed herein are not limited to the specific processing steps or materials disclosed herein, but should be extended to equivalent substitutions of such features as understood by those skilled in the art. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.

[0055] The term "embodiment" in this specification refers to a specific feature or characteristic described in connection with an embodiment that is included in at least one embodiment of the invention. Therefore, phrases or "embodiments" appearing in various places throughout the specification do not necessarily refer to the same embodiment.

[0056] Furthermore, the described features or characteristics can be incorporated into one or more embodiments in any other suitable manner. In the above description, specific details, such as thickness, quantity, etc., are provided to provide a comprehensive understanding of embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented without the aforementioned specific details or may be implemented using other methods, components, materials, etc.

Claims

1. A method for generating AI-powered feature film prompts and managing multi-dimensional assets in a consistent manner, characterized in that... include S1. Structured decomposition and top-level keyword anchoring of long-form scripts: For feature films with a runtime of no less than 90 minutes, the complete script is decomposed into narrative main line units, plot sub-line units, and three-level temporal units of scene-scene-shot, generating top-level anchoring keywords covering the entire film. S2. Full-link hierarchical prompt generation: Based on the top-level anchored prompts, hierarchical sub-prompt clusters are generated according to the full-link dimension of script-storyboard-camera-character-lighting-sound effects-rhythm. Each level of sub-prompt inherits the corresponding weight parameters of the top-level anchored prompts, and each level of sub-prompt is set with independent adjustable weights and association constraint rules. S3. Character Multi-View Feature Asset Anchoring and Consistency Control: Construct a multi-view feature asset package for anchoring each character in the script, specifically including: S31. Based on the character biography of each character, obtain the corresponding set of appearance feature description words, generate a single-view white background baseline image of each character, and store it as the initial feature anchoring asset after manual confirmation. S32. Using the initial feature anchoring assets as graph constraint input, generate at least four view graph sets for the role, and bind them with the role's unique identity ID to form a multi-view feature anchoring asset package; S33. When generating an AI screen containing any temporal unit of the character, automatically call the multi-view feature anchoring asset package of the character and input it as the character appearance constraint into the latent space control layer of the AI ​​generation model. S34. For events in the script where the character's status changes with timeline, such as age changes, injury, or appearance changes, generate new multi-view features to anchor the asset package after the status change based on the current valid asset package, and mark the valid timestamp interval. S35. Based on the timestamp interval of the current time sequence unit, automatically match and call the corresponding version of the character asset package; S4. Scene 360° Surround Video Asset Construction and Consistency Control: This involves constructing a 360° surround video asset package for each independent scene in the script, specifically including: S41. Extract the spatial description text of all independent scenes in the script and generate a single-view scene baseline map for each scene. S42. Using the single-view scene reference map as the initial frame constraint input, generate a 360° surround video sequence with the center point of the scene as the axis, and bind it with the scene's unique spatial identifier ID to form a 360° surround video asset package. S43. When generating AI character images for any shot in the scene, obtain the preset shooting parameters of the current shot, retrieve and extract frames with matching perspectives from the surrounding video asset package as the scene background layer. S44. Establish a mapping between the surround video asset package and all scenes in the script where the scene appears. When generating any shot in any scene, retrieve and automatically call the surround video asset package of the corresponding scene based on the current scene to ensure visual consistency of the same scene throughout the film. S45. For scenes in the script that change over time or are destroyed by external forces, generate a new surround video asset package after the state change based on the base frame of the current valid asset package, and mark the valid timestamp interval. S5. Multi-player positioning complexity hierarchical control and final state frame inheritance: Perform positioning control for scenes with multiple roles, specifically including: S51. Obtain the number N of characters appearing simultaneously in the current session and the interaction data between characters. Based on the value of N and the interaction complexity, automatically divide the positioning control strategy into at least two processing levels: Pre-build cue text containing descriptions of the relative space and position of each character. When N≤3, enable direct position control mode and adjust the positions of multiple characters through the cue text. Pre-construct static positioning reference images generated by AI, video keyframe positioning reference images generated by AI, manually laid-out positioning reference images, or real-life shooting positioning reference images. When N≥4, enable the auxiliary positioning mode and input one of the following as composition constraints into the AI ​​generation model to constrain the spatial distribution of multiple characters in the generated image. S52. When multiple characters undergo plot-driven position changes in a scene, the final frame of the previous shot is obtained as the anchor asset of the position state. This anchor asset of the position state is used as the initial frame reference input or composition constraint input, and the AI ​​generates the next shot to achieve strict inheritance of the position state across shots. S6. Enhanced Character Emotional Performance: Enhances the emotional nuances of character performances within a single shot, specifically including: S61. Extract action sequence data, dialogue text data, behavioral antecedent data, and behavioral consequence data from single-shot storyboard text; S62. Query the preset psychological-emotion mapping knowledge base to obtain a set of emotional representation descriptions that match the psychological changes before and after the behavior; S63. Integrate the emotional representation description set with action sequence data and dialogue text data to generate performance-enhanced storyboard prompts; S64. Input the performance-enhanced storyboard prompts generated in step S63 into the AI ​​generation model to obtain preliminary performance footage. Call the performance emotion recognition model to classify the preliminary performance footage by emotion label and score the emotion intensity. If the recognition result deviates from the emotion representation description set set in step S62 by more than the preset threshold, the weight parameters of the micro-expression or body language description in the prompts will be automatically adjusted and regenerated until the verification is passed. S65. For the performance output of the same character in consecutive shots, extract the performance emotion recognition results of each shot, construct the emotional intensity change curve of the character in the scene or act, compare the emotional intensity change curve with the pre-set emotional rhythm spectrum of the script, and perform the supplementary enhancement operation of steps S63-S64 for shots that deviate from the pre-set curve to ensure the emotional progression logic of the character's performance in cross-shot narrative. S7. Cross-unit consistency verification and correction for feature films: For the sub-prompt word clusters and generated images of all time units in the whole film, perform multi-dimensional consistency verification across scenes, scenes and shots. Specifically, it includes verification of character setting consistency, visual style consistency, narrative sequence consistency, emotional curve coherence, and lighting logic coherence. For sub-prompt words that do not meet the verification threshold, weight correction and content iteration are automatically triggered. S8, Industrial-grade rendering adaptation and dynamic optimization: The verified layered prompt word cluster is mapped to a standardized parameter set that can be recognized by the AI ​​film and television generation engine, and a mapping relationship between prompt word parameters and rendering engine interface is established. During the production of feature films, the weight parameters and constraint rules of subsequent time-series unit sub-prompt words are dynamically adjusted based on the effect data of the generated segments.

2. The method for generating AI-powered feature film prompts and managing multi-dimensional assets according to claim 1, characterized in that, In step S34, the specific steps for handling events related to the character's changing status over time, such as aging, injury, or appearance, are as follows: S341. Read the currently valid multi-view feature anchored asset package for this role; S342. Obtain the set of target feature descriptive words that describe the state transition event; S343. Input the single-view baseline map in the current valid asset package and the target feature descriptor set into the AI ​​image generation model to generate a new single-view baseline map after the state change. S344. Perform the multi-view expansion operation of steps S31-S32 on the new single-view baseline map to generate a new multi-view feature anchored asset package after the state change. S345. Bind the new asset package to the character ID and mark its valid timestamp range in the corresponding scene or act of the script.

3. The method for generating AI-powered feature film prompts and managing multi-dimensional assets according to claim 2, characterized in that, In step S35, based on the current time sequence unit's position in the script timeline, the system automatically matches and calls the character's valid multi-view feature anchor asset package version within the timestamp interval to ensure that the character's appearance is strictly synchronized with the plot development.

4. The method for generating AI-powered feature film prompts and managing multi-dimensional assets according to claim 1, characterized in that, In step S45, the specific operational steps for the scene in the script that undergoes natural changes or external destruction over time are as follows: S451. Read the base frame from the currently valid 360° surround video asset package for this scene; S452. Obtain the set of target feature descriptive words that describe the natural change or external destruction event; S453. Input the baseline frame and the target feature descriptor set into the AI ​​image generation model to generate a new single-view scene baseline map after the state change. S454. Perform the surround video extension operation of steps S41-S42 on the new single-view scene baseline map to generate a new 360° surround video asset package after the state change. S455. Bind the new asset package to the scene ID and mark its valid timestamp range in the corresponding scene or act of the script.

5. The method for generating AI-powered feature film prompts and managing multi-dimensional assets according to claim 4, characterized in that, When an external force damages the scene, the following steps are performed: The system invokes a pre-defined AI video generation model to generate video clips depicting the dynamic process of the destructive event. The stable final state frame after the destruction is completed is extracted from the video clip of the destruction process and used as the base image for generating the new single-view scene reference map in step S45; The video clip of the destruction process is stored as a dynamic destruction event asset for the scene. When generating subsequent continuous shots of the scene after destruction, the intermediate state frames are extracted as the scene background layer to maintain the consistency of the distribution of debris and fragments across shots.

6. The method for generating AI-powered feature film prompts and managing multi-dimensional assets according to claim 1, characterized in that, For the auxiliary control mode in step S51: Call the preset AI image generation model, input the position description text containing the number of characters and their rough positional relationships, and generate a static position reference image; The preset AI video generation model is invoked. With the functions of dialogue generation and complex action generation turned off, only a panoramic video of the spatial distribution of multiple characters is generated, and key frames are extracted from it as a reference map for the video key frame positions. Use image processing tools to manually arrange the preset character materials to generate a layout and positioning reference diagram; We obtained real-life footage of the pre-rehearsed positioning as a reference image for the actual shooting.

7. The method for generating AI-powered feature film prompts and managing multi-dimensional assets according to claim 1, characterized in that, In step S7, cross-cell consistency verification specifically includes: Character consistency check: Based on the character's entire life cycle, check the deviation values ​​of appearance, clothing, age, personality and behavioral logic of characters across shots, with the deviation threshold set to no more than 5%; Visual style consistency check: Based on the overall visual tone of the film, check the deviation values ​​of color system, art style and picture quality of cross-scene shots, with the deviation threshold set to no more than 8%; Narrative timeline consistency verification: Based on the overall timeline of the film, verify the timeline logic of cross-scene plots, the continuity of props, and the compliance of causal relationships between events; Emotional curve coherence verification: Based on the overall narrative rhythm spectrum of the entire film, verify the matching degree between the fluctuation of the emotional curve across scenes and the preset overall spectrum, with the matching degree threshold set to no less than 90%; Lighting and shadow logic consistency verification: Based on the overall lighting and shadow rules of the whole film, verify the logical consistency of light source direction, day and night sequence, and ambient lighting across shots.

8. The method for generating AI-powered feature film prompts and managing multi-dimensional assets according to claim 1, characterized in that, In step S7, the automatic correction iteration is as follows: when the verification result of the sub-prompt word exceeds the preset deviation threshold, the system automatically locks the corresponding parent-level related prompt word of the sub-prompt word, and automatically adjusts the weight parameters and content description of the sub-prompt word based on the top-level anchor prompt word until the verification result meets the threshold requirements, and generates a correction log for archiving.

9. The method for generating AI-powered feature film prompts and managing multi-dimensional assets according to claim 1, characterized in that, In step S8, the standardized parameter set includes shot parameter set, character model parameter set, rendering environment parameter set, audio parameter set, and timing control parameter set. Each parameter set has a unique mapping relationship with the corresponding level of prompt words, and supports direct access to the native rendering interface of mainstream AI film and television generation engines.

10. The method for generating AI-powered feature film prompts and managing multi-dimensional assets according to claim 1, characterized in that, In step S8, dynamic adjustment specifically involves: during the feature film generation process, after each segment-level unit is generated and its effect is evaluated, the quantitative data of the unit's visual consistency, narrative matching degree, and rendering quality are automatically extracted. Based on this data, the weight parameters of the sub-cue words in subsequent segment-level units are dynamically fine-tuned within a range of ±10%. The fine-tuning process does not change the fixed weight parameters of the top-level anchor cue words.