A video generation method and system based on face consistency constraint
By generating and freezing the Face ID of the target person, and combining it with the storyboard description information to generate video, the problems of poor facial consistency and low generation efficiency are solved, and the controllability and high-quality output of video generation are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU YUZE TECHNOLOGY CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-06-05
AI Technical Summary
In existing technologies, facial consistency is poor, video generation quality is difficult to control precisely, and human features are highly coupled with variable elements, resulting in unstable identity features and low generation efficiency.
By generating a baseline face image of the target character, extracting the unique identifier Face ID, and freezing the first prompt word, baseline face image, and Face ID, and combining the storyboard description information in the plot text to generate a storyboard image package, and finally synthesizing the video, the consistency of the character's identity and the controllability of the generation process are ensured.
It achieves stability of character identity and controllability of the generation process in video generation, supports local regeneration, reduces computational resource consumption, and improves generation efficiency and quality.
Smart Images

Figure CN122160588A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence-generated content technology, and in particular to a video generation method and system based on face consistency constraints. Background Technology
[0002] With the development of image and video generation models, solutions for generating character images and video content based on text prompts have been widely used. However, in practical applications, existing solutions still have the following technical problems.
[0003] Poor facial consistency: When generated multiple times, in different shots, with different styles, or in different scenes, the facial features and face shape of the same person are prone to drift, making it difficult to maintain a stable identity; Video generation quality is difficult to control precisely: Existing solutions usually generate images or videos directly from text, and the generated results are difficult to control. When users need to modify a shot or part of the scene, they usually have to regenerate the complete video, which consumes a lot of computing power and is inefficient.
[0004] In other words, existing solutions typically mix elements such as character features, scene, and actions into a single input. For example, if a single prompt word is used to generate multiple times, the results of these multiple generation are highly different and have poor repeatability. Moreover, this design leads to a high degree of coupling between identity features and variable elements, and any adjustment of parameters may affect the consistency of the character.
[0005] Currently, no effective solution has been proposed for improving the video generation quality of large language models in related technologies. Summary of the Invention
[0006] This application provides a video generation method and system based on face consistency constraints, which at least addresses the problem of how to improve the video generation quality of large language models in related technologies.
[0007] In a first aspect, embodiments of this application provide a video generation method based on face consistency constraints, the method comprising: Based on the first cue word, a baseline face image of the target person is generated using a large language model, wherein the first cue word is only used to describe the role information of the target person; Based on the reference face image, a unique identity identifier Face ID for the target person is generated, and the first prompt word, the reference face image, and the unique identity identifier Face ID for the target person are frozen; Based on the scene description information corresponding to each scene in the plot text, combined with the frozen unique identity identifier FaceID and the second cue word, the corresponding scene image package is generated through a large language model; Based on each storyboard image package, a corresponding video segment is generated, and the video segments are synthesized into the final video according to the storyboard order.
[0008] In some embodiments, the method includes: The first prompt word is only used to describe the character's role information, wherein the role information includes basic setting information and facial structure information. The facial features in the facial structure information are divided into multiple mutually exclusive sub-dimensions. Only one value is allowed for the same sub-dimension. Different sub-dimensions only take effect within their respective facial features and do not combine across facial features. The first prompt does not contain descriptions of scene information, style information, or plot information.
[0009] In some embodiments, generating a unique identity identifier (Face ID) for the target person based on the reference face image includes: Feature extraction is performed on the reference face image to obtain the facial feature vector of the target person; Based on the facial feature vector, a unique identity identifier, Face ID, is constructed and generated for the target person.
[0010] In some embodiments, freezing the first prompt word, the reference face image, and the unique identifier Face ID of the target person includes: The first prompt word, the reference face image, and the unique identifier FaceID of the target person are frozen, and the first prompt word, the reference face image, and the unique identifier FaceID cannot be modified after freezing; That is, if it is necessary to modify the unique identity identifier Face ID of the target person, the first prompt word, the reference face image and the unique identity identifier Face ID must all be deleted and recreated, and there is no intermediate state.
[0011] In some embodiments, based on the scene description information corresponding to each scene in the plot text, combined with the frozen unique identifier Face ID and the second cue word, the corresponding scene image package is generated through a large language model, including: Obtain the plot text for video generation, and perform storyboard decomposition on the plot text to obtain the storyboard description information corresponding to each storyboard in the plot text; Under the face consistency constraint of the frozen unique identifier Face ID, the corresponding segment image package is generated by a large language model based on the segment description information and the second prompt word for each segment.
[0012] In some embodiments, under the face consistency constraint of the frozen unique identifier Face ID, based on the segment description information and the second cue word for each segment, the corresponding segment image package is generated through a large language model, including: The unique identity identifier Face ID of the character referenced in each storyboard is bound to the corresponding storyboard; Under the face consistency constraint of the unique identifier Face ID of the referenced person, a segmentation image package corresponding to each segment is generated through a large language model based on the segmentation description information and the second prompt word. The second prompt word is used to describe non-face elements in the image generation.
[0013] In some embodiments, after freezing the first cue word, the reference face image, and the unique identifier Face ID of the target person, the method includes: Enter the target person's role asset expansion process, and generate a multi-view expansion map of the target person based on the target person's unique identity identifier FaceID. The multi-view expansion map includes a front view expansion map, an oblique side view expansion map, and a front side view expansion map.
[0014] In some embodiments, under the face consistency constraint of the frozen unique identifier Face ID, generating the corresponding segment image package through a large language model based on the segment description information and the second cue word for each segment further includes: If a storyboard references a multi-view extended map, then under the face consistency constraint of the frozen unique identifier Face ID corresponding to the multi-view extended map, a storyboard image package corresponding to each storyboard is generated through a large language model based on the storyboard description information and the second prompt word. The storyboard images in the storyboard image package are saved in the order of the storyboard and the order of the numbers within the storyboard.
[0015] In some embodiments, after generating a baseline face image of the target person using a large language model based on a first cue word, the method includes: The quality of the generated reference face image is checked. If no face is detected in the reference face image, or multiple faces are detected, or key facial features are not recognizable, the generation of the reference face image is determined to have failed.
[0016] In a second aspect, embodiments of this application provide a video generation system based on face consistency constraints. The system is used to execute the method described in the first aspect above. The system includes a role identity definition module, a storyboard image generation module, and a video generation module. The role identity definition module is used to generate a baseline face image of the target person based on a first prompt word using a large language model, wherein the first prompt word is only used to describe the role information of the target person; The role identity definition module is used to generate a unique identity identifier Face ID for the target person based on the reference face image, and freeze the first prompt word, the reference face image and the unique identity identifier Face ID for the target person; The storyboard image generation module is used to generate a corresponding storyboard image package based on the storyboard description information corresponding to each storyboard in the plot text, combined with the frozen unique identity identifier Face ID and the second prompt word, through a large language model. The video generation module is used to generate corresponding video segments based on each storyboard image package, and to synthesize the video segments into the final video according to the storyboard order.
[0017] Compared to related technologies, this application provides a video generation method and system based on face consistency constraints. The method generates a reference face image of a target character using a large language model based on a first prompt word, where the first prompt word is only used to describe the character's role information. Based on the reference face image, a unique identity identifier (Face ID) for the target character is generated, and the first prompt word, reference face image, and Face ID are frozen. Based on the scene description information corresponding to each scene in the script text, combined with the frozen Face ID and a second prompt word, a corresponding scene image package is generated using the large language model. A corresponding video segment is generated based on each scene image package, and the video segments are synthesized into a final video according to the scene order. This method utilizes the Face ID of the reference face image to isolate the character's facial identity from the video generation process, effectively ensuring the consistency of the character's identity in the generated video. Furthermore, splitting the script into segments for generation makes the video generation process structured, controllable, and reversible, solving the problem of how to improve the video generation quality using a large language model.
[0018] Furthermore, based on the specific technical solutions provided in the above embodiments, it can be seen that the core architectural advantages of this application are reflected in the following aspects: Independent identity layer architecture: An independent identity layer is established at the very beginning of the generation process; identity information is solidified into an immutable identifier through a freezing mechanism, ensuring identity stability across tasks and shots from an architectural perspective.
[0019] Decoupling design: Separating fixed identity characteristics from variable scene, action, and emotional information; adjustments to scene and action do not affect the character's identity, improving the controllability and flexibility of the generation process.
[0020] Compliance assurance: Fictional characters are generated through structured semantics, without relying on real people's photos, thus avoiding portrait rights risks from the technical source.
[0021] Storyboard management mechanism: Supports partial regeneration, so that when modifying a single shot, it is not necessary to regenerate the entire video, which improves iteration efficiency and reduces computational resource consumption.
[0022] Asset reusability: Establish a reusable role asset library, allowing the same role to be reused in multiple projects, effectively reducing creation costs.
[0023] Multi-view adaptation: Pre-generate character assets from multiple perspectives to adapt to different camera angle requirements and improve the generation quality of scenes such as side shots. Attached Figure Description
[0024] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart illustrating the steps of a video generation method based on face consistency constraints according to an embodiment of this application. Figure 2 This is a schematic diagram of the internal structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.
[0026] Obviously, the accompanying drawings described below are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without any inventive effort. Furthermore, it is understood that although the efforts made in this development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, any changes to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.
[0027] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0028] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following objects are in an "or" relationship. The terms "first," "second," and "third" used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.
[0029] This application provides a video generation method based on face consistency constraints. Figure 1 This is a flowchart illustrating the steps of a video generation method based on face consistency constraints according to an embodiment of this application, as follows: Figure 1 As shown, the method includes the following steps: Step S102: Based on the first prompt word, generate a reference face image of the target person through a large language model, wherein the first prompt word is only used to describe the role information of the target person; Specifically, in step S102, the first prompt word is a "face prompt word" (FACE_PROMPT), which is only used to describe the character's role information. This role information includes basic setting information (such as gender, skin color, age, etc.) and facial feature structure information (such as eyes, nose, lips, chin / jawline, eyebrows). The facial features in the facial feature structure information are divided into multiple mutually exclusive sub-dimensions. Only one value is allowed for the same sub-dimension, and different sub-dimensions only take effect when combined within their respective facial features and do not combine across facial features. The first prompt word does not contain descriptions of scene information, style information, or plot information and is not directly used for plot text or video generation.
[0030] It should be noted that FACE_PROMPT adopts a "three-layer semantic structure": facial feature dimension (organ level) -> sub-dimension (internal breakdown of facial features) -> value (specific concept). Within the same facial feature dimension: different sub-dimensions are allowed to be combined; each sub-dimension is only allowed to select one value. Between different facial features: sub-dimensions cannot be used interchangeably; the system does not perform cross-facial feature semantic inference or appearance completion. It can be seen that the compound semantics are decomposed into multiple sub-dimension mappings at the input stage. Mutually exclusive constraints ensure that conflicting descriptions are structurally avoided at the input stage, prohibiting cross-facial feature inference (e.g., not inferring eyebrow shape from eye shape; not inferring chin contour from lip thickness). This construction rule is used to transform the description of a person's appearance from unstructured prompts into controllable structured semantics, thereby improving the consistency, stability, and repeatability of the baseline face and FaceID.
[0031] In short, the technical challenge of step S102 lies in the ambiguity, subjectivity, and conflict inherent in natural language descriptions of facial features (e.g., simultaneously requiring "large eyes" and "slender eyes"), leading to uncontrollable generation results. Therefore, step S102 employs a three-layer semantic structure: facial feature dimension → sub-dimension → specific value; only one value is allowed for each sub-dimension, structurally avoiding semantic conflicts; sub-dimensions of different facial features do not overlap, avoiding uncontrollable associative inferences; thus ensuring the clarity and consistency of the input semantics, improving the quality and stability of the baseline face generation.
[0032] It should be further explained that this video generation method can be divided into five layers according to function and data boundaries: Face Identity Layer – generates a baseline face and extracts / freezes Face ID; Character Asset Layer – generates multi-view extended assets and reusable sets around Face ID; Shot Layer – generates a script and automatically segments the video according to plot actions and duration nodes; Shot Image Layer – generates multiple consecutive images for each shot to form a shot package; Video Layer – generates video clips in shot / number order and synthesizes them into a video, supporting regeneration by shot.
[0033] Step S104: Based on the reference face image, generate a unique identity identifier Face ID for the target person, and freeze the target person's first prompt word, reference face image, and unique identity identifier Face ID; Specifically, in step S104, the unique identity identifier Face ID is composed of the facial feature vector (embedding) extracted from the reference face image. The first prompt word, reference face image, and unique identity identifier Face ID of the target person are frozen. The system sets a read-only status for the frozen first prompt word, reference face image, and unique identity identifier Face ID. When subsequent operation requests access, the system automatically intercepts the modification operation. That is, if it is necessary to modify the unique identity identifier Face ID of the target person, the first prompt word, reference face image, and unique identity identifier Face ID must all be deleted and recreated. There is no intermediate state.
[0034] In short, the technical challenge of step S104 lies in the fact that identity information is easily affected by other parameters during the generation process, causing the character's features to drift across shots and scenes. Therefore, step S104 extracts the facial feature vector of the reference face as a unique identity identifier; after user confirmation, this identity identifier is frozen to prevent modification. In all subsequent generation processes, this identity identifier serves as a strong constraint, ensuring the uniqueness and stability of the identity throughout the entire generation process from an architectural perspective, and avoiding identity drift caused by parameter changes. Specifically, the frozen Face ID is converted into a feature constraint, which serves as a continuous input during image generation. A feature matching mechanism ensures that the generated facial features of the person are consistent with the identity features represented by the Face ID.
[0035] It should be noted that steps S102 and S104 above belong to the Face Identity Layer. The Face Identity Layer is at the very beginning of the generation process. It generates a stable, unique, and reusable Face ID by combining structured facial semantics (first cue words) with a baseline face image, serving as the identity constraint basis for all subsequent generation stages. Specifically: Core Identity Definition. The unique identity identifier Face ID is composed of a facial feature vector (embedding) extracted from a baseline face image. Users are not allowed to upload real faces as the source of Face ID; and the role identity definition layer does not provide "modifiable logical identity numbers" such as role_id and role state machine to the outside world; identity changes are only allowed to be "deleted and rebuilt".
[0036] The first cue word for structured facial sense semantics is (FACE_PROMPT). Facial senses are broken down into multiple mutually exclusive sub-dimensions. Each sub-dimension is allowed only one value, and combinations between different sub-dimensions are permitted. Sub-dimensions only apply within their respective facial senses and cannot be combined across different facial senses. Therefore, compound semantics are decomposed into multiple sub-dimension mappings during the input phase. This mutual exclusion constraint ensures that conflicting descriptions are structurally avoided during the input phase.
[0037] The set of required fields for the first prompt word. Table 1 is an example table of the minimum set of required fields for the first prompt word according to an embodiment of this application, as shown in Table 1.
[0038] Table 1
[0039] Completion determination and freezing. The role identity definition layer uses explicit user confirmation as the sole criterion for determining the completion of FACE_PROMPT. Confirmation adopts a two-layer confirmation mechanism. After user confirmation, FACE_PROMPT, the baseline face image, and FaceID are frozen. Once frozen, none of the three can be modified. If identity information needs to be modified, the role must be deleted and recreated. There is no intermediate state.
[0040] Baseline face quality verification and resource processing. After the baseline face is generated, a quality verification is performed. Failure occurs if any of the following conditions are met: no face detected, multiple faces detected, severe occlusion causing key facial features to be unrecognizable, obvious distortion, or generation error. No resources are deducted for a failed quality verification task; the cost of the failed task is refunded, and the user is prompted to regenerate the baseline face. If the user wants to discard a baseline face that has passed the quality verification, they must go to the role library, delete the current record, and recreate the role. Modifying identity information on the current object is not allowed.
[0041] Output and Inheritance Rules. The only output of the character identity definition layer is Face ID, which serves as a global identity constraint signal. It continues to be a required input in the subsequent character asset extension layer (S2), storyboard generation layer (S3), storyboard image generation layer (S4), and video generation layer (S5). Each layer must call Face ID for identity constraint when performing generation operations, and no identity-related information may be modified. Identity consistency is guaranteed by Face ID throughout the process.
[0042] Furthermore, the first prompt word (FACE_PROMPT) has a semantic constraint mechanism: the system performs semantic analysis on user input, and when it detects features such as real people or public figures' names, it rejects the input and returns a prompt message, guiding the user to use structured facial semantics for input. In 2D usage scenarios, the FACE_PROMPT constraint can be omitted, allowing users to upload a single baseline image to enter the 2D character asset process; in 3D realistic scenarios, structured FACE_PROMPT must be used, and the minimum required field set must be met; otherwise, the process will not proceed to baseline face image generation and Face ID freezing.
[0043] After step S104, the method includes step S105, which enters the target person's role asset expansion process, and generates a multi-view expansion map of the target person based on the target person's unique identity identifier Face ID. The multi-view expansion map includes a frontal view expansion map, an oblique side view expansion map, and a frontal side view expansion map.
[0044] In short, the technical challenge of step S105 lies in the fact that when the same person appears in the video from different angles, a single-angle reference cannot provide sufficient spatial information, leading to a decrease in the quality of non-frontal shots. Therefore, step S105 pre-generates character images from multiple perspectives (such as frontal, side, and 45-degree side views) based on the character's identity; automatically selects the most suitable perspective image as a reference according to the camera angle requirements; and reduces the amplitude of perspective transitions through multi-view coverage. This significantly improves the generation quality of shots from different angles and reduces facial distortion caused by perspective mismatch.
[0045] It should be noted that step S105 belongs to the Character Asset Layer. After the baseline face image is generated and Face ID is frozen, the character asset expansion process must be entered; otherwise, if the core data is insufficient, there will be no assets available. Specifically: Character asset content. Character assets should include at least: Face ID, base face, generation record (time / model version / parameter summary), asset status (available / deletable), multi-view extended map, feature vector (embedding), tags / notes, and corresponding prompts for each map.
[0046] Multi-view expansion and qualification criteria. Generate multi-view expansion maps to support subsequent head turning, movement, and camera changes. The multi-views should include at least the following categories: frontal view (0°), oblique side view (approximately ±45°), and frontal side view (approximately ±90°). The qualification criteria for the multi-view expansion map should include at least the following: single face, clear and unobstructed facial features, consistent with Face ID, and valid viewpoint.
[0047] Regeneration Mechanism. When the quantity or quality of multi-view expanded maps is insufficient to meet the minimum usability requirements, a "partial pass + automatic regeneration" processing method is allowed: retained qualified maps are kept, and failed views continue to be regenerated until the minimum usability requirements are met or the maximum number of regeneration attempts is reached; and after multiple failures, users are allowed to manually regenerate or abandon and delete the character's assets.
[0048] Deletion and Reconstruction. If a character is deleted from the character library, a soft deletion is performed by default: the character asset / Face ID is marked as unreferenceable and prohibited from being used in newly generated tasks; if historical references (storyboards / video tasks) still exist, the reference chain is preserved and physical deletion is not performed. Physical deletion and resource reclamation are allowed when the reference count is 0.
[0049] External API references: When downstream modules reference role assets, the required input is Face ID, and the optional input is view_set_id (extended view set); apart from the above API, embedding, underlying path, or specific model parameters are not exposed to maintain implementation substitutability. The Agent is only responsible for whether to pass view_set_id and does not change the API.
[0050] Step S106: Based on the scene description information corresponding to each scene in the plot text, combined with the frozen unique identifier Face ID and the second cue word, the corresponding scene image package is generated through the big language model; Step S106 specifically includes the following steps: Step S1061: Obtain the plot text for video generation, and decompose the plot text into storyboards to obtain the storyboard description information corresponding to each storyboard in the plot text. Step S1062: Under the face consistency constraint of the frozen unique identifier Face ID, the corresponding segment image package is generated by the large language model based on the segment description information and the second prompt word of each segment.
[0051] Optionally, in step S1062, the unique identity identifier Face ID of the character referenced in each storyboard is bound to the corresponding storyboard; under the face consistency constraint of the frozen unique identity identifier Face ID of the character referenced, a storyboard image package corresponding to each storyboard is generated through a large language model based on the storyboard description information and the second prompt word, wherein the second prompt word is used to describe the non-face elements in the image generation.
[0052] Optionally, in step S1062, if the storyboard references a multi-view extended map, then under the face consistency constraint of the frozen unique identifier Face ID corresponding to the multi-view extended map, based on the storyboard description information and the second prompt word, a storyboard image package corresponding to each storyboard is generated through a large language model, wherein the storyboard images in the storyboard image package are saved in the order of the storyboard and the order of the numbers within the storyboard.
[0053] In short, the technical challenge of step S106 lies in the fact that identity features, scene, action, and emotion are mixed in a single input, causing mutual interference, and modifying the scene may affect the character's features. Therefore, step S106 divides the input into two layers of prompts: an identity layer (fixed character features) and a scene layer (variable scene, action, and emotion). These two layers act independently in the generation process. The identity layer imposes constraints through identity identifiers, while the scene layer controls scene changes. This allows for free adjustment of scene, action, and emotion without affecting the character's identity, significantly improving controllability.
[0054] It should be noted that step S106 belongs to the Shot Layer and the Shot Image Layer. The methods of obtaining the plot text include temporarily generating scripts (system-generated) and using existing scripts (provided by the user), automatically dividing the plot into scenes according to plot actions and duration nodes, and generating multiple consecutive images for each scene to form a scene package. That is, step S106 includes two sub-level processing flows: (1) Shot Layer: responsible for automatically dividing the plot text according to scene changes, action switching, and duration nodes, generating structured descriptive information (including scene description, action description, emotion description, camera language, etc.) for each scene, and generating the second prompt word corresponding to each scene. (2) Shot Image Layer: responsible for generating 2-8 consecutive keyframe images for each scene based on the descriptive information of each scene, combined with the frozen Face ID and the second prompt word, through a large language model, and organizing these images into a scene image package to express the action continuity and screen progression within the scene. Each storyboard image package corresponds one-to-one with a storyboard, and the images in each storyboard image package are numbered according to their generation order. Specifically: Automatic storyboarding. The script is divided into multiple storyboards based on plot actions and time nodes; the storyboards have a fixed order and cannot be arbitrarily adjusted; each storyboard includes: shot description (scene, action, emotion, shot size / camera position, etc.) and time / pacing node information; each storyboard is bound to the character it references (using Face ID as the identity input, and modifiable identity information cannot be introduced).
[0055] Storyboard Package Rules. A one-to-one correspondence exists between storyboard image packages and storyboards: one storyboard corresponds to one storyboard image package. The storyboard image package is used to express the camera movement and action continuity of that storyboard. Each storyboard image package contains N consecutive images, where N can be 2-8, preferably 3-5. The images are numbered in the order they were generated. Images within a storyboard image package are not allowed to be reused across storyboards to ensure the manageability and traceability of the storyboard structure.
[0056] The organization and invocation boundaries of cue messages. Cue messages in the storyboard image package are used to express scene, action, camera movements, and non-face elements (such as clothing, environment, props, lighting, and atmosphere); a single image only invokes cue messages for the "non-face portion" (i.e., the second cue word) during generation; face consistency constraints come from Face ID and the extended viewpoint set `iew_set_id`. LoRA can optionally be loaded as a capability plugin, but it must not bypass the identity constraints of Face ID.
[0057] Storyboard Image Verification and Saving. Consistency verification of storyboard images can be performed based on Face ID. Anomaly types include, but are not limited to: no faces, multiple faces, severe occlusion, obvious distortion, or equivalent anomalies. Storyboard images are saved in the order of the storyboard scenes and their corresponding numbers within the storyboard image package.
[0058] Step S108: Generate corresponding video segments based on each storyboard image package, and synthesize the video segments into the final video according to the storyboard order.
[0059] It should be noted that step S108 belongs to the video generation layer. Video clips are generated in sequence according to the storyboard / number and then composited into a final cut. Specifically: Video generation input and order rules. The storyboard package is consumed in the order of the storyboard shots, and video clips are generated in the order of image numbers within the storyboard shots. The video generation layer does not participate in the generation of prompts and does not change the storyboard shot order. It supports regenerating video clips by storyboard shot and compositing them into a single video. To ensure the integrity of the narrative structure, the number of storyboard shots must be no less than M, where M can be 2-5, with M=3 being the preferred value.
[0060] The method and steps provided in this application generate and freeze Face IDs at the forefront, ensuring the uniqueness, stability, and reusability of character identities and reducing identity drift across tasks / cameras. By expanding character assets and covering multiple perspectives, the generation stability under scenarios involving head turns, movement, and camera changes is improved. Through continuous images and sequential numbering of storyboard packages, the storyboard structure becomes manageable and traceable, and local regeneration is supported. This method utilizes Face IDs from a reference face image to isolate character facial identities from the video generation process, effectively ensuring the consistency of character identities in the generated video. Furthermore, by splitting the script into segments for generation, the video generation process becomes structured, controllable, and rollback-capable, solving the problem of how to improve the video generation quality of large language models.
[0061] It should be noted that the steps shown in the above process or in the flowchart of the accompanying figures can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0062] This application provides a video generation system based on face consistency constraints. The system is used to execute the method proposed in the above embodiments. The system includes a role identity definition module, a storyboard image generation module, and a video generation module. The role identity definition module is used to generate a baseline face image of the target person based on the first prompt word using a large language model. The first prompt word is only used to describe the role information of the target person. The role identity definition module is used to generate a unique identity identifier FaceID for the target person based on the reference face image, and freeze the target person's first prompt word, reference face image and unique identity identifier FaceID; The storyboard image generation module is used to generate corresponding storyboard image packages based on the storyboard description information corresponding to each storyboard in the plot text, combined with the frozen unique identity identifier Face ID and the second cue word, through a large language model. The video generation module is used to generate corresponding video segments based on each storyboard image package, and to synthesize the video segments into the final video according to the storyboard order.
[0063] The system provided in this application embodiment realizes the separation of facial identity from video generation process by using Face ID of reference face image, effectively ensuring the consistency of facial identity in generated video. Furthermore, by splitting the script into segments for generation, the video generation process becomes structured, controllable, and rollbackable, thus solving the problem of how to improve the video generation quality of large language models.
[0064] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.
[0065] This embodiment provides an electronic device including a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0066] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0067] Optionally, the electronic device may further include a processor, memory, network interface, display screen, and input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a video generation method based on face consistency constraints. The display screen may be a liquid crystal display (LCD) or an e-ink display. The input device may be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.
[0068] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.
[0069] Furthermore, in conjunction with the video generation method based on face consistency constraints in the above embodiments, this application embodiment can provide a storage medium for implementation. This storage medium stores a computer program; when executed by a processor, the computer program implements any of the video generation methods based on face consistency constraints in the above embodiments.
[0070] In one embodiment, Figure 2 This is a schematic diagram of the internal structure of an electronic device according to an embodiment of this application, such as... Figure 2 As shown, an electronic device is provided, which can be a server, and its internal structure diagram can be as follows. Figure 2As shown, the electronic device includes a processor, a network interface, internal memory, and non-volatile memory connected via an internal bus. The non-volatile memory stores the operating system, computer programs, and a database. The processor provides computing and control capabilities, the network interface communicates with external terminals via a network, the internal memory provides an environment for the operating system and computer programs to run, the computer programs are executed by the processor to implement a video generation method based on face consistency constraints, and the database stores data.
[0071] Those skilled in the art will understand that Figure 2 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0072] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0073] Those skilled in the art should understand that the technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0074] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A video generation method based on face consistency constraints, characterized in that, The method includes: Based on the first cue word, a baseline face image of the target person is generated using a large language model, wherein the first cue word is only used to describe the role information of the target person; Based on the reference face image, a unique identity identifier Face ID for the target person is generated, and the first prompt word, the reference face image, and the unique identity identifier Face ID for the target person are frozen; Based on the scene description information corresponding to each scene in the plot text, combined with the frozen unique identity identifier Face ID and the second cue word, the corresponding scene image package is generated through a large language model; Based on each storyboard image package, a corresponding video segment is generated, and the video segments are synthesized into the final video according to the storyboard order.
2. The method according to claim 1, characterized in that, The method includes: The first prompt word is only used to describe the character's role information, wherein the role information includes basic setting information and facial structure information. The facial features in the facial structure information are divided into multiple mutually exclusive sub-dimensions. Only one value is allowed for the same sub-dimension. Different sub-dimensions only take effect within their respective facial features and do not combine across facial features. The first prompt does not contain descriptions of scene information, style information, or plot information.
3. The method according to claim 1, characterized in that, Based on the reference face image, generating the unique identity identifier Face ID for the target person includes: Feature extraction is performed on the reference face image to obtain the facial feature vector of the target person; Based on the facial feature vector, a unique identity identifier, Face ID, is constructed and generated for the target person.
4. The method according to claim 3, characterized in that, Freezing the first prompt word, the reference face image, and the unique identifier Face ID of the target person includes: The first prompt word, the reference face image, and the unique identity identifier Face ID of the target person are frozen, and the first prompt word, the reference face image, and the unique identity identifier Face ID cannot be modified after freezing; That is, if it is necessary to modify the unique identity identifier Face ID of the target person, the first prompt word, the reference face image and the unique identity identifier Face ID must all be deleted and recreated, and there is no intermediate state.
5. The method according to claim 1, characterized in that, Based on the scene description information corresponding to each scene in the plot text, combined with the frozen unique identifier Face ID and the second cue word, the corresponding scene image package is generated through a large language model, including: Obtain the plot text for video generation, and perform storyboard decomposition on the plot text to obtain the storyboard description information corresponding to each storyboard in the plot text; Under the face consistency constraint of the frozen unique identifier Face ID, the corresponding segment image package is generated by a large language model based on the segment description information and the second prompt word for each segment.
6. The method according to claim 5, characterized in that, Under the face consistency constraint of the frozen unique identifier Face ID, based on the segment description information and the second cue word for each segment, the corresponding segment image package is generated through a large language model, including: The unique identity identifier Face ID of the character referenced in each storyboard is bound to the corresponding storyboard; Under the face consistency constraint of the unique identifier Face ID of the referenced person, a segmentation image package corresponding to each segment is generated through a large language model based on the segmentation description information and the second prompt word. The second prompt word is used to describe non-face elements in the image generation.
7. The method according to claim 5, characterized in that, After freezing the first prompt word, the reference face image, and the unique identifier Face ID of the target person, the method includes: Enter the target character's role asset expansion process, and generate a multi-view expansion map of the target character based on the target character's unique identity identifier Face ID. The multi-view expansion map includes a front view expansion map, an oblique side view expansion map, and a front side view expansion map.
8. The method according to claim 7, characterized in that, Under the face consistency constraint of the frozen unique identifier Face ID, the generation of the corresponding segment image package based on the segment description information and second cue word of each segment, through a large language model, also includes: If the storyboard references a multi-view extended map, then under the face consistency constraint of the frozen unique identifier FaceID corresponding to the multi-view extended map, based on the storyboard description information and the second prompt word, a storyboard image package corresponding to each storyboard is generated through a large language model. The storyboard images in the storyboard image package are saved in the order of the storyboard and the order of the numbers within the storyboard.
9. The method according to claim 1, characterized in that, After generating a baseline face image of the target person based on the first cue word using a large language model, the method includes: The quality of the generated reference face image is checked. If no face is detected in the reference face image, or multiple faces are detected, or key facial features are not recognizable, the generation of the reference face image is determined to have failed.
10. A video generation system based on face consistency constraints, characterized in that, The system is used to perform the method according to any one of claims 1 to 9, and the system includes a role identity definition module, a storyboard image generation module, and a video generation module; The role identity definition module is used to generate a baseline face image of the target person based on a first prompt word using a large language model, wherein the first prompt word is only used to describe the role information of the target person; The role identity definition module is used to generate a unique identity identifier Face ID for the target person based on the reference face image, and freeze the first prompt word, the reference face image and the unique identity identifier Face ID for the target person; The storyboard image generation module is used to generate a corresponding storyboard image package based on the storyboard description information corresponding to each storyboard in the plot text, combined with the frozen unique identity identifier Face ID and the second prompt word, through a large language model. The video generation module is used to generate corresponding video segments based on each storyboard image package, and to synthesize the video segments into the final video according to the storyboard order.