System for closed image and video generation with character consistency

TWI934877BActive Publication Date: 2026-08-01智見科技股份有限公司 PRINTAGE INC
View PDF 11 Cites 0 Cited by

Patent Information

Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
智見科技股份有限公司 PRINTAGE INC
Filing Date
2026-01-29
Publication Date
2026-08-01

AI Technical Summary

Technical Problem

Existing image generation models are highly random and feature-coupled, leading to issues such as identity drift, semantic ambiguity, and lack of privacy asset management, making them unsuitable for high-quality digital human production and professional narrative videos, and lacking effective data isolation and encryption protection.

Method used

A closed-loop image and video generation system utilizing semantic regularization, identity feature anchoring, and a privatization management architecture, which includes facial detection, geometric correlation verification, and a privacy control layer to ensure character consistency and secure data management.

Benefits of technology

The system stabilizes character characteristics, maintains consistency under stylistic influences, and ensures secure asset management by locking identity features and preventing data leakage, resulting in high-quality and stable character content output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TB001904349_001
    Figure TWG2TB001904349_001
  • Figure TWG2TB001904349_002
    Figure TWG2TB001904349_002
  • Figure TWG2TB001904349_003
    Figure TWG2TB001904349_003
Patent Text Reader

Abstract

This invention aims to provide a highly consistent character image and video generation system, addressing issues such as identity drift, semantic ambiguity, and lack of privacy asset management in generative AI. The system architecture includes a privacy control layer to define a closed execution environment; a baseline image selection module that automatically selects high-confidence baseline character images from an initial image set and stores them in an identity feature vault for identity anchoring; and a semantic conversion module that translates ambiguous original descriptions into standardized instructions with explicit feature constraints. The semantic conversion module only processes textual information that does not contain character identity features and does not access any identity data in the identity feature vault. By fusing baseline images, standardized instructions, and an optional set of reference images through the image generation control module, the system can produce static images that maintain consistent character features and can further generate time-stable dynamic videos through a cross-modal generation module. Finally, a closed album management interface and a real-time cache destruction mechanism achieve highly private closed-loop management of character assets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a highly consistent closed-loop character image and video generation system, and more particularly to a technical solution that utilizes semantic regularization, identity feature anchoring, and a privatization management architecture to solve problems such as identity drift, semantic ambiguity, and lack of privacy asset management in the image generation process. Prior Technology

[0002] In existing technologies, image generation models (such as diffusion models) are highly random and feature-coupled. When users attempt to convert static images into dynamic videos in different scenes, actions, or other situations, the models are heavily influenced by the input images, making it difficult to lock biometric features: the models cannot accurately lock the microscopic details of the subject (such as nose bridge height, eye distance, and pupil color), making them unsuitable for high-quality digital human production or professional narrative videos with continuity.

[0003] Generative models often exhibit high ambiguity in natural language prompts. Traditional systems directly input raw text into the model, frequently resulting in the model's inability to distinguish the importance of the generated content, which elements "must be changed (e.g., actions, lighting)" and which "must be preserved (e.g., face shape, background style)." This semantic confusion leads to unnecessary element changes in the generated results and an inability to accurately execute complex visual instructions (e.g., specific 3 / 4 perspectives or precise object interactions).

[0004] On the other hand, most image generation systems have an open architecture, where user input images and generated character feature data (embedding) are often processed in an open environment, lacking effective data isolation and encryption protection. For digital character assets with commercial value or private nature, existing technologies lack a closed management mechanism that can integrate "feature extraction, content generation, and private archiving," leading to the risk of user assets being leaked or unable to be systematically retrieved.

[0005] To facilitate examination, understanding, and retrieval, the applicant investigated publicly available information prior to the application date. Some relevant publicly available documents are listed below: (1) US20250308117A1 discloses a face replacement based on low-rank adaptation and subject-independent technology. Its main focus is on obtaining identity features of faces in real-world scenarios, without addressing character feature extraction under strong artistic styles; nor does it perform baseline image anchoring on the input static image to achieve a stable and high-quality character generation system. (2) US11983806B1 involves a system for generating images using a machine learning model, covering the core logic of how to process text input and convert it into image output; however, it does not disclose the characteristics of further standardizing and translating the input text and fusing identity-anchored images to achieve the generation of maintaining the consistency of character style and maintaining character identity features. (3) TWM652806U discloses a system for creating interactive virtual characters; however, its main focus is on using an artificial intelligence facial modeling module for character synthesis, without disclosing the creation of virtual characters with non-realistic styles, nor the security management architecture for the generated results.

[0006] Given that existing technologies for generating unrealistic characters generally suffer from technical flaws such as style drift in identity characteristics, influence from text descriptions and baseline image quality, and a lack of security filtering mechanisms in the generated results, this invention provides an image and video generation system with consistent character representation. Its key feature is the use of a standardized text translation module in conjunction with dynamic anchoring technology. This not only stabilizes character characteristics under strong stylistic influences but also ensures security through closed-loop management of character assets during the generation process, achieving high-quality and highly stable character content output. Summary of the Invention

[0007] This invention provides a closed-loop image and video generation system and its closed-loop management architecture for character consistency generation. It addresses the issues of unclear initial image features and overly vague original natural language descriptions, which lead to unpredictable randomness in the generated model's output, making it difficult to accurately align with user intent and maintain character consistency. Furthermore, in open generation environments, there is a risk of leakage or illegal caching of user biometric data (such as character facial images). The privacy control layer set up in this invention defines the execution environment, reduces the risk of privacy asset leakage, and also facilitates users' systematic retrieval of digital character assets.

[0008] To achieve the above objectives, this invention, in the online real-time phase, performs facial detection, basic quality filtering, and geometric correlation verification on the set of character images through a benchmark image selection module. It automatically determines and selects the image with the highest credibility as the benchmark character image. This benchmark character image possesses distinct facial features, making it suitable for subsequent image generation modules for identity anchoring. This benchmark character information is also stored in a closed vault for privacy management.

[0009] This invention establishes highly stable semantic translation rules, semantically decomposes the ambiguous text input by the user, and performs proactive feature completion. Explicit feature constraints are directly implanted for undefined visual parameters in the description, encapsulating the original narrative into standardized instructions and forcibly locking character features and specified local conditions for modification. This improves the quality and stability of character results when the model understands the fusion of text and image character features.

[0010] To prevent asset leakage risks, this invention establishes a privacy control layer to define a closed execution environment at the logical or entity level. This restricts core modules such as benchmark selection, image fusion, and timing generation to execution within this boundary, ensuring that the high-dimensional features of the subject do not interact with untrusted external environments. This invention establishes a closed-loop process from screening and generation to archiving, assisting users in systematically managing privacy-protected role assets and improving the efficiency and security of digital content management. Simple Explanation of the Diagram

[0011] [Figure 1] is a system architecture diagram of one embodiment of the present invention. [Figure 2] is a reference image selection logic flow diagram of one embodiment of the present invention. [Figure 3] is a flowchart of semantic transformation and instruction standardization processing according to one embodiment of the present invention (Prompt Transformation Flow). [Figure 4] is a multi-modal feature fusion diagram for image generation according to one embodiment of the present invention. [Figure 5] is a schematic diagram of cross-modal temporal consistency video generation according to one embodiment of the present invention. [Figure 6] is a diagram illustrating the relationship between privatized asset management and privacy protection in one embodiment of the present invention (Secure Data siloing Diagram). Implementation

[0012] The following illustrations illustrate specific embodiments of the present invention. The interactive target is not limited to a person, but may also be an animal, a two-dimensional character, a product, or other visible object; however, the following embodiments are merely illustrative, using a person as an example for explanation. The scope of protection of this invention is determined by the claims, and all equivalent changes or substitutions made in accordance with the spirit of this invention should be included within the scope of this invention.

[0013] Please refer to Figure 1, which is a system block diagram of one embodiment of the present invention. The present invention provides an image and video generation system for character consistency generation and its closed management architecture. First, the system includes an input display panel 100 for receiving user input information. The input display panel 100 includes: an initial character image set 101, which consists of a plurality of candidate subject images; a raw natural language description 102 for defining the intent of the generation task; and a reference image set 103 (optional) for providing specific action, pose, or scene layout references. In the data processing layer, the present invention sets a privacy control layer 430 (shown in the dashed box in the figure) to define a closed execution environment to ensure that the character's biometric data and generation core are not leaked. Within the privacy control layer 430, a baseline image selection module 210 is provided, which, after receiving the initial character image set 101, calculates a confidence index through a built-in object detection model and automatically selects the optimal "baseline character image". The features of the baseline character image are transmitted and stored in an identity feature vault 410, establishing an encrypted identity asset associated with the user account. At the same time, the input raw natural language description 102 is transmitted to the semantic conversion module 220 (located outside or at the boundary of the privacy control layer 430), which translates, expands and normalizes the ambiguous description into a "standardized instruction".

[0014] The core of this invention lies in an image generation control module 230, located within a privacy control layer 430. This module simultaneously receives a "baseline character image" from a baseline image selection module 210, "standardization instructions" from a semantic conversion module 220, and (if provided) a reference image set 103 from the input terminal. The image generation control module 230 locks in the subject's features through an identity anchoring unit and produces a character-consistent static image 320. This image visually inherits all the features of the baseline image and conforms to the instruction requirements. To achieve cross-modal dynamic output, the system further includes a cross-modal generation module 310, which receives the character-consistent static image 320 and uses it as a first frame or temporal reference, thereby driving a temporal generation model to produce a character-consistent dynamic video 330. The video generation process is also constrained by the privacy control layer 430, ensuring feature stability between dynamic frames. Ultimately, the generated character-consistent static images 320 and character-consistent dynamic videos 330 are integrated into a closed album management interface 420. This interface 420 uses an associative indexing mechanism to bind the generated content with the original baseline images and standardized commands, providing users with asset retrieval, management, and subsequent iterative generation.

[0015] Please also refer to Figure 2, which is a detailed flowchart of the benchmark image selection module 210 of the present invention. This module aims to automatically determine the optimal identity recognition benchmark from multiple initial images. The specific steps are as follows: Initial input receiving step S1: Receive the initial character image set 101 from the input terminal 100; Face detection and recognition step S2: Use the object detection model to perform face category recognition and localization on each image in the image set 101; Basic filtering and judgment step S3: Determine whether a single image contains only a single facial subject, and whether its detection confidence (clarity, whether it is occluded, deflection angle, etc.) is higher than a preset threshold; if not, the image is removed; Geometric correlation verification step S4: For the images filtered through step S3, verify whether the facial feature points conform to the preset anatomical geometric proportions; Benchmark selection step S5: Compare the confidence index of each image, and select the image with the highest score as the "benchmark character image"; Output benchmark character image step S6: Output the selected benchmark character image to the image generation control module 230, and simultaneously perform encrypted storage step S7 to store it in the identity feature safe 410. As shown in Figure 1, the safe is located within the protection range of the privacy control layer 430 to ensure that biometric data is not leaked.

[0016] Please refer to Figure 3, which is a flowchart of the semantic conversion module 220 of the present invention performing semantic conversion and instruction standardization processing. After receiving the original description, the module initiates processing in the original description receiving step S11; subsequently, in the semantic decomposition and tag mapping step S12, semantic decomposition is performed, mapping the user's ambiguous intent to specific technical tags (posture, clothing, background, etc.); in the visual detail completion procedure step S13, the system performs feature expansion, logically completing undefined spatial information (such as perspective, environmental details) in the description; the key constraint mechanism occurs in the explicit feature constraint implantation step S14. In this step, the system implements a technical control measure, namely, for visual attributes not mentioned or defined by the user in the original description (e.g., subtle facial geometry, precise pupil color, or skeletal proportions), actively extracting parameters from the reference image and implanting corresponding mandatory constraint instructions. The core function of this mechanism is to solve the identity drift problem caused by the random in-painting of undefined parameters in traditional generative models. By using this explicit feature constraint, it is ensured that when the generated model undergoes significant changes in clothing, scene, or action, these core identity feature parameters are strictly locked and unaffected by the randomness of the model. Finally, in the standardized instruction encapsulation step S15, the module integrates the above logic into a standardized instruction that conforms to a preset format (e.g., starting with "Change the Character's⋯⋯") and outputs it to the image generation control module 230, thereby achieving the technical effect of eliminating the randomness of model generation and maintaining the consistency of the character.

[0017] Please refer to Figure 4, which is a schematic diagram of the logical relationship of the image generation control module 230 of the present invention performing multi-dimensional input feature fusion. This module aims to receive input data from different modalities and integrate their features through a specific processing unit to produce high-quality images with character consistency. In this embodiment, the image generation control module 230 simultaneously receives multi-dimensional control inputs: a reference character image (the optimal identity reference selected by the reference image selection module 210, including core facial and biometric features), standardized instructions (translated by the semantic conversion module 220, including visual detail completion, action description, and feature constraint statements), and a reference image set 103: providing optional spatial structure references, such as body posture, environmental layout, or specific object features. Subsequently, a feature decoupling and fusion mechanism is implemented. The image generation control module 230 is equipped with a multimodal dedicated processing unit capable of processing text and images and converting them to the same high-dimensional space. The identity anchoring unit 231 (which receives the "baseline character image") and the multi-reference unit 232 (which receives the "reference image set 103") are used as a graphics encoder to convert the images into feature vectors. Together with the text feature vectors of the "standardization instructions", they use an adaptive attention mechanism to integrate the identity features of the baseline character image and the semantic features corresponding to the standardized generation instructions in a common representation space, ensuring that no character shift occurs during the generation process. The "standardization instructions" and the multi-reference unit 232 also use the adaptive attention mechanism to focus on the reference structure information (such as scene, pose, objects, depth map or skeleton map) that needs to be modified, guiding the model to generate the correct local modifications and outputting the final character-consistent static image 320. The relationship graph process in Figure 4 all occurs in the privacy control layer 430, reducing the risk of character asset leakage.

[0018] Please refer to Figure 5, which is a schematic diagram of cross-modal temporally consistent video generation by the cross-modal generation module 310. This module employs the following mechanism: the system receives a consistent static image 320 and extracts its visual features as a global bias injection. This global bias, along with "standardization instructions," applies temporally consistent constraints to the video sequence (Frame 1, Frame 2, Frame 3...), ensuring that each frame stably inherits the identity and style details of the first frame, ultimately producing 330 character-consistent dynamic videos. The process shown in Figure 5 all occurs within the privacy control layer 430, reducing the risk of character asset leakage.

[0019] Please refer to Figure 6, which is a diagram illustrating the relationship between asset management and privacy protection in this invention. This invention employs strict data isolation measures: Feature locking: Identity features are provided by the identity feature safe 410 and are only processed within the execution cache 230 or 310 within the privacy control layer 430; Instant destruction mechanism (X): Once the static image 320 or dynamic video 330 is generated, the system immediately triggers a destruction command (represented by X) to clear intermediate feature data in the cache, preventing feature residue. Associated indexing: The generated finished products (320, 330) can only be accessed within the closed album management interface 420, establishing an encrypted association with the original features in the safe, achieving secure and systematic asset management.

[0020] In summary, this invention provides a system architecture that integrates "identity asset management" and "high-fidelity role consistency generation." Through the closed environment defined by the privacy control layer 430, this invention effectively isolates sensitive biometric data and establishes a robust asset protection network through the identity feature safe 410 and an instant destruction mechanism. At the generation technology level, this invention utilizes the credibility index screening technology of the benchmark image selection module 210, combined with the instruction standardization and feature constraints of the semantic conversion module 220, to stably provide excellent AI generation input from the source.

[0021] 100: Input Display Panel 101: Initial Character Image Set 102: Original Natural Language Description 103: Reference Image Set (Optional) 200: Core Processing Layer 210: Reference Image Selection Module: Contains a target detection model and a confidence calculation unit. 220: Semantic conversion module: Contains a label mapping table and feature expansion / constraint units. 230: Image Generation Control Module 231: Identity Anchor 232: Multi-reference Unit 300: Generation and Output Layer 310: Cross-modal generation module: Receives still images and drives time-series models 320: Consistent static image 330: Character-Style Dynamic Video 400: Closed-loop management environment (closed system) 410: Identity Vault: Encrypted and linked to the user account. 420: Closed album management interface: includes a linked index database 430: Privacy Control Layer: Dashed boundary surrounding the safe and generation engine

Claims

1. A system for generating images and videos for character consistency, characterized by comprising: an identity vault for storing a plurality of character identity images extracted by a reference image selection module, wherein the system can only retrieve authorized identity data from the identity vault; a reference image selection module for receiving a plurality of initial character images and calculating the credibility index of each image through an object detection model, thereby selecting the image with the highest credibility as the reference image; a semantic conversion module for receiving a natural language description and executing an explicit feature constraint implantation procedure to actively complete and lock undefined character geometric feature parameters in the natural language description, thereby translating it into a standardized instruction format; and an image generation control module, comprising a privacy control layer, the image generation control module including at least one identity anchor unit for retrieving the identity data and locking character features based on the reference image, and generating a consistent static image by combining the standardized instruction with a feature weight allocation mechanism; wherein... The privacy control layer ensures that the character's identity features are not accessed by the public model, and that intermediate feature data generated during the generation process is destroyed immediately after the task is completed; and a cross-modal generation module is used to take the consistent static image as input to drive a temporal generation model to produce a dynamic video that maintains the consistency of the character's features; and a closed album management interface is used to display the generated content and to perform relational indexing on the generated images and videos with the original baseline image and corresponding semantic commands.

2. The system as described in claim 1, wherein the reference image selection module uses the initial character images as a plurality of candidate images, each candidate image containing an animated character face with realistic or non-realistic features; the reference image selection module calculates a confidence index for each candidate image through an object detection model, which is based on a comprehensive evaluation of the clarity, occlusion degree, deflection angle and preset geometric correlation of facial features in the image, and then selects the one with the highest confidence from the candidate images as the reference image.

3. The system as described in claim 1, wherein the image generation control module locks the core features of the character from the identity feature safe through the identity anchoring unit, and receives external reference images including objects, body postures or scene layouts through the multi-reference unit; characterized in that the image generation control module executes an asymmetric feature weight allocation mechanism to ensure that the weight of the core features of the character is higher than the weight of the external reference images, so as to forcibly maintain the visual consistency of the animated character's face and art style while introducing external reference features.

4. The system as described in claim 1, wherein the cross-modal generation module uses the consistent still image as a reference point for an initial latent space; characterized in that the cross-modal generation module executes a temporal continuity constraint algorithm to decouple the action description in the standardized instructions from the core character features from the identity feature safe, so as to force each frame of the video to maintain the same style, non-realistic brushstroke distribution, character features and their geometric proportions as the consistent still image when performing cross-temporal generation, thereby suppressing style drift caused by action changes.