Multi-image reference number life body generation method, device, equipment and storage medium
By using a multi-image reference digital life generation method, and optimizing the generation model through feature extraction and comprehensive constraints, the problems of insufficient personalization, difficulty in multi-source information fusion, and poor video consistency in existing technologies are solved, thus achieving high-quality and personalized digital life generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- LIANGSHENG DIGITAL CREATIVE DESIGN (HANGZHOU) CO LTD
- Filing Date
- 2026-02-28
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies for generating digital life forms with multiple reference images suffer from insufficient personalization, difficulty in fusing multi-source information, limited generation quality, and poor consistency of dynamic video frames, making it difficult to generate high-fidelity, personalized, and realistic digital life forms.
By acquiring multiple target reference images, feature vectors of appearance, clothing, scene, posture, and style are extracted respectively. Comprehensive constraints are determined and input into the target generation model for image generation. Combined with quality optimization and iterative updates, until a digital life form image that meets the requirements is generated.
It achieves high-fidelity and personalized digital life form generation, ensuring that the character's features, clothing style, background scene, and posture are highly consistent with the target reference image. The generated results are realistic and natural, with good consistency between video frames, and flexible and reliable background and clothing replacement.
Smart Images

Figure CN121747158B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence, computer vision, and digital image generation technology, and in particular to a method, apparatus, device, and storage medium for generating multi-image reference digital life forms. Background Technology
[0002] With the development of virtual reality (VR), augmented reality (AR), and metaverse technologies, digital lifeforms are increasingly being used in entertainment, education, customer service, and other fields. Digital lifeforms refer to virtual images generated by computers in a virtual environment, possessing specific human appearances and behaviors. However, existing technologies still have many limitations in generating multi-image reference digital lifeforms, mainly in the following aspects.
[0003] (1) Insufficient personalization and controllability. Existing image generation methods (such as typical text-to-image models) can only generate stylized images of people based on text descriptions, making it difficult to accurately reproduce the appearance and features of specific real people. Some AI face-swapping or avatar generation technologies can only change the face and cannot simultaneously change the clothing or background of the person as the user expects. There are also some few-sample personalized face generation methods that fix a specific face by providing a few photos to fine-tune the model (for example, some studies have proposed fine-tuning the internal representation of the diffusion generation model through 3 to 5 photos to enable the model to learn to generate images of that person), but these methods usually only focus on preserving the face and lack control over clothing style or background environment. Moreover, each time a new control is introduced (such as changing clothes), it is often necessary to redesign or retrain the model, which is a cumbersome process.
[0004] (2) Difficulty in multi-source information fusion. When generating digital life forms by combining multiple reference images, existing solutions lack effective feature fusion and control methods. If multiple reference images are simply input directly into the model (e.g., splicing a portrait and clothing image into a neural network), feature confusion and distortion are likely to occur. Images from different sources contain information of different dimensions—such as facial features, clothing styles, and shooting scenes. Direct mixing and processing can make it difficult for the model to distinguish the source of each piece of information, potentially resulting in errors such as the mixing of facial features and clothing textures, or incorrect background replacement. Currently, some multi-condition generation attempts can only handle two conditions and are difficult to stably superimpose multiple conditions: for example, existing pose-controlled portrait generation methods can generate a portrait with a specified pose based on a portrait and a pose skeleton image, but the clothing and background remain fixed; another example is virtual try-on technology, which allows the model to generate an image with different clothing based on a portrait and clothing image, but the background and the person's identity cannot be changed. When multiple elements such as the face, clothing, and background are required to be changed simultaneously, traditional models often fail to address one aspect, resulting in distortion or condition conflicts in one aspect. For example, changing the background results in noticeable inconsistencies at the edges of the character, or changing clothing alters the character's appearance. This reveals the technical challenges of feature decoupling and fusion under multi-source conditions.
[0005] (3) Limitations on Generation Quality and Realism. Generating high-fidelity, lifelike digital lifeforms that closely resemble real-life photographs remains a technical challenge in existing solutions. While early GAN models could generate clear facial images, they suffered from deficiencies in detail preservation and overall consistency, such as blurred clothing textures and potentially distorted limbs. In recent years, diffusion generation models have significantly improved image quality through progressive refinement, but directly applying them to personalized multi-image reference digital lifeform generation still presents challenges. On the one hand, the fidelity of the person's identity and the realism of the image need to be balanced: undesigned generation models easily sacrifice the person's recognizability for overall realism, resulting in generated faces that do not closely resemble the reference person. On the other hand, when attempting to change the background or place a virtual scene, improper handling can lead to mismatches in lighting and proportions between the digital lifeform and the background, reducing realism. For example, simply overlaying a person's image onto a new background often results in harsh edge transitions and inconsistent lighting, while if the generation model does not clearly distinguish between the person and background features, it may output an image with incorrect lighting. Existing technologies lack generation solutions that decouple characters from backgrounds, making it difficult to achieve a natural visual integration of characters and new scenes simultaneously.
[0006] (4) The consistency problem of dynamic video frames. In the generation of videos of digital life forms, the traditional method is usually to generate them independently frame by frame and then stitch the frame sequence into a video. Due to the lack of inter-frame constraints, the appearance of the character in consecutive frames is prone to drift, which manifests as flickering, jittering and other discontinuous artifacts. Especially when the character is speaking or moving, if the expressions and actions of each frame are not closely connected, the viewer will feel that it is unrealistic. Although some studies have tried to improve the continuity by adding temporal loss or post-processing optical flow correction during generation, the overall effect is still limited. At present, most personalized digital life form generation technologies are still at the stage of static images, and there are very few solutions that can generate videos of continuous digital life forms. If it is necessary to make the digital life form "move" (for example, to speak or perform as a virtual anchor), the existing technology either relies on a pre-made 3D animation skeleton driving model (which is costly and the image is not realistic enough) or uses a general video generation model but cannot guarantee that the generated character or clothing is specific.
[0007] In summary, existing technologies have not yet provided a complete solution for the image generation of digital life forms: either the degree of personalization is insufficient (unable to accurately specify identity, clothing, scene, etc.), or the generated results lack realism, or the video continuity is poor. These shortcomings limit the further application of digital life form technology in scenarios such as virtual anchors, intelligent customer service, film and television production, and virtual try-on. Therefore, it is necessary to provide a new technical solution to solve the above problems and realize the generation of realistic, personalized multi-image reference digital life forms based on multiple image inputs. Summary of the Invention
[0008] This invention provides a method, apparatus, device, and storage medium for generating multi-image reference digital life forms, in order to solve the problem that there are still many limitations in the existing technology for generating multi-image reference digital life forms, so as to achieve the generation of images corresponding to personalized digital life forms of real people.
[0009] A method for generating digital life forms using multi-graph references, comprising:
[0010] Obtain multiple target reference images corresponding to different attributes, wherein the multiple target reference images include at least two of the following: a person's appearance reference image, a clothing reference image, a scene reference image, a pose reference image, and a style reference image;
[0011] Feature extraction is performed on the multiple target reference images to determine multiple target feature vectors. The multiple target feature vectors include at least two of the following: facial feature vector, clothing feature vector, scene feature vector, posture feature vector, and style feature vector.
[0012] Determine comprehensive constraints, which include at least a comprehensive condition vector and condition injection parameters. The comprehensive condition vector includes attribute sub-vectors corresponding to at least two attributes among appearance, clothing, scene, posture, and style. The condition injection parameters include weights of multiple attributes at different resolution layers of a pre-trained target generation model. The weights are used to indicate the degree of influence of the attributes on different resolution layers.
[0013] The comprehensive constraints are input into the target generation model to generate images, and the initial generated image corresponding to the digital life form is determined. The generation features of different attributes in the initial generated image are constrained by the comprehensive constraints.
[0014] The initial generated image corresponding to the digital life form is optimized for quality, and the quality detection results corresponding to multiple attributes are determined. When the quality detection result corresponding to any attribute is a failure, the comprehensive constraint conditions are iteratively updated based on the quality detection result and the image is regenerated until the quality detection results corresponding to all attributes are all successful, and then the target generated image corresponding to the digital life form is determined.
[0015] A multi-image reference digital life form generation device, comprising:
[0016] The reference image acquisition module is used to acquire multiple target reference images corresponding to different attributes. The multiple target reference images include at least two of the following: a person's appearance reference image, a clothing reference image, a scene reference image, a pose reference image, and a style reference image.
[0017] The feature extraction module is used to extract features from the multiple target reference images respectively and determine multiple target feature vectors. The multiple target feature vectors include at least two of the following: facial feature vector, clothing feature vector, scene feature vector, posture feature vector, and style feature vector.
[0018] A constraint determination module is used to determine comprehensive constraints. The comprehensive constraints include at least a comprehensive condition vector and condition injection parameters. The comprehensive condition vector includes attribute sub-vectors corresponding to at least two attributes among appearance, clothing, scene, posture, and style. The condition injection parameters include weights of multiple attributes at different resolution layers of a pre-trained target generation model. The weights are used to indicate the degree of influence of the attributes on different resolution layers.
[0019] The image generation module is used to input the comprehensive constraints into the target generation model to generate images and determine the initial generated image corresponding to the digital life form. The generation features of different attributes in the initial generated image are constrained by the comprehensive constraints.
[0020] The quality optimization module is used to optimize the quality of the initial generated image corresponding to the digital life form and determine the quality detection results corresponding to multiple attributes. When the quality detection result corresponding to any attribute is a failure, the comprehensive constraint conditions are iteratively updated based on the quality detection result and the image is regenerated until the quality detection results corresponding to all attributes are all passed, and then the target generated image corresponding to the digital life form is determined.
[0021] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the above-described method for generating multi-graph reference digital life forms.
[0022] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for generating multi-graph reference digital life forms.
[0023] The multi-image reference digital life form generation method, apparatus, device, and storage medium provided in this invention can utilize multiple target reference images for feature extraction and fusion to obtain a high-fidelity digital life form corresponding to the user's personalized requirements. This ensures that the facial features, clothing style, background scene, posture, and other elements of the generated target image are highly consistent with the target reference image, thus balancing the fidelity of the digital life form's identity and scene adaptation. Users do not need cumbersome model training or manual editing; they only need to provide a few photos to quickly obtain a lifelike personal digital life form image. This can be used in virtual interaction, digital human modeling, immersive applications, and other fields, and has broad application prospects. Attached Figure Description
[0024] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart of a method for generating multi-image reference digital life forms in an embodiment of the present invention;
[0026] Figure 2 yes Figure 1 A flowchart of step S102;
[0027] Figure 3 yes Figure 1 Another flowchart of step S102;
[0028] Figure 4 yes Figure 1 A flowchart preceding step S103;
[0029] Figure 5 yes Figure 1 A flowchart of step S104;
[0030] Figure 6 yes Figure 5 A flowchart of step S501;
[0031] Figure 7 yes Figure 1 A flowchart following step S104;
[0032] Figure 8 yes Figure 1 A flowchart of step S105;
[0033] Figure 9 This is a schematic diagram of a multi-image reference digital life form generation device in an embodiment of the present invention. Detailed Implementation
[0034] To make the technical problems solved, the technical solutions, and the beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0035] This invention provides a method for generating multi-image reference digital life forms, which can be applied to mobile phones, computers, or other electronic devices, such as... Figure 1 As shown, the methods for generating multi-image reference digital life forms include:
[0036] S101: Obtain multiple target reference images corresponding to different attributes. The multiple target reference images include at least two of the following: a person's appearance reference image, a clothing reference image, a scene reference image, a pose reference image, and a style reference image.
[0037] S102: Perform feature extraction on multiple target reference images respectively to determine multiple target feature vectors. The multiple target feature vectors include at least two of the following: facial feature vector, clothing feature vector, scene feature vector, posture feature vector, and style feature vector.
[0038] S103: Determine the comprehensive constraints. The comprehensive constraints include at least a comprehensive condition vector and condition injection parameters. The comprehensive condition vector includes attribute sub-vectors corresponding to at least two attributes among appearance, clothing, scene, posture and style. The condition injection parameters include the weights of multiple attributes in different resolution layers of the pre-trained target generation model. The weights are used to indicate the degree of influence of the attributes on different resolution layers.
[0039] S104: Input the comprehensive constraints into the target generation model to generate images, determine the initial generated image corresponding to the digital life form, and the generation features of different attributes in the initial generated image are constrained by the comprehensive constraints.
[0040] S105: Optimize the quality of the initial generated image corresponding to the digital life form and determine the quality detection results corresponding to multiple attributes; when the quality detection result corresponding to any attribute is a failure, iteratively update the comprehensive constraint conditions based on the quality detection results and regenerate the image until the quality detection results corresponding to all attributes are all passed, and then determine the target generated image corresponding to the digital life form.
[0041] In this context, the original reference image refers to the reference image generated by the user-input auxiliary multi-image reference digital lifeform. The target reference image is a preprocessed version of the original reference image. In this example, the multiple target reference images include at least two of the following: a person's appearance reference image, a clothing reference image, a scene reference image, a pose reference image, and a style reference image. The person's appearance reference image specifies the person's appearance. The clothing reference image specifies the clothing worn. The scene reference image specifies the environmental background. The pose reference image specifies the person's pose. The style reference image specifies the style tone. Here, the reference images are used to assist in the generation of the multi-image reference digital lifeform.
[0042] The target generation model is a pre-trained model used to generate images corresponding to digital life forms. The initial generated image is an image generated by the target generation model, and the target generated image is an image after quality optimization of the initial generated image. Here, the image can be either an image or a video.
[0043] As an example, in step S101, the electronic device may receive multiple target reference images input by the user. These multiple target reference images include at least two of the following: a person's appearance reference image for specifying a person's face, a clothing reference image for specifying clothing, a scene reference image for specifying the environmental background, a pose reference image for indicating a person's posture, and a style reference image for indicating a style tone. After acquiring the multiple target reference images, the electronic device performs preprocessing on each target reference image. The preprocessing includes at least one of the following: size or resolution normalization, color space and channel normalization, noise reduction and sharpening, portrait detection and alignment cropping, foreground / background segmentation, clothing region segmentation, and keypoint or skeletal point detection.
[0044] As an example, in step S102, after acquiring multiple target reference images, the electronic device can preprocess each target reference image to obtain preprocessed target reference images. Then, different feature extraction algorithms or models are used to extract features from the preprocessed target reference images with different attributes, determining the target feature vector corresponding to each target reference image. The multiple target feature vectors include the facial feature vector corresponding to the person's facial reference image. Clothing feature vectors corresponding to clothing reference images Scene feature vector corresponding to scene reference image Attitude feature vectors corresponding to the attitude reference image Style feature vectors corresponding to style reference images At least two of the following. In this example, feature extraction is performed on target reference images with different attributes to achieve independent representation and decoupling of different attribute information, so that the feature vectors of different targets do not confuse or interfere with each other. In this example, preprocessing includes at least one of the following: size or resolution normalization, color space and channel normalization, denoising and sharpening, portrait detection and alignment cropping, foreground and background segmentation, clothing region segmentation, and key point or skeletal point detection.
[0045] As an example, in step S103, after acquiring multiple target feature vectors, the electronic device can use a pre-set feature fusion mechanism to fuse the multiple target feature vectors and determine a comprehensive condition vector. This comprehensive condition vector is the output of the fusion processing of multiple target feature vectors based on the pre-set feature fusion mechanism. It can serve as a comprehensive latent vector or multi-channel condition signal for generation conditions. It condenses and processes the key information of multiple target reference images provided by the user, making the multiple target feature vectors compatible and synergistic, unlike the rigid stitching of multiple inputs, which helps ensure the high quality of the subsequently generated image. Next, the electronic device also acquires the condition injection parameters corresponding to the pre-trained target generation model. The target generation model includes multiple resolution layers. The condition injection parameters can be the system default condition injection parameters or the condition injection parameters input by the user. The condition injection parameters include the weights of multiple attributes in different resolution layers of the pre-trained target generation model. The weights are used to indicate the strength of the influence of the attributes on different resolution layers. For example, if the target generation model includes three resolution layers, the conditional injection parameters obtained by the electronic device can include the weights of the attributes of appearance, clothing, scene, pose, and style in each resolution layer, ensuring that the sum of the weights of all attributes within the same resolution layer is 1. For instance, if the multiple target reference images input by the user only include target reference images corresponding to the three attributes of appearance, clothing, and scene, the sum of the weights of these three attributes in each resolution layer must be 1; if the target reference images input by the user include target reference images corresponding to the five attributes of appearance, clothing, scene, pose, and style, the sum of the weights of these five attributes in each resolution layer must be 1. This allows the electronic device to autonomously adjust the influence of different attributes on the generated image of the digital life form according to actual needs. Finally, the electronic device can determine comprehensive constraints based on the comprehensive condition vector and conditional injection parameters. This comprehensive constraint condition is used to instruct the target generation model to generate images.
[0046] As an example, in step S104, the electronic device processes the comprehensive condition vector using a pre-trained target generation model. Specifically, the comprehensive condition vector is used as a condition input to the target generation model, so that the target generation model generates an initial generated image corresponding to the digital life form under the guidance of the comprehensive condition vector. In this example, the initial generated image can be an initial generated image or an initial generated video formed by multiple frames of initial generated images. The initial generated image includes multiple generated features of different attributes. The generated features of each attribute are constrained by comprehensive constraints to make the generated features of that attribute highly similar to the target reference image corresponding to that attribute. In this example, when multiple target reference images include a person's appearance reference image, a clothing reference image, and a scene reference image, under the guidance of comprehensive constraints, the appearance, clothing, and scene attributes in each initial generated image generated by the target generation model are highly similar to the person's appearance reference image, clothing reference image, and scene reference image, respectively.
[0047] As an example, in step S105, after the electronic device acquires the initial generated image corresponding to the digital life form, it needs to optimize the quality of the initial generated image. Specifically, different quality detection algorithms corresponding to different attributes can be used to perform quality detection on each initial generated image and determine the quality detection results corresponding to different attributes. When any quality detection result is a failure, the comprehensive constraint conditions need to be adjusted. Specifically, the condition injection parameters in the comprehensive constraint conditions can be adjusted, and then step S104 is repeated to generate images until the quality detection results corresponding to all attributes are all passed. Then, the initial generated image corresponding to it is determined to be the target generated image that meets the quality requirements.
[0048] In this example, after acquiring the target-generated image corresponding to the digital life form, the electronic device can save or present the result in a corresponding format according to user needs. For example, the static digital life form image can be saved as a JPEG / PNG image, or the processed frame sequence can be encapsulated and output using a standard video encoding (such as MP4). If used in a real-time interactive scenario, the generated image stream can be pushed to the front-end application, allowing users to see their digital life form appear in the target scene in real time.
[0049] Thus, the entire process from multi-source input to the presentation of a digital life form is fully realized. The above-described method steps of this invention can be executed by software on general computing devices (such as servers, PCs, or mobile terminals), or the various functional modules can be implemented by hardware circuits.
[0050] This invention provides a method for generating digital life forms with multiple reference images, which achieves multiple objectives that cannot be simultaneously achieved by existing technologies, and has the following significant advantages.
[0051] (1) Highly personalized customization capability: This solution can generate digital life forms that perfectly match the user's requirements in all aspects based on multiple target reference images provided by the user, realizing the customized generation of personalized digital life forms. Compared with methods that require repeated debugging of text prompts to approximate the target image, this invention directly uses images as the generation condition, greatly simplifying the personalized customization process, and providing a WYSIWYG experience. For example, traditional pure text-guided image generation models often require users to cleverly conceive prompts and conduct multiple trials to generate the desired clothing or scene, while this invention only requires providing the corresponding clothing reference image and scene reference image, and the generated result is highly consistent with the reference in one go. Another example is that some model-based fine-tuning methods require time-consuming and computationally expensive training of the model whenever a new target character is added, while this invention can instantly generate the corresponding character image using pre-extracted facial feature vectors, without the need for additional model training, resulting in fast response speed and low usage threshold. It can be seen that when creating a user's personal digital life form, this invention greatly improves efficiency and controllability, enabling ordinary users to easily obtain highly personalized digital life form images and fully satisfying users' personalized expression needs.
[0052] (2) The multi-condition fusion effect is realistic and natural. Specifically, through strict feature decoupling and attention fusion mechanisms, this invention overcomes the problem of information mixing under multi-source conditions, and the quality of the generated result is significantly better than that of simple superposition conditions. Some existing attempts to fuse multiple images often suffer from distortion problems, such as the deviation of facial features due to the influence of background color, and the loss of clothing details during synthesis. However, this invention strictly distinguishes and processes various features independently in the algorithm architecture, and then intelligently fuses them within the model to ensure that the effects of each condition are fully reflected. Therefore, in the target generated image corresponding to the digital life form, the facial details, clothing texture, and scene atmosphere are consistent with their respective references and are coordinated with each other: there is no phenomenon of misalignment between human features and background, nor is there blurring of human edges or double image superposition as in some GAN fusions. Especially in multi-person or complex scenes, this solution can maintain the clear independence of each human figure and will not cause confusion of identity features as in some multi-person generation methods. Overall, this invention significantly improves the credibility of multi-condition image synthesis, and the output image is realistic and natural, meeting the requirements of high realism for commercial applications.
[0053] (3) The image realism and detail far surpass previous methods. Thanks to the powerful modeling capabilities of the diffusion generation model and the combination of local detail guidance technology, the digital life forms generated by this invention have achieved an unprecedented level of overall clarity, texture, and detail restoration. In contrast, traditional GAN-generated human images often suffer from limited resolution and blurred details; however, this invention, through multi-step iterative denoising, meticulously depicts facial expressions, hair strands, and clothing textures, achieving an effect close to that of photographs. For example, in experiments, the target generated image corresponding to the digital life form generated by this invention is almost indistinguishable from real photos in terms of detail sharpness and color fidelity, while the same human image generated by a mainstream GAN model has unnatural skin texture and obvious missing clothing patterns. Furthermore, for clothing materials, this invention integrates the provided real clothing texture, making even subtle elements such as lace clearly visible, which is difficult to achieve with previous automatic clothing-changing technologies. Thus, this invention sets a new benchmark for the image quality of virtual digital life forms, making it suitable for high-requirement application scenarios such as high-definition promotional videos and posters.
[0054] (4) The video output is consistent and natural between frames. Specifically for the generation of digital life form video sequences, this invention ensures high frame consistency through special temporal modeling. The "shaking" of the characters caused by traditional frame-by-frame generation is almost imperceptible in the videos generated by this invention. For example, in the virtual anchor test, the facial posture and expression changes between any two frames of the speaker video output by this invention are smooth and continuous, without sudden jumps; while the method without temporal constraints shows discontinuity multiple times at blinking and mouth shape changes. Objective evaluation shows that the videos generated by this invention improve the inter-frame SSIM (structural similarity) index by an average of more than 30%, and the trajectory of human key points is more stable. This means that when viewers watch the digital life form videos generated by this invention, the characters' movements are natural and smooth, almost indistinguishable from real shooting, enabling digital life forms to perform in real time and for long periods of time. This advantage significantly expands the practical value of digital life forms.
[0055] (5) Flexible and reliable background and clothing replacement: This invention allows for separate control and free replacement of the background scene and clothing appearance of digital life forms, while ensuring "changing the background without changing the face" and "changing clothes without changing the person," avoiding common problems in existing technologies. Users can use any background image to change the scene, and the electronic device will automatically adjust the integration of the person with the new environment (such as lighting, proportions, etc.) to make the result harmonious and natural; different clothing photos can also be provided to change the clothing of the digital life form, while the person's body shape and facial features remain unchanged. Traditional methods either simply cut out the image and replace the background, resulting in harsh edges of the person, or change the background through style transfer, which changes the overall color tone of the person. For example, in the field of clothing replacement, existing solutions often require pre-shooting portraits of people in specific poses before applying clothes, while this invention can directly generate a portrait of a person wearing the corresponding clothing based on any clothing image, without being limited by the pose of the original photo. Therefore, this invention achieves truly flexible and versatile background and clothing control, significantly improving the scene adaptability of digital life forms. Users can customize the environment and appearance of digital life forms at will, and the composite effect is professional and natural.
[0056] (6) The architecture is general-purpose and easily expandable. This invention adopts a modular design and is compatible with various deep learning technologies, possessing strong adaptability and upgrade potential. On the one hand, it is compatible with various existing pre-trained models (such as face recognition models, pose estimation algorithms, different versions of diffusion generation models, etc.), and can continuously improve performance by replacing better sub-modules as technology advances. On the other hand, the architecture can easily and smoothly expand to new control dimensions. For example, a voice-driven module can be added to enable digital life forms to synchronize lip movements based on voice, or text label control can be added to further refine the style. That is, only the corresponding feature extraction and fusion modules need to be added to the existing framework, without needing to start from scratch. Therefore, compared to closed, dedicated algorithms, this invention is more flexible, less prone to obsolescence, and can continuously enrich its functionality. This means that in commercial applications, this solution has a longer lifespan and a higher return on investment, not only currently holding a leading position but also maintaining competitiveness in the future.
[0057] In summary, this invention proposes solutions to key challenges such as multi-image fusion, image realism, and video continuity, achieving results that comprehensively surpass existing digital life form generation technologies. This solution enables the generation of high-quality, personalized multi-image reference digital life forms, significantly reducing production barriers and costs. Various practical application tests demonstrate the superior performance of this solution, highlighting its technological advantages and practical value, and possessing significant industrial value. As the underlying technology of a multi-image reference digital life form generation platform, this invention has important strategic significance. These advantages will be fully reflected and protected in the appended claims.
[0058] In one embodiment, such as Figure 2As shown, step S102, which involves extracting features from multiple target reference images to determine multiple target feature vectors, includes at least two of the following steps:
[0059] S201: Perform face region cropping and alignment processing on the reference image of the human face to determine the face region image, extract facial features from the face region image, and determine the facial feature vector;
[0060] S202: Remove the background from the clothing reference image, determine the clothing region image, extract clothing features from the clothing region image, and determine the clothing feature vector;
[0061] S203: Scale the scene reference image to obtain a scene reference image of a preset size, extract scene features from the scene reference image of the preset size, and determine the scene feature vector;
[0062] S204: Perform human figure segmentation on the pose reference image, determine the human foreground image, extract pose features from the human foreground image, and determine the pose feature vector.
[0063] S205: Perform content contour blurring and style texture enhancement processing on the style reference image to obtain a style texture image, extract style features from the style texture image, and determine the style feature vector.
[0064] As an example, in step S201, the electronic device can perform face region cropping and alignment processing on the person's facial reference image to determine the face region image. Specifically, the person's facial reference image can be cropped and aligned to a face region image of approximately 128×128 pixels to remove interference information and ensure the accuracy and efficiency of subsequent processing. Using a pre-trained face recognition or portrait coding model, facial features are extracted from the face region image to extract facial feature vectors. This facial feature vector This refers to the unique facial features (such as facial proportions, face shape, and skin tone) of a person in a reference image representing their appearance. Preferably, a fixed-dimensional facial feature vector can be obtained using a mature face recognition algorithm. Facial feature vector The dimensions can be 128 or 512. In this example, the facial feature vector determined based on the human face reference image can ensure that the appearance of the subsequently generated digital life form is highly similar to the human face reference image provided by the user.
[0065] In one embodiment, in step S202, the electronic device can first use a pre-set clothing segmentation model to separate the clothing region from the clothing reference image, obtaining a clothing region image after removing the background; the clothing region image may include the segmentation mask corresponding to the clothing region and clothing texture information, which are used for subsequent clothing feature extraction and deformation adaptation. Then, the clothing region image is input into a pre-trained deep neural network for feature extraction to obtain clothing feature vectors. Clothing feature vector This includes information such as the garment's cut, pattern, and color. If necessary, it can be combined with human anatomy algorithms, and the clothing region image can be deformed and adapted based on the target person's body shape or posture. The deformation adaptation algorithm can employ Thin Plate Spline (TPS) transformation or a human skeleton-driven approach to match the garment's cut or other features in the clothing region image with the target person's body shape or posture. Next, the deformed and adapted clothing region image is input into a pre-trained deep neural network for feature extraction, ensuring that the extracted clothing feature vectors are adapted to the target person's body shape or posture, thus obtaining a texture mapping of the clothing worn by the target person. In this example, the clothing feature vectors determined based on the clothing reference image ensure that the subsequently generated digital lifeform not only wears the specified clothing but also has detailed textures highly similar to the user-provided clothing reference image.
[0066] In one embodiment, in step S203, the electronic device can scale the scene reference image to obtain a scene reference image of a preset size (e.g., 512×512 pixels) to match the output resolution of the final generated digital life form; then, scene features are extracted from the scene reference image of the preset size to determine the scene feature vector. In this example, scene recognition or image encoding models are used to extract the feature representation of the background environment from the scene reference image of the preset size and determine it as the scene feature vector. There are two ways to achieve this: (1) Direct background usage: The scene reference image provided by the user is directly used as the background of the final generated background. At this time, the size of the scene reference image is mainly adjusted, and the background and the character are separated by foreground / background layer rendering during subsequent synthesis; (2) Background regeneration: The semantic and stylistic features of the scene reference image are extracted to guide the regeneration of the background. Specifically, a multimodal model or variational autoencoder (VAE) can be used to encode the scene reference image into a scene feature vector. Alternatively, semantic tags of the scene (such as "beach," "office," etc.) can be extracted as conditional prompts. In some embodiments, depth estimation models can be further used to generate depth maps or segmentation maps for the scene, providing additional information on scene layout and object positions. Regardless of the approach used, the scene feature vector determined based on the scene reference image is a description of the characteristics of the background environment, allowing digital life forms to be placed in scenes with corresponding atmospheres.
[0067] As an example, in step S204, the electronic device can use a pre-set portrait segmentation model to segment the pose reference image, obtaining an initial foreground image by removing background information and ensuring foreground information, so as to avoid background contour interference with key point detection. Then, the initial foreground image is processed to grayscale to eliminate interference from color and decorative patterns, obtaining a human foreground image that retains only the brightness contour information of the human body. Next, the electronic device can extract pose features from the human foreground image to determine the pose feature vector. Specifically, this can be achieved by extracting the skeletal key point information of the pose in the human foreground image using a pose estimation algorithm, and determining it as the pose feature vector. For example, common pose estimation algorithms can be used to extract key points from a human foreground image to determine the human body's key points. The response values of all human key points are calculated, and key points with response values less than a preset threshold are deleted as invalid key points. Key points with response values not less than the preset threshold are retained and identified as valid key points. Based on the valid key points, geometric features independent of size or position are extracted to avoid distortion of the extracted features due to image scaling or translation. Finally, all extracted geometric features are concatenated into a pose feature vector, which is used to guide the digital life form to assume a specified action pose. Alternatively, a pre-trained human pose estimation model can be used to extract key points from a human foreground image, determine skeletal key point features, and identify them as the pose feature vector. In this example, if no pose reference image is provided, this step can be skipped, and the electronic device will generate a target person in a standing posture or other preset posture by default.
[0068] Furthermore, after acquiring the posture feature vector, the electronic device can also adjust the posture to improve controllability and stability, specifically including: (1) Key point normalization. Using the root key point of the human body (e.g., the center of the pelvis or the center of the torso) as the origin, the coordinates of each key point are normalized by translation; and the length of the main skeleton of the human body (e.g., shoulder width or hip-shoulder distance) is used as the scale for scaling normalization, so that the posture representation is independent of the image resolution and the distance of the person from the camera. (2) Key point repair and smoothing. For key points with low response values or missing key points, interpolation repair is performed based on the topological relationship of adjacent skeletons; when the output is a video or posture sequence, the key points of adjacent frames are smoothed by exponential moving average or Kalman filtering to suppress key point jitter, thereby improving the motion stability of the generated video. (3) Posture intensity adjustment. Set the posture intensity coefficient. Based on the user's control requirements for "strong / weak attitude constraints", the attitude feature vector is mixed and adjusted, for example: , The pose feature vector is extracted from the pose reference image. This is the default pose feature vector. The adjusted attitude feature vector. (4) Construction and injection of attitude constraint signals. The adjusted attitude feature vector is... The pose constraint signal is converted into a pose constraint signal (e.g., key point heatmap, skeleton connection map or its encoded vector), and the pose constraint signal is written into the integrated condition vector Fc (as part of the pose sub-vector) and / or injected into the medium resolution layer of the diffusion generation model (lightly injected into the high resolution layer if necessary) to stably constrain the limb structure and motion layout during the generation process and avoid limb distortion and pose drift.
[0069] As an example, in step S205, the electronic device can perform preprocessing operations on the style reference image, such as format and channel normalization, size normalization, and edge-preserving denoising. Next, the preprocessed style reference image undergoes content contour blurring to slightly blur the content contours and prevent complex content (such as people, landscapes, or buildings) from being mixed in, thus obtaining an initial texture image. Then, a style texture enhancement processing can be performed on the initial texture image using, but not limited to, a Laplacian edge enhancement algorithm, to determine the style texture image. Finally, a pre-trained neural network model (including but not limited to CNN) is used to extract style features from the style texture image, obtaining style information such as brushstrokes, composition, and tone blending, to determine the style feature vector. In this example, the style feature vector determined based on the style reference image ensures the overall visual style of the subsequently generated digital life form's images or videos.
[0070] Furthermore, after acquiring the style feature vector, the electronic device can also adjust the style to achieve controllable style intensity and more stable style, specifically including: (1) Style intensity adjustment. Setting the style intensity coefficient. Based on the user's desired level of stylization, the style feature vectors are mixed and adjusted, for example: ,in, The style feature vector is extracted from the style reference image. It is a neutral style vector (e.g., statistical mean style or preset realistic style). The adjusted style feature vector. (2) Hue and brightness normalization. To avoid color cast of the style reference image causing overall color cast, color normalization processing can be introduced into the style reference image or style feature vector during the style constraint construction stage. For example, the brightness L channel is normalized in Lab space, and histogram matching or color transfer is performed on the a / b color channels to improve the stability of style transfer. (3) Style constraint signal construction and injection. The adjusted style feature vector Write the integrated conditional vector Fc (as part of the style sub-vector) and / or inject it into the low-resolution layer of the diffusion generation model (lightly injecting medium-resolution layers if necessary) to prioritize the constraint of global style information such as overall composition, tone atmosphere and brushstroke texture, and avoid style details from interfering with the fine structure of the character's appearance.
[0071] In this embodiment, different target reference images are converted into high-level numerical feature representations, including facial feature vectors. Clothing feature vectors Scene feature vector Attitude feature vector and style feature vector At least two of these elements enable independent representation and decoupling of different attribute information. Each feature is extracted by an independent model, emphasizing its corresponding attribute, without confusion or interference between them. For example, facial feature vectors... Extracting the facial features from the reference image, excluding the influence of clothing and background; clothing feature vectors. The focus is on extracting clothing and other apparel, excluding information about people. These resulting features each serve a specific purpose, laying the foundation for subsequent multimodal fusion.
[0072] In one embodiment, such as Figure 3 As shown, in step S102, feature extraction is performed on multiple target reference images to determine multiple target feature vectors, including:
[0073] S301: Perform person detection and instance segmentation on the person facial reference image to determine the facial feature vectors, segmentation masks and bounding box information corresponding to multiple target persons in the target reference image;
[0074] S302: Determine the location code for each target person based on the segmentation mask and bounding box information corresponding to each target person;
[0075] S303: Associate the facial feature vectors and location codes of multiple target persons to determine the person feature vectors of multiple target persons;
[0076] S304: Based on the character feature vectors corresponding to multiple target characters, and at least one of the clothing feature vector, scene feature vector, posture feature vector and style feature vector, determine the comprehensive condition vector corresponding to multiple target characters.
[0077] S305: Inject the comprehensive condition vectors corresponding to multiple target characters into different generation regions or different generation branches of the target generation model, and constrain the generation region of each target character based on the character segmentation mask and / or pose skeleton information.
[0078] In one embodiment, when multiple target reference images include multiple target figures or when multiple target figures need to be composited in the same frame, it is necessary to perform positional encoding or label differentiation on multiple target figures in the facial reference images during the feature extraction process. Specifically, positional encoding and other techniques can be combined to ensure that the reference information of multiple target figures is not confused. When there are multiple figures or multiple similar references (for example, a user provides two photos of figures and wants to generate a group photo), different positional information can be added to the features of different reference sources to label the features of each figure separately: different positional identifiers are added when injecting into the generation model, thereby ensuring that the features of each figure are not confused. In addition, for the spatial layout of figures in the background, this invention can also introduce explicit spatial positional information into the fused features or provide constraints with the help of a dedicated control network. For example, the extracted pose skeleton map can be input into a dedicated pose control sub-network so that the target generation model strictly arranges the position of the figure's limbs according to the skeleton; or a segmentation mask of the figure can be provided to guide the model to generate the foreground figure in a specified area of the image and the background in another area. This layout control method solves the problem of uncontrollable figure position in the past, allowing users to not only specify "what to generate" but also partially specify "where to generate".
[0079] As an example, in step S301, the electronic device needs to perform a person detection and instance segmentation step, performing person detection and instance segmentation on a person facial reference image containing multiple target persons to obtain the segmentation mask and bounding box information corresponding to each target person. For example, the instance segmentation network can adopt a Mask R-CNN network or other equivalent instance segmentation networks, let the first... The segmentation mask for each target person is Bounding box information is .
[0080] As an example, in step S302, the electronic device needs to perform a location code generation step, which can be based on the step... Bounding box of a target person (e.g., center point coordinates, width and height, normalized coordinates relative to the image scale) and / or segmentation mask c generate positional codes. , Location coding is used to describe the spatial position and occupied area of each target person in the reference image of the person's appearance.
[0081] As an example, in step S303, the electronic device needs to perform tag differentiation and specific feature extraction operations. To avoid confusion between the appearances / clothing of multiple people, person tags are set for different target people. (or character index), and based on segmentation mask For each target person, facial feature extraction is performed to determine the facial feature vectors corresponding to multiple target people. Then, the features of each target person are combined with other target reference images to extract features such as clothing and posture, resulting in a set of target feature vectors (such as facial feature vectors, clothing feature vectors, and posture feature vectors) exclusive to each target person.
[0082] As an example, in step S304, the electronic device needs to associate the person tag corresponding to each target person with its comprehensive condition vector. A comprehensive condition vector is constructed for each target person separately. and encode each position. , Corresponding exclusive comprehensive condition vector (For example, embedding the positional encoding vector with...) (Adding or splicing) thus achieves "independent positioning and attribute restoration of multiple characters in the same scene" in the subsequent diffusion generation stage, and significantly reduces the risk of identity feature confusion when generating multiple characters.
[0083] As an example, in step S305, during the diffusion generation stage, in order to avoid crosstalk between the identities and poses of multiple people, the electronic device can inject the comprehensive condition vectors corresponding to multiple target people into different generation regions or different generation branches of the target generation model, and perform spatial constraints and position locking on the generation region of each target person based on the person segmentation mask and / or pose skeleton information (key point heatmap, skeleton connection map, etc.) obtained from instance segmentation, so that the spatial position, pose action and identity facial features of each target person in the target generated image do not interfere with each other, thereby improving the positioning accuracy and visual consistency of multiple people generated in the same frame.
[0084] This invention preferably employs an attention-based fusion strategy, which, while preserving the independence of each feature, allows the model to emphasize corresponding features at different generation stages, thereby achieving fine-grained control over the attributes of the digital life form. Furthermore, relative position encoding is introduced into the fusion process to incorporate the spatial location information of the features; for scenarios requiring the generation of multiple characters within the same frame, a label embedding mechanism is used to distinguish features from different reference image sources, ensuring that the feature representation of each character or object remains independent and does not interfere with each other.
[0085] In one embodiment, such as Figure 4 As shown, before step S103, i.e. before determining the comprehensive constraints, the multi-graph reference digital life form generation method further includes:
[0086] S401: Concatenate and splice multiple target feature vectors to determine the comprehensive feature vector;
[0087] S402: Perform linear transformation or normalization on the comprehensive feature vector to determine the comprehensive condition vector.
[0088] As an example, in step S401, the electronic device concatenates multiple target feature vectors of different attributes along the channel dimension to form a comprehensive feature vector. For example, in the facial feature vector... The dimension is Clothing feature vector The dimension is Scene feature vector The dimension is The concatenation yields the comprehensive feature vector. , dimension The multiple target feature vectors also include attitude feature vectors. and style feature vectors The dimensions are respectively and The concatenation yields the comprehensive feature vector. , dimension .
[0089] As an example, in step S402, since the numerical range and statistical characteristics of different target feature vectors may differ, it is necessary to perform a linear transformation or normalization on the concatenated comprehensive feature vector to align the scales of each feature component and determine the comprehensive condition vector. Specifically, this involves... To perform fusion, where Concat(·) represents vector concatenation. Let be the projection weight matrix. For bias, Norm(·) indicates LayerNorm or L2 normalization. For the comprehensive condition vector, d represents the dimension of the comprehensive condition vector. For example, in the comprehensive feature vector... To fuse target feature vectors representing appearance, clothing, and setting—for example, when multiple target feature vectors include those corresponding to appearance, clothing, setting, posture, and style—these vectors are concatenated and spliced to form a comprehensive feature vector. Perform linear transformation or normalization to determine the comprehensive condition vector. If no pose reference image or style reference image is provided, it can be set to... and / or It is an empty vector or a zero vector, or skips the corresponding item in Concat.
[0090] In this embodiment, multiple extracted target feature vectors are fused to form a fused condition vector representing the comprehensive conditions for image generation. The fusion process aims to effectively integrate information such as appearance, clothing, background, and pose, ultimately yielding a comprehensive condition vector that serves as the generation condition. This comprehensive condition vector represents and encapsulates the key information from each reference image provided by the user, and is processed to ensure compatibility and synergy among them. Unlike simply stitching together multiple inputs in a rigid manner, the conditions obtained by this invention are structurally bound to correspondences such as identity, appearance, and environment, laying the foundation for high-quality generation in the next step.
[0091] In one embodiment, the target generation model includes a diffusion generation model, which includes a U-Net network, a multi-layer attention module, and a decoder;
[0092] like Figure 5 As shown, step S104 involves inputting comprehensive constraints into the target generation model to generate images, determining the initial generated image corresponding to the digital life form, including:
[0093] S501: Iterative sampling and generation are performed using a diffusion generation model, including: inputting the input latent vector corresponding to the current iteration step into the U-Net network of the diffusion generation model, injecting comprehensive constraints into the multi-layer attention module of the diffusion generation model for condition guidance, determining the current noise prediction corresponding to the current iteration step; denoising the current noise prediction corresponding to the current iteration step, and determining the output latent vector corresponding to the current iteration step.
[0094] S502: Update the current iteration step number to the new current iteration step number, where the new current iteration step number = current iteration step number - 1;
[0095] S503: When the new current iteration step is not 0, the output latent vector corresponding to the current iteration step is determined as the input latent vector corresponding to the new current iteration step, and the iterative sampling and generation using the diffusion generation model is repeated.
[0096] S504: If the current iteration step is 0, the decoder corresponding to the diffusion generation model is used to decode the output latent vector generated by the last iteration sampling to determine the initial generated image corresponding to the digital life form.
[0097] The input latent vector refers to the latent vector input during each iteration of sampling. The output latent vector refers to the latent vector output during each iteration of sampling. The latent vector is a numerical feature vector or feature map used in the diffusion generation model for adding or removing noise; it is also called a latent feature or latent space vector.
[0098] As an example, the electronic device uses a diffusion generation model as the target generation model for image generation. This diffusion generation model includes a U-Net network containing approximately six cascaded residual modules and introduces multiple attention modules at multiple scales. Each attention module includes multiple attention heads, for example, 4 to 8 attention heads per layer. In this example, the electronic device uses the diffusion generation model to process the comprehensive conditional vector, and the process of generating the initial generated image corresponding to the digital life form includes the following steps:
[0099] (1) Initial state setting: Let T represent the initial number of steps in the inverse diffusion process, for example, T=50 or T=100, corresponding to the number of iteration steps generated. A random latent vector is obtained by sampling from the standard normal distribution. Its dimensions can be 64×64×4, and this is determined as the input latent vector generated in the first iteration of sampling. For video generation, a latent vector can be initialized for each frame to be generated. Alternatively, you can initialize only the potential vector of the first frame and use it to derive the vectors of subsequent frames (see the extended scheme for video generation below).
[0100] (2) Iterative sampling generation: from Begin iteratively executing the inverse diffusion process until... Every step Perform the following operations:
[0101] (2A) Condition-guided prediction: The input latent vector corresponding to the current iteration step t The U-Net network of the diffusion generation model is input, and the comprehensive constraints are also included. Injecting attention modules into the diffusion generation model can be done, for example, through cross-modal attention layers or other control signal methods. Each conditionally guided prediction outputs the current noise prediction corresponding to the current iteration step. The current noise prediction is used to characterize the diffusion generation model's determination of the input latent vector. The noise component in.
[0102] (2B) Denoising Calculation: Based on the update rule of the selected sampling algorithm (such as DDIM, PNDM, etc.), predict the current noise corresponding to the current iteration step. Perform denoising calculations to determine the output latent vector corresponding to the current iteration step with slightly reduced noise. .
[0103] For example, when using DDIM sampling, let... For the first Step latent variables, In order to meet the conditions The noise in the prediction. One form of DDIM update is: Formula (1); where For the preset noise attenuation coefficient sequence, Let it be the cumulative value. , ; It is a preset noise scheduling coefficient sequence (usually) With iteration steps (incremental) It is by and The estimated value of the initial clear sample obtained from the prediction. It is the coefficient that adds noise. The noise is standard Gaussian noise. The above denoising calculation formula enables the diffusion generation model to gradually reduce the noise component in the latent vector, approximating a clear image. To enhance the conditional guidance effect, a classifier-free guidance method is used, combining conditional and unconditional predictions in a proportional manner: Formula (2); where, This indicates a situation where no conditions are used. For the guiding strength coefficient, we can take... The scope of implementation. In this example, the result calculated in formula (2) is... Substituting into formula (1) can improve the degree to which the generated result conforms to the input features. In this process, comprehensive constraints are considered. Each step guides the current noise prediction. This ensures that the denoising process evolves in a direction that satisfies the comprehensive constraints.
[0104] (3) Repeated iteration: Update the current iteration step number to the new current iteration step number, i.e. Determine the updated Is it equal to 0? If so... Then the output latent vector corresponding to the current iteration step is determined as the new input latent vector corresponding to the current iteration step, and the diffusion generation model is repeatedly used for iterative sampling and generation, that is, the process (2) is repeated; if Then the iterative sampling generation process is complete, and the final iteration of sampling generates a definite output latent vector. That is, a potential representation that approximates the real image.
[0105] (4) Decoding output: After the diffusion iteration is completed, the decoder of the diffusion generation model is used to generate the corresponding output latent vector from the last iteration sample. The image is decoded into a visual image and identified as the initial generated image corresponding to the digital life form. If higher resolution is required, the output can be further magnified using an image super-resolution module. For video outputs, the above process generates a corresponding latent representation for each time step (or keyframe), which is then decoded to obtain consecutive frame images. This yields a preliminary generated digital life form image or frame sequence, ready for subsequent result processing.
[0106] Key Algorithm Structure and Formula Derivation for Decoupling Characters and Background
[0107] To more clearly illustrate the internal mechanism of "multi-attribute decoupling representation + hierarchical condition injection + quality closed-loop update" in this invention, an achievable core mathematical description is given below (exemplary expression; specific parameters and network structures can be equivalently replaced according to implementation requirements).
[0108] (A) Multi-attribute conditional fusion (linear projection example)
[0109] Let the facial feature vector be... The feature vector of clothing is The scene feature vector is (Optional: posture) With style Concatenating the above vectors along the channel dimension yields the comprehensive feature vector:
[0110] The synthesis condition vector is obtained through linear layering and normalization:
[0111]
[0112] in, Let be the projection weight matrix. For bias, This indicates LayerNorm or L2 normalization. This is the composite condition vector used to generate constraints.
[0113] Multi-attribute conditional fusion (attention fusion example)
[0114] To reduce interference between attribute conditions, each attribute vector can be treated as an "attribute token," and cross-attribute attention can be performed using a query / key / value approach.
[0115]
[0116] in, It can be obtained from the feature mapping of the current generation layer. It can be obtained by linear mapping of each attribute subvector; For attention space dimensions.
[0117] (C) Diffusion Reverse Process Sampling (DDIM Example)
[0118] set up For the first Step latent variables, In order to meet the conditions The noise in the prediction. Taking DDIM as an example, first estimate the noise-free samples:
[0119]
[0120] Update according to the sampling rules:
[0121]
[0122] in, Represents the cumulative value of the noise scheduling coefficient (as...). change).
[0123] (D) Classifier-less guidance (CFG example)
[0124] To improve conditional alignment, a classifier-free guide can be used, allowing... For conditional prediction, If the prediction is based on an empty condition, then the guided prediction will be:
[0125]
[0126] in To guide the intensity.
[0127] (E) Cross-frame consistency of video scenes (example of cross-frame attention / similarity constraints)
[0128] When the output is a sequence of video frames, cross-frame attention can be introduced into the denoising network to improve the query vector of the current frame. Simultaneously, pay attention to the key / value pairs of neighboring frames (e.g., concatenate the features of neighboring frames into...). Thus, maintaining continuity of identity and background:
[0129]
[0130] Simultaneously, neighboring frame similarity constraints can be added (e.g., threshold verification of identity embedding or pixel structure similarity). When the similarity is lower than the threshold, iterative correction is performed by updating cross-frame attention weights or related attribute injection weights to improve inter-frame consistency and stability.
[0131] In one embodiment, the U-Net network includes multiple resolution layers with residual connections, and the multiple resolution layers include at least a low-resolution layer, a medium-resolution layer, and a high-resolution layer;
[0132] like Figure 6As shown, step S501 involves injecting comprehensive constraints into the multi-layer attention module of the diffusion generation model to guide the determination of the current noise prediction corresponding to the current iteration step, including:
[0133] S601: Decompose the comprehensive conditional vector into attribute sub-vectors corresponding to multiple attributes, and determine the attention weights corresponding to multiple attributes in each resolution layer based on the conditional injection parameters. The scene has the highest attention weight in the low resolution layer, clothing has the highest attention weight in the medium resolution layer, and appearance has the highest attention weight in the high resolution layer.
[0134] S602: Based on the attribute subvectors corresponding to multiple attributes and the attention weights corresponding to multiple attributes in each resolution layer, determine the attention output of each resolution layer. The attention output is determined by the following formula: , , ,in, For attention output, The query vector is determined based on the original feature map of the i-th resolution layer. and These are the Key and Value values for the i-th resolution layer, respectively. and Let be the linear mapping matrix between the Key and Value values of the i-th resolution layer. The number of attributes, Let j be the attention weight of the j-th attribute in the i-th resolution layer. Let j be the attribute subvector of the j-th attribute. The dimension of the comprehensive condition vector;
[0135] S603: Perform residual connections based on the original feature maps and attention outputs of each resolution layer to determine the enhanced feature maps corresponding to each resolution layer. Based on the enhanced feature maps corresponding to all resolution layers, determine the current noise prediction corresponding to the current iteration step.
[0136] As an example, in step S601, the electronic device can sequentially decompose the comprehensive condition vector into attribute sub-vectors corresponding to multiple attributes. For instance, when multiple target reference images include a person's facial reference image, clothing reference image, scene reference image, pose reference image, and style reference image, the attribute sub-vectors corresponding to the decomposed attributes are respectively facial sub-vectors. Clothing subvector Scene subvectors Attitude subvectors and style vector In this example, a diffusion-based generative model based on the U-Net architecture contains multiple attention modules. We introduce conditional injection parameters into these attention modules to adjust the attention weights of different attributes at different resolution layers. For example, for the high-resolution layer responsible for generating the person's appearance, we increase the attention weight for facial features; for the medium-resolution layer rendering clothing textures, we emphasize the attention weight for clothing features; and for the low-resolution layer rendering the background, we emphasize the attention weight for scene features to control the overall style. By incorporating modal features into the spatial attention calculation, we achieve partitioned control—that is, the model "focuses" on person-related features when rendering the foreground region of the person, and "focuses" on background features when rendering the background region, ensuring that each part of the generated image conforms to the corresponding reference information.
[0137] Furthermore, since pose sub-vectors are primarily connected to the medium-resolution layer and lightly penetrate the high-resolution layer, while style feature vectors are primarily connected to the low-resolution layer and lightly penetrate the medium-resolution layer, the scene attention weights are adjusted when the i-th resolution layer is a low-resolution layer. Highest, style attention weight Secondly; when the i-th resolution layer is a medium resolution layer, the attention weight of clothing... Attention weights for pose are highest. Secondly, the attention weights of different attributes in each resolution layer can be adjusted according to the actual situation.
[0138] As an example, in step S602, the electronic device can determine the Key and Value values corresponding to each resolution layer based on the attribute sub-vectors corresponding to multiple attributes and the attention weights corresponding to multiple attributes in each resolution layer. and Then, based on The attention output for each resolution layer is determined using the following formula. For example, let the attention weights corresponding to the five attributes in the i-th resolution layer be... , , , and Then the calculated , Then, based on the Key value corresponding to the i-th resolution layer. and Value Determine its corresponding attention output .
[0139] As an example, step S603: The electronic device will generate the original feature map for each resolution layer. and attention output Perform residual connections to obtain the enhanced feature map of this resolution layer, i.e. , The enhanced feature map for the i-th resolution layer is then used. Next, based on the enhanced feature maps of all resolution layers, the current noise prediction corresponding to the current iteration step is determined.
[0140] In this embodiment, comprehensive constraints are fully utilized. This guidance ensures that the initially generated image is highly consistent with the reference information in all aspects. For example, during iterative denoising, if the diffusion generation model attempts to deviate from the facial features of the reference image of a person's face, the image will be influenced by the reference image. Facial feature constraints will pull it back; if the background color and atmosphere deviate, it will be affected. Background features guide the diffusion generation model to adjust the overall color tone. Compared to diffusion generation that is unconditional or relies solely on text prompts, this invention significantly improves the controllability and accuracy of the generation results through multi-condition joint guidance. Experiments show that, with the same number of diffusion steps, introducing multiple image conditions can more accurately meet user expectations and reduce the time spent repeatedly adjusting prompts.
[0141] In one embodiment, the conditional injection parameters further include cross-frame attention or neighboring frame similarity constraints corresponding to the cross-frame attention layer;
[0142] The cross-frame attention corresponding to the cross-frame attention layer is determined by the following formula: ,in, For the cross-frame attention of the nth frame at the i-th resolution layer, Let n be the query vector for the nth frame at the i-th resolution layer. The key value of the (n-1)th frame at the i-th resolution layer. The value of the (n-1)th frame at the i-th resolution layer;
[0143] The neighboring frame similarity constraint is determined based on the similarity between the current frame's latent vector and the previous frame's latent vector.
[0144] Step S104 involves inputting the comprehensive constraint conditions into the target generation model to generate images and determine the initial generated image corresponding to the digital life form. This includes: injecting the cross-frame attention or neighboring frame similarity constraints corresponding to the cross-frame attention layer into the diffusion generation model; processing the comprehensive condition vector using the diffusion generation model to determine the initial generated video corresponding to the digital life form; the initial generated image of the nth frame in the initial generated video is determined based on the feature map of the nth frame and the cross-frame attention between the (n-1)th and nth frames, or the initial generated images of two adjacent frames in the initial generated video are constrained by the neighboring frame similarity constraints.
[0145] As an example, when it is necessary to generate the initial generated video corresponding to the digital life form, since the initial generated video generally includes multiple consecutive initial generated images, in the process of outputting the initial generated video, the cross-frame attention or neighboring frame similarity constraint corresponding to the cross-frame attention layer in the time dimension can be introduced into the diffusion generation model. The feature map or latent vector corresponding to the initial generated image of the previous frame is introduced into the generation process of the initial image of the current frame to ensure the inter-frame consistency of the initial generated images of two adjacent frames and reduce flickering and identity drift.
[0146] In one example, a cross-frame attention layer can be introduced into the U-Net network, allowing the current frame feature map and the previous frame feature map to jointly compute multi-head attention in the cross-frame attention layer, thereby obtaining temporally continuous information. For example, let the feature map of the nth frame in the i-th layer (i.e., the feature map of the current frame) be... The feature map of the (n-1)th frame at layer i (i.e., the feature map of the previous frame) is The calculated cross-frame attention is Then, based on the feature map of the nth frame at the i-th layer (i.e., the feature map of the current frame) and the cross-frame attention, a residual connection is performed to obtain the enhanced feature map, i.e. The enhanced feature map is then used as the initial generated image for the nth frame.
[0147] In another example, the neighboring frame similarity constraint of the forward diffusion generation model is specifically determined by calculating the similarity between the latent vector of the current frame and the latent vector of the previous frame. Determining the inter-frame similarity requires ensuring that the inter-frame similarity is greater than a preset similarity, so as to ensure that the initially generated images of similar frames remain consistent, resulting in a smooth and flicker-free transformation of the generated digital life form over time. For example, let... and Representing the generated first Frame and the The cosine similarity is defined based on the latent vector representation of a frame at a specific layer: By maximizing during training or generation (Or minimize their difference) can encourage adjacent frames to maintain consistency, thus making the character's image change smoothly and without flickering over time. Combining the above measures, in generating the initial video corresponding to the digital life form, it is possible to maintain the high fidelity of the initial generated image of a single frame while maintaining the temporal consistency of multiple initial generated images, generating coherent and natural digital human video clips with a unified character identity.
[0148] In this embodiment, during the initial video generation process, cross-frame attention layers or neighboring frame similarity constraints ensure inter-frame consistency. Furthermore, the generated multi-frame initial images can be smoothed. For example, by calculating pixel differences between adjacent frames or changes in key character positions, if slight jumps in scene or character pose are detected, interpolation algorithms can be used to generate transition frames between the current and previous frames, or optical flow technology can be used to perform motion compensation on the current frame to align the character positions between preceding and following frames. Additionally, frame stabilization filtering can be applied to the frame sequence to eliminate minor jitter. Through these processes, the actions and expressions in the initial generated video corresponding to the digital life form are ensured to be coherent and natural. Experiments show that the average SSIM (Structural Similarity Between Frames) of the video processed by this invention can reach over 0.93, while ordinary frame-by-frame diffusion generation is only about 0.68. This indicates that the character video processed by this invention significantly outperforms the frame-by-frame generation method without temporal constraints in terms of SSIM and other metrics, meeting the requirements for film-level coherence.
[0149] In this embodiment, the U-Net network of the diffusion generation model will include a comprehensive condition vector. The combined constraints of the injected parameters are used as conditions, which are injected into different layers of the model through cross-modal attention mechanisms or in the form of control signals. Specifically, multi-layer attention modules and decoders, etc., diffuse generative models receive the combined conditional vector. Subsequent iterative sampling generation: Initial time Starting with random noise, in the comprehensive condition vector Guided by the principle of gradual denoising, samples are generated step by step. After iteration A clear image representation is obtained at that time. The initial generated image corresponding to the synthesized digital life form is then obtained through a decoder. If the target output is video, the diffusion sampling process is repeated to generate multiple consecutive frames, and a cross-frame attention mechanism is introduced into the diffusion generation model to ensure the consistency of adjacent frames. The number of iterations for diffusion sampling can be set according to the required image quality, typically in the range of 20 to 100 steps. For video generation, in addition to ensuring inter-frame consistency through cross-frame attention, a keyframe generation strategy can also be adopted: for example, a keyframe is generated every 5 or 10 frames, and this keyframe is used as a reference to guide the generation of intermediate transition frames, so as to improve generation efficiency and reduce computational overhead while ensuring video continuity.
[0150] In one embodiment, the initially generated image includes an initially generated video; such as Figure 7 As shown, after step S104, that is, after determining the initial generated image corresponding to the digital life form, the multi-image reference digital life form generation method further includes:
[0151] S701: Obtain the target reference speech, perform speech feature analysis on the target reference speech, and determine the phoneme feature sequence or speech feature sequence corresponding to the target reference speech;
[0152] S702: Perform mouth shape conversion on the phoneme feature sequence or speech feature sequence to obtain the pronunciation mouth shape sequence;
[0153] S703: Based on the frame rate synchronization mechanism, it controls the mouth movements of the initial generated video corresponding to the digital life form to make them consistent with the lip-sync sequence.
[0154] Among them, the target reference speech is the reference speech used to specify the character's speech.
[0155] As an example, the electronic device analyzes the input target reference speech to extract speech features. Preferably, a Mel-frequency cepstral coefficient (MFCC) extraction algorithm can be used to obtain a series of coefficients reflecting the spectral characteristics of the speech, or a pre-trained speech recognition model based on the Transformer architecture can be used to perform end-to-end feature extraction on the target reference speech, thereby obtaining a phoneme feature sequence or speech feature sequence corresponding to the target reference speech. Then, the extracted speech feature sequence is converted into a corresponding lip-sync sequence according to a preset lip-sync mapping table, that is, the lip-sync (Viseme) corresponding to each speech segment on the time axis is determined. To ensure that the mouth movements of the digital life form are strictly synchronized with the target reference speech, the electronic device uses a frame rate synchronization mechanism to control the update rate of the lip-sync animation: for example, a lip-sync image can be generated every 40 milliseconds (equivalent to 25 frames / second), and this lip-sync frame can be applied to the generated digital life form character image, so that its mouth shape matches the speech pronunciation in the current time period. Through the above-mentioned speech analysis, lip-syncing and frame rate synchronization strategies, the generated digital life form video can change its lip shape in response to the input speech, achieving a realistic speech synchronization effect and ensuring that the mouth movements and audio rhythm of the digital life form are accurate and consistent when interacting with the user.
[0156] In one embodiment, such as Figure 8 As shown, step S105 involves optimizing the quality of the initially generated image corresponding to the digital life form and determining the quality detection results corresponding to multiple attributes, including at least two of the following steps:
[0157] S801: Calculate the similarity between the facial feature vector in the initially generated image and the facial feature vector in the reference image of the person's face. When the facial similarity between the two is less than the first preset similarity, iteratively update the attention weight corresponding to the face in the comprehensive constraint conditions until the facial similarity between the two is not less than the first preset similarity.
[0158] S802: Perform texture defect and blur detection on the clothing area in the initial generated image. If the detection results do not meet the quality requirements, perform local image regeneration on the clothing area.
[0159] S803: Smooth the edge regions in the initially generated image and adjust the color and brightness of the edge regions using the lighting information in the scene reference image. The edge regions are the areas where the foreground and background of the person meet.
[0160] S804: Calculate the similarity between the skeletal keypoints in the initially generated image and the skeletal keypoints in the pose reference image. When the pose similarity between the two is less than the second preset similarity, iteratively update the attention weights corresponding to the pose in the comprehensive constraint conditions until the pose similarity between the two is not less than the second preset similarity.
[0161] S805: Calculate the similarity between the style feature vector in the initially generated image and the style feature vector corresponding to the style reference image. When the style similarity between the two is less than the third preset similarity, iteratively update the attention weight corresponding to the style in the comprehensive constraint conditions until the style similarity between the two is not less than the third preset similarity.
[0162] As an example, in step S801, the electronic device can use a pre-trained face recognition algorithm to extract facial feature vectors from the reference image of the person's appearance and the initially generated image, respectively. Then, a similarity algorithm is used to calculate the similarity between the two facial feature vectors to determine their corresponding facial similarity. This facial similarity is then compared with a first preset similarity. If the facial similarity is less than the first preset similarity (preferably 0.85 to 0.99), the weights of the facial features in the comprehensive condition vector are adjusted and / or the image generation step is re-executed until the facial similarity is not less than the first preset similarity. The weight adjustment is achieved by scaling the facial feature vector by multiplying it by a predetermined coefficient to ensure the realism and consistency of the digital life form's facial features. In this example, a face recognition algorithm and a similarity algorithm are used to verify the facial similarity between the initially generated image and the reference image of the person's appearance, ensuring that the facial similarity remains consistent. The first preset similarity is a pre-set threshold for evaluating whether facial features meet the similarity standard, and can be set to 0.85 to 0.99.
[0163] As an example, in step S802, the electronic device can use a pre-set defect detection algorithm or fuzzy detection algorithm to detect texture defects or fuzziness in the clothing area of the initially generated image. If a texture defect is found or the fuzziness is greater than a preset fuzziness, the detection result is determined to be inconsistent with the quality requirements. At this time, the non-clothing areas in the initially generated image can be locked unchanged, and local image regeneration can be performed on the clothing area in the initially generated image. The texture details provided by the clothing reference image are then reapplied to the clothing area in the initially generated image so that the clothing details of the generated digital life form are consistent with the clothing reference image. For example, when certain local areas (such as the clothing area) of the initially generated image have texture defects or distortion, local diffusion regeneration can be performed on the local area, performing only a small number of iterations (such as about 10 steps) to enhance the details without affecting other intact areas of the image.
[0164] As an example, in step S803, the electronic device can smooth the edge region corresponding to the boundary between the foreground and background of the person in the initially generated image. Specifically, filtering, blurring, or alpha blending algorithms can be used to eliminate harsh boundary lines. Gaussian blur, bilateral filtering, or guided filtering algorithms are preferred for edge smoothing. The color and brightness of the foreground edge region of the person in the initially generated image are adjusted based on the lighting information of the scene reference image, making the digital life form blend harmoniously with the background environment. For example, for the edge region where the foreground and background of the person in the initially generated image meet, a small-sized smoothing filter can be used to blur the boundary (e.g., using a 3×3 Gaussian blur kernel) to eliminate harsh boundary lines and achieve a natural foreground-background blending effect. Similarly, if there are harsh edges where the person's outline blends with the new background, transition filtering or background segmentation-based alpha blending can be used to smooth them, and the color of the person's edges can be adjusted according to the background lighting to make the person blend more harmoniously with the environment. If necessary, color correction is performed on the entire image to bring the tone style of the generated result closer to the reference scene, avoiding color distortion.
[0165] As an example, in step S804, the electronic device can perform skeletal keypoint verification on the pose features of the initially generated image, optimizing it in the following way: First, extract the skeletal keypoints (such as shoulders, elbows, hips, knees, etc.) of the person in the initially generated image, calculate the Euclidean distance with the skeletal keypoints in the pose reference image, determine the pose similarity, and then compare the pose similarity with a second preset similarity (preferably 0.8 to 0.85). If the pose similarity is less than the second preset similarity, adjust the attention weights corresponding to the pose in the comprehensive constraint conditions based on the skeletal information of the pose reference image. Specifically, increase the attention weights of the pose attributes in the mid-resolution layer of U-Net. For example, increase the pose weights in the mid-resolution layer from 0.3 to 0.5, and regenerate the image. For the initially generated video, additionally calculate the smoothness of the motion trajectory of the skeletal keypoints in adjacent frames. If there are abrupt joint changes (such as an elbow angle change exceeding 30° between frames), use an optical flow interpolation algorithm to correct the pose of the intermediate frames to ensure smooth movement. The second preset similarity is a pre-set threshold for evaluating whether the pose features meet the similarity standard, and can be set to 0.8 to 0.85.
[0166] As an example, in step S805, the electronic device verifies the consistency between the overall style of the initially generated image and the style reference image, optimizing it in the following way: First, it extracts style features (such as tone distribution, brushstroke texture, and contrast) from the initially generated image and the style reference image, and calculates style similarity using a style distance algorithm (such as Gram matrix distance). The style similarity is compared with a preset similarity (preferably 0.75–0.8). If the style similarity is lower than the third preset similarity, the linear transformation coefficients of the style sub-vector in the comprehensive conditional vector are adjusted. For example, the weight of the style sub-vector is adjusted from 0.2 to 0.4, or the injection weight of style attributes is increased in the low-resolution layer of the diffusion generation model (responsible for global style), and the image is regenerated. For the initially generated video, it ensures that the variance of style features in all frames is less than a preset value (e.g., tone deviation does not exceed 10%). If the style of a frame shifts, the facial and clothing features of that frame are locked unchanged, and only local regeneration is performed on the style layer to avoid overall style breakage. The third preset similarity is a pre-set threshold for evaluating whether style features meet the similarity standard, and can be set to 0.75–0.8.
[0167] To more clearly illustrate the technical solution of the present invention, a specific application scenario and accompanying drawings are provided below. Suppose a user wants to generate the image of their digital life form: the user uploads a front-facing portrait photo as a reference image for the person's appearance, a picture of their favorite clothing (e.g., an evening gown) as a clothing reference image, and a picture of a target scene (e.g., a sunset view on a beach) as a scene reference image, hoping the digital life form will be standing and situated within that scene.
[0168] (1) The system reads the user's profile picture and performs feature extraction to obtain the facial feature vector. The system captures the user's facial features, such as facial structure and face shape; it also reads images of evening gowns and extracts features to obtain garment feature vectors, including style and texture characteristics. The dress image is then matched with a pre-defined standing human model to generate a segmentation image of the person wearing the dress; the beach background image is read, and scene feature vectors such as color and environmental features are extracted. The system sets this image as the output background layer; since the user did not provide a specific pose reference image, the system uses the default standing pose skeleton as the background. .
[0169] (2) Integrate the above-mentioned identity, dress and beach scene features into a comprehensive condition vector. These conditions are injected into multiple layers of the diffusion model through an attention mechanism. For example, the model is forced to pay attention during the face generation stage. Features, for reference during the clothing drawing stage Details applied during the background rendering stage Color tone and environmental information.
[0170] (3) The diffusion generation model starts with random noise and gradually generates images over approximately 50 iterations. During this process, the model, on the one hand, determines the image based on random noise. The model "draws" the user's face and, based on the dress's texture, "draws" the corresponding clothing details. The background is either a reference beach image or a similar beach sunset scene redrawn based on its features. Ultimately, the model generates a photo of "the user standing on the beach in an evening gown."
[0171] (4) Result verification and processing: The system compares the face of the person in the generated image with the user's headshot photo. The similarity reaches 0.98, proving that the person's identity features are highly consistent. The system also checks that the dress pattern matches the reference clothing perfectly, and that the lighting direction of the person and the beach background is consistent. Since the result is satisfactory, no further local redrawing or adjustment is needed, and the system directly outputs the image to the user.
[0172] In a comparative experiment, the method of this invention was evaluated against traditional GAN generation models and single-conditional diffusion generation models. The results show that this invention achieves superior performance in key metrics: regarding person identity similarity, the average cosine similarity between the facial features of the generated image and the reference photo reaches over 0.95, significantly higher than the approximately 0.85 of the comparative method; regarding image structure consistency, the structural similarity (SSIM) between the generated result and the source reference exceeds 0.90, also superior to the comparative method (approximately 0.80). Furthermore, in terms of generation efficiency, this invention, by performing diffusion sampling in the latent space, controls the generation time of a single 512×512 resolution image to approximately 5 seconds, comparable to single-conditional diffusion generation models. Although this speed is slightly lower than a single forward inference (<1 second) of a pre-trained GAN model, this approach eliminates the need to retrain the model for each individual user, significantly reducing the preparation time for personalized generation and demonstrating greater overall efficiency. The above comparative experimental results fully verify that the present invention achieves a satisfactory generation speed while ensuring high face similarity and image fidelity, demonstrating a significant improvement over existing GAN methods and single-conditional diffusion generation models.
[0173] Without departing from the core idea of this invention, those skilled in the art can make various modifications and improvements to the above technical solutions. Several possible alternative implementations or equivalent variations are listed below to illustrate that the scope of protection of this invention is not limited to the specific implementations described above:
[0174] Replacement of Feature Extraction Algorithms: This invention does not limit the extraction of various features such as appearance, clothing, and scene to specific models or algorithms. For example, facial features can be replaced by parameter vectors extracted based on 3D face reconstruction models; clothing features can also be extracted using traditional image processing methods to extract manual features such as the main color tone and texture pattern of clothing instead of vectors obtained from deep learning. As long as information representations that can distinguish each attribute can be obtained, they can be used for fusion in this invention. Furthermore, in addition to using human pose estimation algorithms to extract skeletal keypoint coordinates, pose information can also be input in different forms, such as providing a silhouette of the person or a simple sketch; background features can be obtained not only by directly using the output of the image encoder, but also by pre-classifying the scene into several labels (such as "beach," "office," etc.) and using a set of labels to represent the scene conditions. Different feature extraction schemes can be replaced according to application requirements and resource conditions, and these changes are all within the scope of protection of this invention.
[0175] Variations in Multi-Condition Fusion Methods: Besides the aforementioned cross-modal attention-based fusion mechanism, the multi-modal fusion of this invention can also be achieved through other methods. For example, a multi-branch conditional control network scheme can be employed: an independent control sub-network is configured for each condition, mapping its respective features into control signals applied to different layers of the diffusion model. During generation, these control signals work together to satisfy multi-modal constraints. This method is essentially equivalent to fusion performed by multiple parallel modules, with effects similar to the single fusion module of this invention. Furthermore, a noise space decoupling sampling method can be used: partial images are generated separately for different conditions (e.g., first generating a face based on identity features, then refining the clothing based on clothing features), and finally synthesizing the results of each part into a complete image according to rules in the noise latent space (similar to the "noise fusion" technique in some studies, i.e., fusing the results obtained from multiple diffusion sampling segments). Whether fusion is performed within the model through an attention mechanism or by combining multiple pre-trained networks, the purpose of these variant schemes is to integrate multiple conditions to guide generation; therefore, they should all be considered equivalent implementations of the multi-modal fusion of this invention.
[0176] Replacement of Generative Model Type: While a diffusion model is preferred as the generation engine, this invention does not preclude the use of other image generation models. For example, a modified Generative Adversarial Network (GAN) or other probabilistic generative models can be used to achieve similar functionality: fused features are used as input to the GAN generator to generate images that satisfy multiple conditional constraints in a multi-input, multi-output structure. Alternatively, new generative frameworks that emerge in the future (such as the Transformer model based on the diffusion principle, streaming generative models, etc.) with more efficient generation capabilities can also be used to replace the diffusion model module. The ideas of multi-condition processing and control in this invention are also applicable to these new models—simply feed multiple image conditions as generation conditions into the new model. Therefore, this invention is open in its choice of specific generative models and is not limited to a particular algorithm. As long as the goal of generating realistic digital life forms based on multiple reference conditions can be achieved, it falls within the scope of patent protection of this invention.
[0177] Alternative Implementations for Video Generation: For the generation of digital life form videos, there are alternative implementation paths to those described in the embodiments above. For example, a "keyframe + frame interpolation" strategy can be adopted: first, static digital life form images at several key poses are generated using the method of this invention; then, a coherent video sequence is generated between the keyframes through video frame interpolation or animation transition techniques. This is equivalent to ensuring that the keyframes meet the conditional requirements before transitioning them, which reduces the computational load of direct video diffusion generation. Another approach is to use a dedicated video diffusion model, where the model learns to generate a series of coherent frames simultaneously during training, outputting a multi-frame sequence at once after inputting multiple conditions. This invention can be used in conjunction with such models: simply extend the fusion features to time-series conditions and input them into the video diffusion model. Different video generation strategies differ in effect and complexity, but all improve inter-frame consistency; therefore, they are all included within the concept of "digital life form video generation" in this invention.
[0178] Variations in background processing: Regarding the introduction of background scenes, this invention supports both directly using a reference background image and generating a new background that matches the reference style. Alternative solutions include: not using the original background image during model generation, but redrawing the background based on the scene feature vector $f_{bg}$, thus avoiding the foreground / background edge blending problem that may result from direct stitching; or, after generating the character image, using an image synthesis algorithm to blend the character and background, for example, first generating a character image with a transparent background, then overlaying it onto the background image and using a GAN model to adjust the blending area. These methods each have their advantages and disadvantages: directly generating the background from a diffusion model can make the character and background more integrated, but sometimes the reference background cannot be accurately reproduced; later background blending can retain the details of the original background image, but requires additional processing of the character's edge transitions. Regardless of the method used, this invention maintains the consistent idea of "separate control of the character and background," and any flexible combination of background and character images that can be achieved constitutes an equivalent transformation of this invention.
[0179] Variations of Identity Feature Injection: To ensure consistency in personal identity, this invention extracts facial vectors and injects them into the generative model. Of course, other approaches exist. For example, a few-shot model fine-tuning method can be used: after a user provides their photo, the pre-trained diffusion model is quickly fine-tuned, introducing a new feature token representing the user's identity, allowing the model to "remember" the facial features. Then, during generation, using this feature token as a condition, the corresponding person can be directly generated. This method also achieves personalization, but requires a certain training cost each time a new person is generated. Alternatively, text prompts can be used, such as converting user facial features into descriptive text (e.g., "a male who looks similar to someone") and combining it with the generative model. Although these variations follow different paths, their purpose is to integrate specific identity information into the generation process. Therefore, they are also alternative technical means to achieve the goal of ensuring personal identity in this invention and should be included within the scope of protection of this invention.
[0180] Introduction of Other Modalities: This invention currently describes multi-condition input for the image modality, but it can also be extended to incorporate non-image modalities as additional conditions into the generation process. For example, text descriptions can be used as supplementary conditions to specify details that are not easily reflected in images (such as the personality traits, facial expressions, and tone of voice of digital life forms); speech can be used to drive the lip movements of digital life forms, synchronizing with audio to make them appear to be speaking; three-dimensional depth information can be used to more accurately represent the geometric relationship between characters and scenes. If these conditions are added, the implementation only requires adding corresponding feature extraction and fusion processing, and the process is similar to the existing framework. Although these additional modalities are not described in detail in the preferred embodiment, their inclusion will not change the core process; it only adds input sources and adjusts the fusion module. Therefore, they fall within the foreseeable extensions of this invention and should also be within the scope of protection of this patent.
[0181] Hardware Deployment Changes: The method of this invention can be executed on a cloud server or deployed on a terminal device. When implemented on devices with limited computing power, such as mobile devices, lightweight models or step-by-step processing can be used to meet real-time requirements. For example, time-consuming diffusion sampling can be completed in the cloud first, and then the results can be transmitted back to the terminal; or large models can be pruned and quantized to enable them to run at an acceptable speed on mobile devices. Furthermore, dedicated AI acceleration chips or FPGAs can be used to implement the computation of some modules to further improve efficiency. These all fall within the scope of engineering optimizations for implementing this invention and can be considered equivalent implementation methods for patent purposes, without affecting the protection of the technical solution of this invention.
[0182] The multi-image reference digital life form generation method provided in this invention has the following advantages compared with the prior art.
[0183] Multi-source image and multi-dimensional conditions are introduced: This invention, for the first time, comprehensively utilizes multiple reference images from the user (including facial features, clothing, and scenes) as generation conditions to achieve customized generation of personalized digital lifeforms. Unlike traditional solutions that rely solely on a single image or text prompt, this invention can simultaneously accept input from multiple dimensions and transform them into effective constraints on the generation process, ensuring that the generated result fully meets the user's expectations in terms of face, clothing, and background. This multi-condition driven design truly endows the digital lifeform with "personalized customization" attributes, significantly improving the richness and accuracy of the generated results.
[0184] Feature Decoupling Extraction and Independent Representation: To address the problem of mutual interference among multiple source reference inputs, this invention designs a feature extraction and decoupling mechanism categorized by different attributes. Independent modules extract features such as person identity, clothing, environment, and posture, removing irrelevant information during extraction to achieve independent representation of each feature. This effectively avoids feature "cross-contamination" between different conditions. Features extracted by each module do not contaminate each other, and each performs its specific function during fusion. For example, facial features only represent a person's appearance and identity, excluding clothing information; clothing features describe the appearance of clothing but do not carry facial information. This innovation ensures that the person's image is not distorted during multi-condition synthesis, fundamentally solving the problem of feature confusion in previous multi-condition synthesis methods.
[0185] Attention Mechanism for Multimodal Feature Fusion: The fusion module of this invention introduces cross-modal attention and conditional control mechanisms to achieve collaborative binding and regional control of multiple features such as identity, clothing, and environment. By embedding fused features into the attention calculations at different levels of the generation model, the model can dynamically emphasize relevant reference conditions during detailed rendering, achieving a finely controllable effect. Unlike simple feature splicing or concatenating multiple conditions, this invention fuses information through attention weight allocation, resulting in higher expressive flexibility and accuracy. In particular, it introduces a "two-level" attention control scheme (e.g., controlling the posture structure first and then the appearance texture) and a reference source location encoding strategy to ensure that even with multiple reference objects, the model can distinguish their respective sources and render each part of the content in an orderly manner. This multimodal fusion mechanism is the key technological foundation for achieving "what you see is what you get" multi-image reference digital life generation.
[0186] High-fidelity generation based on a diffusion generation model: This invention employs a diffusion probability model as the generation engine, significantly improving the realism and detail quality of digital lifeform images. While diffusion generation is a mature generation technique, this invention has made targeted improvements and optimizations for generating multi-condition personalized multi-image reference digital lifeforms. For example, by introducing a dedicated conditional control mechanism to integrate external conditions into the diffusion process, the model can strictly adhere to input constraints while maintaining its original image detail modeling capabilities. Furthermore, through a multi-scale progressive sampling strategy, a low-resolution overall image is first generated and then refined to high resolution, ensuring that local details such as facial features and clothing textures are fully depicted. Therefore, this invention successfully applies the powerful generation capabilities of the diffusion generation model to complex digital lifeform synthesis scenarios, achieving high fidelity imaging effects close to real photographs while ensuring identity matching—something previously difficult to achieve with methods such as GANs. Combining the diffusion generation model with multi-condition fusion is also a key technical highlight and innovative protection point of this invention.
[0187] Inter-frame consistency in live-action digital life form videos: This invention, based on diffusion generation, solves the problem of discontinuity between frames in digital life form videos by introducing time-series attention constraints and pose continuity guidance. Its core lies in using the generation information of the previous frame to guide the next frame, ensuring a high degree of consistency in the appearance and movements of the character across consecutive frames. For example, a cross-frame attention module is introduced, allowing the features of the current frame to focus on the features of the corresponding parts in the previous frame, thus "inheriting" details from the previous frame; or the edge / depth information of the character in the previous frame is used as an additional condition input to the control module, allowing the model to "remember" the character's state at the previous instant and smoothly transition to the next frame. This greatly reduces inter-frame jitter and identity drift, making the generated digital life form videos stable in identity features and with consistent movements. This invention includes the above-mentioned video consistency technology as an important protected content, filling the technological gap in extending personalized digital life forms from static images to dynamic videos.
[0188] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0189] This invention provides a multi-image reference digital life form generation device, which corresponds one-to-one with the multi-image reference digital life form generation method described in the above embodiments. The multi-image reference digital life form generation device includes:
[0190] The reference image acquisition module 901 is used to acquire multiple target reference images corresponding to different attributes. The multiple target reference images include at least two of the following: a person's appearance reference image, a clothing reference image, a scene reference image, a pose reference image, and a style reference image.
[0191] The feature extraction module 902 is used to extract features from multiple target reference images respectively and determine multiple target feature vectors. The multiple target feature vectors include at least two of the following: facial feature vector, clothing feature vector, scene feature vector, posture feature vector, and style feature vector.
[0192] The constraint determination module 903 is used to determine the comprehensive constraint conditions. The comprehensive constraint conditions include at least a comprehensive condition vector and condition injection parameters. The comprehensive condition vector includes attribute sub-vectors corresponding to at least two attributes among appearance, clothing, scene, posture and style. The condition injection parameters include the weights of multiple attributes in different resolution layers of the pre-trained target generation model. The weights are used to indicate the degree of influence of the attributes on different resolution layers.
[0193] The image generation module 904 is used to input comprehensive constraints into the target generation model to generate images and determine the initial generated image corresponding to the digital life form. The generation features of different attributes in the initial generated image are constrained by the comprehensive constraints.
[0194] The quality optimization module 905 is used to optimize the quality of the initial generated image corresponding to the digital life form and determine the quality detection results corresponding to multiple attributes. When the quality detection result corresponding to any attribute is a failure, the comprehensive constraint conditions are iteratively updated based on the quality detection results and the image is regenerated until the quality detection results corresponding to all attributes are all passed, and then the target generated image corresponding to the digital life form is determined.
[0195] Specific limitations regarding the multi-image reference digital life form generation device can be found in the limitations of the multi-image reference digital life form generation method described above, and will not be repeated here. Each module in the aforementioned multi-image reference digital life form generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in an electronic device, or stored in the memory of an electronic device in software form, so that the processor can call and execute the operations corresponding to each module.
[0196] This invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the multi-image reference digital life form generation method described in the above embodiments. To avoid repetition, it will not be described again here.
[0197] This invention provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the multi-graph reference digital life form generation method described in the above embodiments. To avoid repetition, it will not be described again here.
[0198] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for generating a multi-image reference digital life form, characterized in that, include: Obtain multiple target reference images corresponding to different attributes. The multiple target reference images include at least two of the following: person appearance reference image, clothing reference image, scene reference image, pose reference image, and style reference image. Feature extraction is performed on the multiple target reference images to determine multiple target feature vectors. The multiple target feature vectors include at least two of the following: facial feature vector, clothing feature vector, scene feature vector, posture feature vector, and style feature vector. Determine comprehensive constraints, which include at least a comprehensive condition vector and condition injection parameters. The comprehensive condition vector includes attribute sub-vectors corresponding to at least two attributes among appearance, clothing, scene, posture, and style. The condition injection parameters include weights of multiple attributes at different resolution layers of a pre-trained target generation model. The weights are used to indicate the degree of influence of the attributes on different resolution layers. The comprehensive constraints are input into the target generation model to generate images, and the initial generated image corresponding to the digital life form is determined. The generation features of different attributes in the initial generated image are constrained by the comprehensive constraints. The initial generated image corresponding to the digital life form is optimized for quality, and the quality detection results corresponding to multiple attributes are determined. When the quality detection result corresponding to any attribute is a failure, the comprehensive constraint conditions are iteratively updated based on the quality detection result and the image is regenerated until the quality detection results corresponding to all attributes are all successful, and then the target generated image corresponding to the digital life form is determined.
2. The method according to claim 1, characterized in that, The step of extracting features from multiple target reference images to determine multiple target feature vectors includes at least two of the following steps: The face region of the reference image is cropped and aligned to determine the face region image. Facial features are extracted from the face region image to determine the facial feature vector. Background removal is performed on the clothing reference image to determine the clothing region image, and clothing features are extracted from the clothing region image to determine the clothing feature vector; The scene reference image is scaled to obtain a scene reference image of a preset size. Scene features are extracted from the scene reference image of the preset size to determine the scene feature vector. Human image segmentation is performed on the pose reference image to determine the human foreground image, and pose features are extracted from the human foreground image to determine the pose feature vector; The style reference image is subjected to content contour blurring and style texture enhancement processing to obtain a style texture image. Style features are extracted from the style texture image to determine the style feature vector.
3. The method according to claim 1, characterized in that, The step of extracting features from the multiple target reference images to determine multiple target feature vectors further includes: The target reference image is subjected to person detection and instance segmentation to determine the facial feature vectors, segmentation masks and bounding box information corresponding to multiple target persons in the target reference image. Based on the segmentation mask and bounding box information corresponding to each target person, the location code corresponding to each target person is determined; Associate the facial feature vectors and location codes corresponding to multiple target individuals to determine the person feature vectors corresponding to multiple target individuals; The multi-image reference digital life form generation method further includes: determining a comprehensive condition vector corresponding to multiple target characters based on character feature vectors corresponding to multiple target characters, as well as at least one of clothing feature vectors, scene feature vectors, posture feature vectors and style feature vectors; The comprehensive condition vectors corresponding to multiple target characters are injected into different generation regions or different generation branches of the target generation model, and the generation region of each target character is constrained based on the character segmentation mask and / or pose skeleton information.
4. The method according to claim 1, characterized in that, Before determining the comprehensive constraints, the multi-graph reference digital life form generation method further includes: Multiple target feature vectors are concatenated and spliced together to determine a comprehensive feature vector; The comprehensive feature vector is subjected to linear transformation or normalization to determine the comprehensive condition vector.
5. The method according to claim 1, characterized in that, The target generation model includes a diffusion generation model, which includes a U-Net network, a multi-layer attention module, and a decoder. The step of inputting the comprehensive constraints into the target generation model to generate images and determine the initial generated image corresponding to the digital life form includes: Iterative sampling generation using a diffusion generation model includes: inputting the input latent vector corresponding to the current iteration step into the U-Net network of the diffusion generation model; injecting the comprehensive constraint conditions into the multi-layer attention module of the diffusion generation model for condition guidance; determining the current noise prediction corresponding to the current iteration step; and performing denoising processing on the current noise prediction corresponding to the current iteration step to determine the output latent vector corresponding to the current iteration step. Update the current iteration step number to a new current iteration step number, where the new current iteration step number = current iteration step number - 1; When the new current iteration step is not 0, the output latent vector corresponding to the current iteration step is determined as the input latent vector corresponding to the new current iteration step, and the iterative sampling generation using the diffusion generation model is repeated. If the new current iteration step is 0, the decoder corresponding to the diffusion generation model is used to decode the output latent vector generated by the last iteration sampling to determine the initial generated image corresponding to the digital life form.
6. The method according to claim 5, characterized in that, The U-Net network includes multiple resolution layers with residual connections, and the multiple resolution layers include at least a low-resolution layer, a medium-resolution layer, and a high-resolution layer; The step of injecting the comprehensive constraints into the multi-layer attention module of the diffusion generation model for conditional guidance, and determining the current noise prediction corresponding to the current iteration step, includes: The comprehensive condition vector is decomposed into attribute sub-vectors corresponding to multiple attributes. Based on the condition injection parameters, the attention weights corresponding to multiple attributes in each resolution layer are determined. The scene has the highest attention weight in the low resolution layer, clothing has the highest attention weight in the medium resolution layer, and appearance has the highest attention weight in the high resolution layer. Based on the attribute subvectors corresponding to multiple attributes and the attention weights corresponding to multiple attributes in each resolution layer, the attention output of each resolution layer is determined by the following formula: , , ,in, For attention output, The query vector is determined based on the original feature map of the i-th resolution layer. and These are the Key and Value values for the i-th resolution layer, respectively. and Let be the linear mapping matrix between the Key and Value values of the i-th resolution layer. The number of attributes, Let j be the attention weight of the j-th attribute in the i-th resolution layer. Let j be the attribute subvector of the j-th attribute. The dimension of the comprehensive condition vector; Residual connections are made based on the original feature maps of each resolution layer and the attention output to determine the enhanced feature maps corresponding to each resolution layer. Based on the enhanced feature maps corresponding to all resolution layers, the current noise prediction corresponding to the current iteration step is determined.
7. The method according to claim 1, characterized in that, The conditional injection parameters also include cross-frame attention or neighboring frame similarity constraints corresponding to the cross-frame attention layer. The cross-frame attention corresponding to the cross-frame attention layer is determined by the following formula: ,in, For the cross-frame attention of the nth frame at the i-th resolution layer, Let n be the query vector for the nth frame at the i-th resolution layer. The key value of the (n-1)th frame at the i-th resolution layer. The value of the (n-1)th frame at the i-th resolution layer; The neighboring frame similarity constraint is determined based on the similarity between the current frame's latent vector and the previous frame's latent vector; The step of inputting the comprehensive constraints into the target generation model to generate images and determine the initial generated image corresponding to the digital life form includes: The cross-frame attention or neighboring frame similarity constraint corresponding to the cross-frame attention layer is injected into the diffusion generation model. The diffusion generation model is used to process the comprehensive condition vector to determine the initial generated video corresponding to the digital life form. The initial generated image of the nth frame in the initial generated video is determined based on the feature map of the nth frame and the cross-frame attention between the (n-1)th frame and the nth frame. Alternatively, the initial generated images of two adjacent frames in the initial generated video are constrained by the neighboring frame similarity constraint.
8. The method according to claim 1, characterized in that, The initially generated image includes the initially generated video; After determining the initial generated image corresponding to the digital life form, the multi-image reference digital life form generation method further includes: Obtain target reference speech, perform speech feature analysis on the target reference speech, and determine the phoneme feature sequence or speech feature sequence corresponding to the target reference speech; Perform lip-sync conversion on the phoneme feature sequence or the speech feature sequence to obtain the pronunciation lip-sync sequence; Based on the frame rate synchronization mechanism, the mouth movements of the initial generated video corresponding to the digital life form are controlled to make them consistent with the pronunciation mouth shape sequence.
9. The method according to claim 1, characterized in that, The process of optimizing the quality of the initially generated image corresponding to the digital life form and determining the quality detection results corresponding to multiple attributes includes at least two of the following steps: The similarity between the initial generated image facial feature vector and the facial feature vector in the reference image of the person's face is calculated. When the facial similarity between the two is less than the first preset similarity, the attention weight corresponding to the face in the comprehensive constraint condition is iteratively updated until the facial similarity between the two is not less than the first preset similarity. Texture defects and blurring are detected in the clothing area of the initial generated image. If the detection results do not meet the quality requirements, the clothing area is locally regenerated. The edge regions in the initially generated image are smoothed, and the color and brightness of the edge regions are adjusted using the lighting information in the scene reference image. The edge regions are the areas where the foreground and background of the person meet. The similarity between the skeletal key points in the initially generated image and the skeletal key points in the pose reference image is calculated. When the pose similarity between the two is less than the second preset similarity, the attention weights corresponding to the poses in the comprehensive constraint conditions are iteratively updated until the pose similarity between the two is not less than the second preset similarity. The similarity between the style feature vector in the initially generated image and the style feature vector corresponding to the style reference image is calculated. When the style similarity between the two is less than the third preset similarity, the attention weight corresponding to the style in the comprehensive constraint condition is iteratively updated until the style similarity between the two is not less than the third preset similarity.
10. A multi-image reference digital life form generation device, characterized in that, include: The reference image acquisition module is used to acquire multiple target reference images corresponding to different attributes. The multiple target reference images include at least two of the following: a person's appearance reference image, a clothing reference image, a scene reference image, a pose reference image, and a style reference image. The feature extraction module is used to extract features from the multiple target reference images respectively and determine multiple target feature vectors. The multiple target feature vectors include at least two of the following: facial feature vector, clothing feature vector, scene feature vector, posture feature vector, and style feature vector. A constraint determination module is used to determine comprehensive constraints. The comprehensive constraints include at least a comprehensive condition vector and condition injection parameters. The comprehensive condition vector includes attribute sub-vectors corresponding to at least two attributes among appearance, clothing, scene, posture, and style. The condition injection parameters include weights of multiple attributes at different resolution layers of a pre-trained target generation model. The weights are used to indicate the degree of influence of the attributes on different resolution layers. The image generation module is used to input the comprehensive constraints into the target generation model to generate images and determine the initial generated image corresponding to the digital life form. The generation features of different attributes in the initial generated image are constrained by the comprehensive constraints. The quality optimization module is used to optimize the quality of the initial generated image corresponding to the digital life form and determine the quality detection results corresponding to multiple attributes. When the quality detection result corresponding to any attribute is a failure, the comprehensive constraint conditions are iteratively updated based on the quality detection result and the image is regenerated until the quality detection results corresponding to all attributes are all passed, and then the target generated image corresponding to the digital life form is determined.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the multi-graph reference digital life generation method according to any one of claims 1 to 9.
12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the multi-graph reference digital life generation method according to any one of claims 1 to 9.