Image generation method based on single-image geometric texture joint control and terminal
Through the joint control method of single-graph geometric texture, the diffusion transformer model and low-rank adaptation training strategy are used to solve the problems of high training costs and multi-graph data sets in the existing technology, and the high-quality identity-maintaining image generation driven by single-graph is realized, improving the structural rationality and texture details of the generated images.
Patent Information
- Application Number
- CN202510740949.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The prior art has problems such as high training costs, time overhead, the need for multiple image data sets, the generated portrait body is incoordinated or lacks body components, and the ignorance of the shape and texture priors of the original input face in the generation of identity-keeping images, especially in methods without fine-tuning.
The method of joint control of single-graphic geometric texture is adopted, and by receiving the original image and text prompt information, a face sub-picture, a key point diagram and a frontal face image are generated, a face identity hidden vector, a frontal face control hidden vector and a key point control hidden vector are extracted, and a layered selective feature injection is combined with the diffusion transformer model, and a low-rank adaptation training strategy is adopted to generate an identity-keeping image.
The high-quality identity-keeping image generation driven by single-picture is realized, which reduces the computational cost, improves the structural rationality, identity consistency and texture details of the generated images, and breaks through the limitations of the traditional method's multi-picture or additional control model.
Smart Images

Figure CN120259479A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of model-based image generation, and particularly relates to an image generation method and a terminal for jointly controlling single-image geometry and texture. Background Art
[0002] Identity-preserving image generation technology has received extensive attention in the field of computer vision, especially identity-preserving methods based on diffusion models. It is mainly divided into two categories: fine-tuning-based methods and non-fine-tuning-based methods.
[0003] (1) Fine-tuning-based methods For fine-tuning-based methods, there are mainly three approaches. One is a scheme similar to Textual Inversion, which requires optimizing a new word embedding for each identity (input portrait) to enable the generation model to learn specific portrait features. This scheme needs to train a word embedding for each portrait. The second mainstream scheme is a scheme like DreamBooth: fine-tuning the entire generator to improve identity fidelity. The third scheme is to perform LoRA fine-tuning through pictures of the same person from different angles and in different scenarios, save a specific LoRA for each person, and combine the LoRA during the inference process to generate pictures with the characteristics of the input person.
[0004] (2) Non-fine-tuning-based methods To reduce the resource requirements of fine-tuning-based methods, many non-fine-tuning-based methods have emerged, such as the IP-Adapter series: using a CLIP or face recognition model to extract facial features and embedding the features into the generation network through cross-attention, so that the generation model can generate images with specific identities. Later, to further improve identity consistency, PuLID introduced semantic alignment and layout alignment losses based on the principle of progressive distillation to further enhance identity consistency. In addition, InstantID introduced an additional key point control network to control the face position, enabling the generated face to be generated with reference to the input face position, improving the controllability of the generated face. In addition to the one-step generation scheme using a single picture, there are also multi-step generation schemes to improve the similarity of the generated face, such as PhotoMaker, which embeds multiple facial images into the Stable Diffusion model, enabling the Stable Diffusion model to refer to the identity information of multiple pictures when generating images. Figure 1 However, (1) existing fine-tuning-based methods have high training costs and time overheads, and require multiple pictures of the same person from different angles and in different scenarios, resulting in limitations in practical scenarios.
[0005]
[0006] (2) Most of the existing methods without fine-tuning are based on the SD-XL model, and the generated human figures have uncoordinated bodies and fingers, or even lack body components.
[0007] (3) Existing human figure generation methods without fine-tuning and with a control network need to introduce an additional control network, resulting in a significant increase in the training and inference time.
[0008] (4) Existing human figure generation methods without fine-tuning and with a control network ignore the shape and texture priors of the original input face, resulting in insufficient similarity. Summary of the Invention
[0009] The technical problem to be solved by the present invention is to provide an image generation method and a terminal for jointly controlling single-image geometry and texture, so as to realize controllable generation of identity-preserving images driven by a single image and with multiple poses and styles.
[0010] To solve the above technical problems, the technical solution adopted by the present invention is as follows: An image generation method for jointly controlling single-image geometry and texture, comprising the steps of: S1. Receive the original image and text prompt information input by the user, and generate a face sub-image, a key point image, and a frontalized face image based on the original image; S2. Extract and process the features of the face sub-image to generate a face identity latent vector; generate a latent space control vector based on the key point image and the frontalized face image; generate a text embedding feature based on the text prompt information; S3. Inject the face identity latent vector, the text embedding feature, and the latent space control vector after splicing with a noise vector into a diffusion transformer model in a hierarchical selective feature injection manner, and generate an identity-preserving image in combination with a decoder; The diffusion transformer model is trained using a low-rank adaptation training strategy.
[0011] To solve the above technical problems, another technical solution adopted by the present invention is as follows: An image generation terminal for jointly controlling single-image geometry and texture, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps in any one of the above-mentioned image generation methods for jointly controlling single-image geometry and texture are implemented.
[0012] The beneficial effects of the present invention are as follows: A single-image geometric texture joint control image generation method and terminal of the present invention achieve precise control of "image generation from an image" by extracting face identity latent vectors (identity information), frontal face control latent vectors (texture information and shape information), and key point control latent vectors (geometric features), and combining text prompts. The hierarchical selective injection mechanism ensures the decoupling of identity features (injected into even layers) and style / pose features (injected into all layers), while the low-rank adaptation training strategy significantly reduces the computational cost, making it possible to generate high-quality single-image-driven generation and breaking through the limitations of traditional methods that require multiple images or additional control models. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a schematic diagram of a brief process example of a single-image geometric texture joint control image generation method according to an embodiment of the present invention; Figure 2 It is a schematic diagram of an overall view of a scheme of a single-image geometric texture joint control image generation method according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the cross-joint attention principle of a single-image geometric texture joint control image generation method according to an embodiment of the present invention; Figure 4 It is a schematic diagram of the structure of a conventional MM-DiT Block of a single-image geometric texture joint control image generation method according to an embodiment of the present invention; Figure 5 It is a schematic diagram of the structure of an MM-DiT Block with CAJ of a single-image geometric texture joint control image generation method according to an embodiment of the present invention; Figure 6 It is a schematic diagram of the structure of a Single-DiT module of a single-image geometric texture joint control image generation method according to an embodiment of the present invention; Figure 7 It is a schematic diagram of the structure of a single-image geometric texture joint control image generation terminal according to an embodiment of the present invention; Reference Numeral Explanation: 1. A single-image geometric texture joint control image generation terminal; 2. A processor; 3. A memory. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] To describe in detail the technical content, achieved objectives, and effects of the present invention, the following is described in conjunction with the embodiments and accompanied by the drawings.
[0015] Please refer to Figure 1 , A single-image geometric texture joint control image generation method includes the steps of: S1. Receive the original image and text prompt information input by the user, and generate a face sub-image, a key point image, and a frontal face image based on the original image; S2. Extract and process the features of the face sub - graph to generate a face identity latent vector; generate a latent space control vector based on the key - point graph and the frontalized face image; generate a text embedding feature based on the text prompt information; S3. Inject the face identity latent vector, the text embedding feature, and the latent space control vector after concatenating the noise vector into the diffusion transformer model in a hierarchical selective feature injection manner, and generate an identity - preserving image in combination with the decoder; The diffusion transformer model is trained using a low - rank adaptation training strategy.
[0016] As can be seen from the above description, the beneficial effects of the present invention are as follows: A single - image geometric texture joint control image generation method and terminal of the present invention realize precise control of "image - to - image" by extracting a face identity latent vector (identity information), a frontalized face control latent vector (texture information and shape information), and a key - point control latent vector (geometric features), and combining text prompts. The hierarchical selective injection mechanism ensures the decoupling of identity features (injected into even layers) and style / pose features (injected into all layers), while the low - rank adaptation training strategy significantly reduces the computational cost, making it possible to generate high - quality single - image - driven generation, breaking through the limitations of traditional methods that require multiple images or additional control models.
[0017] Further, step S1 includes the following steps: S11. Receive the original image and text prompt information input by the user; S12. Obtain the face region bounding box of the original image through a face recognition algorithm, and crop out the face sub - graph; S13. Perform face key - point detection on the face sub - graph, and generate a key - point graph according to the detection result; S14. Remove the background of the face sub - graph, and perform pose correction processing. Use the spatial transformation algorithm to align the face orientation and optimize the pose parameters to generate a frontalized face image.
[0018] As can be seen from the above description, the original image processing flow is refined. Through face detection (S12), key - point extraction (S13), and pose correction (S14), a complete geometric prior system is constructed. Precise face cropping avoids background interference, the key - point graph provides structural constraints for subsequent geometric control, and the frontalization process standardizes the pose reference, laying a foundation for the stable extraction of face identity features, and improving the structural rationality and identity consistency of the generated image.
[0019] Further, step S14 includes the following steps: S141. Segment the face sub - graph through the BiSeNet segmentation model and remove the background; S142. Calculate the pitch angle information between the face sub - graph and the standard frontalized face using the implicit key - point method in Live Portrait; S143. Input the pitch angle information and the face sub - graph into the Stitching module in Live Portrait to obtain an intermediate - state face at the target position; S144. Input the intermediate - state face into the Warpping module and decoder in Live Portrait to generate a frontalized face image.
[0020] As can be seen from the above description, the frontalization process is further optimized. Background noise is removed through BiSeNet segmentation (S141), and pose alignment is achieved using the implicit key - points (S142) and Stitching / Warping modules (S143 - S144) of Live Portrait. Compared with the traditional 3DMM method, this scheme does not require explicit 3D modeling, and can efficiently generate frontalized results through 2D image transformation, correcting the pose while retaining texture details, providing a standardized input for subsequent feature extraction.
[0021] Furthermore, the generation of the face identity latent vector includes the steps of: Input the face sub - graph into a face feature extractor, and extract face features through the ElasticFace algorithm; Process the extracted face features through an MLP layer and a LayerNorm layer to generate the face identity latent vector.
[0022] As can be seen from the above description, the ElasticFace algorithm is used to extract face features, and its dynamic angular boundary mechanism effectively solves the problems of over - fitting of simple samples and under - fitting of difficult samples in traditional metric learning, improving the robustness of identity features. The multi - layer perceptron MLP + layer normalization LayerNorm processing further enhances the feature expression ability, enabling the generated image to maintain identity consistency during pose and expression changes.
[0023] Furthermore, the generation of the latent space control vector includes the steps of: Encode the key - point map and the frontalized face image through a VAE encoding model, and splice the encoded vectors to obtain the latent space control vector.
[0024] As described above, by fusing the key point map (geometry) and the frontalized face (texture and shape) through VAE encoding, the generated latent space control vector realizes the decoupled representation of geometry and texture. The splicing operation enables the two features to cooperate in the latent space, retaining both the face structure constraints (such as the positions of facial features) and transmitting details such as face shape and skin texture, providing a comprehensive and controllable conditional signal for subsequent diffusion generation and enhancing the realism and diversity of the generated images.
[0025] Further, the diffusion transformer model consists of 19 MM-DiT blocks and 38 Single-DiT blocks, and the cross-attention mechanism is injected into the MM-DiT blocks at even positions. Step S3 includes the steps: The latent space control vector after splicing the noise vector is input into the diffusion transformer model. The face identity latent vector is injected into the MM-DiT blocks at even positions in the diffusion transformer model, and the text embedding features are injected into all the MM-DiT blocks in the diffusion transformer model to obtain the model output. The model output is input into the VAE decoder to obtain the final identity-preserving image containing the features of the person in the original image.
[0026] As described above, a specific network architecture (19 MM-DiT + 38 Single-DiT) and a hierarchical injection strategy are designed. The cross-attention mechanism (injecting the identity latent vector) of the MM-DiT blocks at even positions ensures strong constraints on identity features, while the injection of text embeddings in all MM-DiT blocks and the injection of key point latent vectors achieve flexible control of style / pose. This structure balances identity preservation and generation diversity.
[0027] Further, the loss function of the diffusion transformer model adopts MSE Loss, which is expressed as follows: ; ; Among them, x 1 is the latent vector after the original image is encoded by VAE, x 0 represents noise, t represents time, c n represents the conditions, including text prompt information, key point map, and frontalized face image.
[0028] As described above, the MSE Loss is used as the loss function of the diffusion model to directly optimize the noise prediction, enabling the model to learn the gradient flow of the data distribution. Combined with conditional inputs (text, key points, frontalized images), this loss function guides the model to precisely follow geometric and texture constraints during the generation process, improving image quality and controllability.
[0029] Furthermore, the training of the diffusion transformer model using the low-rank adaptation training strategy includes: Fix the main body of the diffusion transformer model and only train the low-rank decomposition matrices for the cross-attention mechanism and attention weights, and combine the flow matching algorithm to accelerate convergence.
[0030] As described above, the low-rank adaptation strategy freezes the parameters of the DiT main body and only trains the low-rank decomposition matrices of the cross-attention and attention weights, reducing the number of parameters and video memory occupancy. Combined with the flow matching algorithm, it improves the training speed, supports deployment on consumer-grade GPUs, and at the same time avoids catastrophic forgetting caused by full-parameter fine-tuning, retaining the visual understanding ability of the pre-trained model.
[0031] Furthermore, the noise vector is randomly sampled from a Gaussian distribution and has the same shape and size as the latent space control vector.
[0032] As described above, the noise vector is randomly sampled from a Gaussian distribution and, after being concatenated with the latent space control vector, is used as the input of the diffusion model, introducing randomness into the generation process. The design of the same shape and size ensures that the noise matches the dimension of the conditional signal, enabling the model to generate diverse texture details (such as expressions, lighting changes) while maintaining identity and geometric constraints, enhancing the richness and authenticity of the generation results.
[0033] An image generation method and terminal for joint control of single-image geometry and texture according to the present invention are applicable to identity-preserving image generation based on single-image input.
[0034] Please refer to Figures 1 to 6 , and the first embodiment of the present invention is: An image generation method for joint control of single-image geometry and texture, comprising the steps of: S1. Receive the original image and text prompt information input by the user, and generate a face sub-image, a key point image, and a frontalized face image based on the original image; Step S1 includes the steps of: S11. Receive the original image and text prompt information input by the user; S12. Obtain the face region bounding box of the original image through a face recognition algorithm, and crop out the face sub-image.
[0035] In this embodiment, the input original image y is input into the face detection module to obtain the bounding box of the face region; The face sub-image f is cropped according to the bounding box and the original image y c , and is sent to the subsequent module.
[0036] S13. Perform face key point detection on the face sub-image, and generate a key point map according to the detection result.
[0037] In this embodiment, face key point detection is further performed on the result of face detection; for the 5 points (left and right eyes, nose, left and right mouth corners) in the key point detection result, with the nose as the center, the connection lines between the nose and the other four points are respectively drawn to obtain the key point map k c .
[0038] S14. Remove the background from the face sub-image, and perform pose correction processing. Use the spatial transformation algorithm to align the face orientation and optimize the pose parameters to generate a frontalized face image.
[0039] Step S14 includes the steps: S141. Segment the face in the face sub-image through the BiSeNet segmentation model and remove the background; S142. Calculate the pitch angle information between the face sub-image and the standard frontalized face using the implicit key point method in Live Portrait; S143. Input the pitch angle information and the face sub-image into the Stitching module in Live Portrait to obtain the intermediate face at the target position; S144. Input the intermediate face into the warpping module and decoder in Live Portrait to generate a frontalized face image.
[0040] In this embodiment, the face sub-image y c is segmented through BiSeNet, the background is removed, and replaced with a white background; Calculate the pitch, yaw, and roll of the face sub-image y c and the standard frontalized face. In particular, the implicit key point method in Live Portrait is adopted here; Further input the Pitch, Yaw, Roll information and the segmented result into the stitching module in Live Portrait to obtain the intermediate face at the target position; Input the intermediate face into the Warpping module and decoder, and finally obtain the frontalized face image f w ; The formula for the entire frontalization process is as follows: ; ; ; ; ; where, represents the face sub - image after cropping, represents the BiSeNet segmentation model, M represents the segmentation mask, ⊙ represents the Hadamard product, and 1 represents a vector of all 1s with the same shape as corresponding to white, represents the segmented image. represents the pitch angle estimation network, represents the estimated pitch angle parameter, Stitching, Warping, respectively represent the Stitching module, Warping module and decoder processing of Live Portrait.
[0041] S2. Extract and process the features of the face sub - image to generate a face identity latent vector; generate a latent space control vector based on the key point map and the frontalized face image; generate a text embedding feature based on the text prompt information; The generation of the face identity latent vector includes the steps of: Input the face sub - image into a face feature extractor, and extract face features through the ElasticFace algorithm; In this embodiment, the face sub - image y c is input into the face feature extractor to extract face features. Specifically, ElasticFace is used here to extract face features e c ; its formula can be expressed as follows: .
[0042] Process the extracted face features through an MLP layer and a LayerNorm layer to generate the face identity latent vector.
[0043] In this embodiment, the extracted face features e c pass through the MLP layer and the LayerNorm layer to obtain the face identity latent vector , and its calculation process can be described by the formula as follows: .
[0044] The generation of the latent space control vector includes the steps: Encode the key point map and the frontalized face image through the VAE encoding model, and splice the encoded vectors to obtain the latent space control vector.
[0045] In this embodiment, the frontalized face image f w and the key point map k c are respectively encoded through the VAE encoding model, and the latent space control vector c N obtained by splicing the encoded vectors can be described by the following formula: ; ; .
[0046] S3. Inject the face identity latent vector, the text embedding feature, and the latent space control vector after splicing the noise vector into the diffusion Transformer model in a hierarchical selective feature injection manner, and combine the decoder to generate an identity-preserving image; The diffusion Transformer model is trained using a low-rank adaptation training strategy.
[0047] The diffusion Transformer model consists of 19 MM-DiT blocks and 38 Single-DiT blocks, and the cross-attention mechanism is injected into the MM-DiT blocks at even positions; Step S3 includes the steps: Input the latent space control vector after splicing the noise vector into the diffusion Transformer model, inject the face identity latent vector into the MM-DiT blocks at even positions in the diffusion Transformer model, and inject the text embedding feature into all MM-DiT blocks in the diffusion Transformer model to obtain the model output; Input the model output into the VAE decoder to obtain the final identity-preserving image containing the characteristics of the person in the original image.
[0048] In this embodiment, referring to Figure 2 , splice the latent space control vector c N with the noise vector to obtain the final vector z N input into the DiT model, and its process can be described by the following formula, and the noise vector is randomly sampled from the Gaussian distribution, and its shape and size are the same as those of the control vector: .
[0049] At the same time, in this embodiment, the face identity latent vector after passing through the embedding projection layer , which is injected into the DIT module through Cross Joint Attention (CJA), and the composition of CJA is as Figure 3 shown. Three input vectors, the latent space control vector z N , the text latent vector e t and the face identity latent vector . Among them, the text latent vector is the vector obtained after passing the input text prompt information through the text encoder of Flux. The face identity latent vector is obtained by the face feature extractor extracting the face feature e c from the face subgraph, and passing through the MLP layer and the LayerNorm layer. After passing through CJA, the image latent vector embedded with identity information is obtained.
[0050] Among them, reference can be made to Figure 4 and Figure 5 . CJA injects the identity latent vector only when the multi-modal DiT block (MM-DiT Block) is even (i, i + 2,... i + 2N), but the text latent vector is injected into all MM-DiT modules.
[0051] The output z0 of the DiT model is sent into the VAE decoder to obtain the final picture containing the input person's features. The DiT model is composed of 19 MM-DiT blocks and 38 Single-DiT blocks. The composition of the ordinary MM-DiT can be described by Figure 4 , the MM-DiT module with the CAJ attention mechanism is described by Figure 5 , and the Single-DiT module is described by Figure 6 .
[0052] Among them, the VAE uses the VAE model officially open-sourced by Flux.
[0053] Embodiment 2 of the present invention is as follows: An image generation method for single-image geometric texture joint control, which is different from Embodiment 1 in that the training of the diffusion transformer model is described in this embodiment.
[0054] The diffusion transformer model is trained using a low-rank adaptation training strategy, including: Fixing the main body of the diffusion transformer model, and only training the cross joint attention mechanism and the low-rank decomposition matrix of the attention weight, and combining the flow matching algorithm to accelerate convergence.
[0055] In this embodiment, during the training process of the entire model, the DiT model is fixed, and only the parameters of the added LoRA layer, as well as the parameters in the embedding projection layer and CAJ, are trained. The parameters to be trained in the LoRA layer are as follows:
[0056] The loss function of the diffusion transformer model uses MSE Loss, which is expressed as follows: ; .
[0057] Among them, x 1 is the latent vector after the original image is encoded by VAE, x 0 represents noise, t represents time, c n represents the condition, including text prompt information, key point map, and frontalized face image.
[0058] Please refer to Figure 7 , and the third embodiment of the present invention is: An image generation terminal 1 for joint control of single - image geometry and texture, including a processor 2, a memory 3, and a computer program stored in the memory 3 and executable on the processor 2. When the processor 2 executes the computer program, it implements the steps in the method for generating an image with joint control of single - image geometry and texture described in the above - mentioned embodiment 1 or 2.
[0059] In summary, the method and terminal for generating an image with joint control of single - image geometry and texture according to the present invention extract the face identity latent vector (identity information), the frontalized face control latent vector (texture information and shape information), and the key point control latent vector (geometric features), and combine text prompts to achieve precise "image - to - image" control. The hierarchical selective injection mechanism ensures the decoupling of identity features (injected into even layers) and style / pose features (injected into all layers), while the low - rank adaptation training strategy significantly reduces the computational cost, making it possible to generate high - quality single - image - driven images, breaking through the limitations of traditional methods that require multiple images or additional control models.
[0060] The present invention can bring the following advantages: Advantage 1. It can achieve identity - preserving image generation with only a single reference image, without the need to collect data of different scenes and angles of the same person in multiple images, without additional training time consumption, and without the need to store LoRA models or DreamBooth models for each person redundantly. The single - inference time is shortened to the second level.
[0061] Advantage 2. The frontalized face and key points are jointly used as control conditions and input into the Flux DiT model, making the model generate precisely controllably and be able to refer to the texture of the frontalized face, improving the consistency of the generated image and the input image in terms of high-level semantics and low-level texture.
[0062] Advantage 3. During the training process of the Flux control network, no additional control network is introduced. Instead, through the LoRA injection method, the control conditions are directly input into DiT, significantly reducing the training cost and improving the training efficiency.
[0063] The solution of the present invention can solve the drawbacks existing in the prior art in the background art of this application. The reason is as follows: 1. The generalized hierarchical identity feature extraction network learns the common representation rules across characters through pre-training, directly breaking through the limitation of traditional methods that rely on multi-angle data. After pre-training, the network only needs a single reference image to analyze the essence of identity features, without the need to fine-tune the model parameters for specific characters, fundamentally eliminating the need for multi-data collection and model storage.
[0064] 2. The joint control condition injection strategy jointly encodes the frontalized face and key points as the control signal of Flux DiT, enabling the model to synchronously constrain high-level semantics (such as the positions of facial features) and low-level texture (such as skin texture) during the generation process. This dual guidance mechanism directly acts on the generation path, avoiding the semantic-texture inconsistency splitting problem caused by single-condition control in traditional methods.
[0065] 3. The LoRA lightweight control architecture directly embeds the control conditions into the attention layer of DiT instead of an external control network, and realizes the model control ability through efficient parameter adaptation. Experiments prove that this design reduces the number of training parameters by 99%, and no additional network modules need to be loaded during the inference stage, forming a closed loop with the characteristics of low storage and high efficiency.
[0066] The above are only embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. All equivalent transformations made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in related technical fields, are equally included in the patent protection scope of the present invention.
Claims
1. An image generation method with joint control of single - map geometric texture, characterized in that, Including the steps: S1. Receive the original image and text prompt information input by the user, and generate a face sub-image, a key point map, and a frontalized face image based on the original image; S2. Extract and process the features of the face sub-image to generate a face identity latent vector; generate a latent space control vector based on the key point map and the frontalized face image; generate a text embedding feature based on the text prompt information; S3. Inject the face identity latent vector, the text embedding feature, and the latent space control vector after splicing with a noise vector into a diffusion transformer model in a hierarchical selective feature injection manner, and generate an identity-preserving image in combination with a decoder; The diffusion transformer model is trained using a low-rank adaptation training strategy.
2. The image generation method for jointly controlling single - map geometric texture according to claim 1, wherein, Step S1 includes the steps: S11. Receive the original image and text prompt information input by the user; S12. Obtain the face region bounding box of the original image through a face recognition algorithm, and crop out the face sub-image; S13. Perform face key point detection on the face sub-image, and generate a key point map according to the detection result; S14. Remove the background of the face sub-image, and perform pose correction processing. Use a spatial transformation algorithm to align the face orientation and optimize the pose parameters to generate a frontalized face image.
3. The image generation method for single-graph geometric texture joint control according to claim 2, wherein Step S14 includes the steps: S141. Segment the face of the face sub-image through a BiSeNet segmentation model and remove the background; S142. Calculate the pitch angle information between the face sub-image and the standard frontalized face using the implicit key point method in Live Portrait; S143. Input the pitch angle information and the face sub-image into the Stitching module in Live Portrait to obtain an intermediate face at the target position; S144. Input the intermediate face into the Warpping module and decoder in Live Portrait to generate a frontalized face image.
4. A method for generating an image by jointly controlling a single image geometric texture according to claim 1, wherein, The generation of the face identity latent vector includes the steps: Input the face sub-image into a face feature extractor, and extract face features through the ElasticFace algorithm; Process the extracted face features through an MLP layer and a LayerNorm layer to generate the face identity latent vector.
5. A method for generating an image by jointly controlling a single-image geometric texture according to claim 1, wherein, The generation of the latent space control vector includes the steps: Encode the key point map and the frontalized face image through a VAE encoding model, and splice the encoded vectors to obtain the latent space control vector.
6. The image generation method with combined control of single - graph geometric texture according to claim 1, wherein, The diffusion transformer model consists of 19 MM-DiT blocks and 38 Single-DiT blocks, and the cross-attention mechanism is injected into the MM-DiT blocks at even positions; Step S3 includes the steps: Input the latent space control vector after splicing with a noise vector into the diffusion transformer model, inject the face identity latent vector into the MM-DiT blocks at even positions in the diffusion transformer model, and inject the text embedding feature into all MM-DiT blocks in the diffusion transformer model to obtain the model output; Input the output of the model into the VAE decoder to obtain the final identity-preserving image containing the human features in the original image.
7. A method for generating an image by jointly controlling a single map geometric texture according to claim 6, characterized in that, The loss function of the diffusion transformer model adopts MSE Loss, which is expressed as follows: ; ; Among them, x 1 is the latent vector after the original image is encoded by VAE, x 0 represents noise, t represents time, c n represents the conditions, including text prompt information, key point maps, and frontalized face images.
8. A method for generating an image by jointly controlling a single map geometric texture according to claim 6, characterized in that, The training of the diffusion transformer model using the low-rank adaptation training strategy includes: Fix the main body of the diffusion transformer model, and only train the cross-attention mechanism and the low-rank decomposition matrix of the attention weights, and combine the flow matching algorithm to accelerate convergence.
9. A method for generating an image by jointly controlling a single map geometric texture according to claim 6, wherein, The noise vector is randomly sampled from a Gaussian distribution and has the same shape and size as the latent space control vector.
10. An image generation terminal with combined control of single - image geometric texture, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps in the image generation method for single-image geometric texture joint control according to any one of claims 1-9 above.
Citation Information
Patent Citations
Human body image generation method and system based on freehand sketch
CN112862920A
Image processing method based on few-sample learning and related equipment
CN116310008A
High-precision modeling method for three-dimensional face of digital teacher
CN116958420A
Face image synthesis method and system
JP2011060289A
Image processing method and apparatus, device, and medium
WO2022257766A1
Cited By
Image annotation control method
CN121330687A
Head portrait splicing method and device based on extension model, and storage medium
CN121353074A
Avatar splicing method and device based on extended model and storage medium
CN121353074B
Personalized LoRA model construction method and system and storage medium
CN121458529A
A personalized LoRA model construction method and system, and a storage medium
CN121458529B