Digital person photo reloading method and device

By combining the VL visual model and segmentation technology, the problems of operational complexity and low detail restoration in existing dressing technologies are solved, and a high-quality, facial-consistent dressing effect is achieved. It is suitable for scenarios such as digital human dressing, virtual fitting, and social media content creation.

CN120706557APending Publication Date: 2025-09-26XIAMEN CHANJING TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510808864.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing dressing-up technology is complex to operate, has low clothing detail restoration, cannot maintain facial consistency, and has poor adaptability to complex scenes. It is difficult to meet users' needs for high-quality, high-freedom, and high-restoration dressing-up.

Method used

The VL visual model is used to identify clothing types, and the SAM segmentation model and G-DINO semantic segmentation technology are combined to segment the model image. The FLux Fill model and CLIP model are used to redraw clothing features. Semantic constraints and pose estimation are introduced to generate high-quality clothing change images.

Benefits of technology

It achieves efficient, flexible and high-quality dressing effects, simplifies the operation process, can automatically process various postures and backgrounds, and generate dressing images with realistic clothing details and consistent faces, which is suitable for a variety of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706557A_ABST
    Figure CN120706557A_ABST
Patent Text Reader

Abstract

The invention provides a digital person photo changing method and device, and relates to the technical field of photo changing, and the core of the method lies in that 1, a clothing feature extraction technology is adopted, a visual encoder model is utilized to carry out high-precision feature extraction on a clothing image, and a semantic segmentation technology is combined to realize automatic and precise shielding of a clothing area; according to the dynamic redrawing technology, natural adaptation and detail consistency control of the garment contour are achieved through attitude estimation and a low-rank adapter mechanism; and a context-aware semantic segmentation technology is provided, so that the clothing area can be accurately positioned under a complex background, and the mask accuracy is improved. In addition, a visual language model is introduced, through accurate garment prompt word generation, the mask effect of the segmentation model is further optimized, and the reduction degree of garment details is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of changing the costume of a person's photo, and in particular to a method and device for changing the costume of a digital human's photo. Background Art

[0002] In the fields of digital content creation and image processing, with the rapid development of artificial intelligence (AI) technology, the demand for personalized customization of digital avatars and human images is growing. In particular, in scenarios such as social media, e-commerce marketing, and virtual avatar creation, users are increasingly demanding the ability to quickly and efficiently change the costumes of digital avatars or real-life photos. However, existing costume-changing technologies face numerous challenges in practical application, struggling to meet user demands for high-quality, high-definition customization, and high fidelity.

[0003] Traditional dressing technologies rely primarily on manual input and simple image processing methods. For example, users need to upload separate images of tops and bottoms, and often need to provide full-body model images, which greatly limits user usage scenarios and flexibility. Furthermore, existing dressing technologies have significant shortcomings in generating clothing details. Relying on simple image feature extraction or cue word inference, these methods struggle to accurately capture details such as clothing wrinkles, materials, and textures, resulting in less than refined clothing and potentially unnatural texture misalignment and edge issues. The limitations of existing technologies are particularly pronounced when dealing with complex model poses, such as the extremely poor edge generation of occluded areas like hands, hair, and accessories.

[0004] Furthermore, when a user is very satisfied with the facial features of a digital person or photo of a person, but dissatisfied with the clothing or other features, existing technology cannot fine-tune the face while maintaining consistency. When users attempt to adjust the clothing, they often need to regenerate the entire image, causing changes in facial features and failing to maintain the desired effect. This limitation is particularly prominent in female content and live-streaming scenarios, as these scenarios require a high level of sophistication in clothing, and existing clothing-changing technology struggles to meet user expectations.

[0005] In practice, existing costume-changing technology faces other challenges. For example, images on some e-commerce platforms lack clarity, background clarity, and some online images carry watermarks. These issues further limit the application and effectiveness of existing costume-changing technology. Furthermore, existing technology is less than ideal when dealing with women with long hair or those with unusual sitting postures (such as shrugging shoulders). These issues prevent existing costume-changing technology from meeting users' demands for high-quality, flexible, and accurate costume-changing.

[0006] Therefore, the market urgently needs a new technical solution that can solve the problems existing in existing dressing-up technologies, such as complex operation, low detail restoration, inability to maintain facial consistency, and poor adaptability to complex scenes, so as to provide users with a simple, efficient, and high-quality dressing-up experience to meet the diverse needs of users in different scenarios.

[0007] In view of this, this application is filed. Summary of the Invention

[0008] The present invention provides a method and device for changing the costume of a digital human portrait, which can at least partially improve the above-mentioned problem.

[0009] To achieve the above object, the present invention adopts the following technical solutions: A method for changing the appearance of a digital human portrait, comprising: Acquire a clothing image, identify the clothing image using a VL visual model, and generate a clothing type recognition result; Based on the clothing type recognition result, the clothing image is extracted and processed to generate a cross-modal feature vector and a high-dimensional feature vector, and semantic constraints corresponding to the clothing image are introduced to obtain clothing features; Obtain a model image, perform segmentation processing on the model image by combining the SAM segmentation model and the G-DINO semantic segmentation technology, and associate the model image with the human body posture during the segmentation processing based on the CLIP model to obtain a redrawing area of ​​the model image; The clothing features and the model image redrawing area are redrawn according to the FLux Fill model to generate the final dressing image.

[0010] The present invention also provides a device for changing the appearance of a digital human portrait, which comprises: A clothing type recognition unit is used to obtain a clothing image, identify the clothing image using a VL visual model, and generate a clothing type recognition result; a clothing feature extraction unit, configured to extract and process the clothing image based on the clothing type recognition result, generate a cross-modal feature vector and a high-dimensional feature vector, and introduce semantic constraints corresponding to the clothing image to obtain clothing features; A redrawing region segmentation unit is used to obtain a model image, perform segmentation processing on the model image by combining the SAM segmentation model and the G-DINO semantic segmentation technology, and associate the model image with the human body posture during the segmentation processing based on the CLIP model to obtain a redrawing region of the model image; The redrawing unit is used to redraw the clothing features and the model image redrawing area according to the FLux Fill model to generate the final dressing image.

[0011] In summary, one of the core innovations of the described digital human costume replacement method is dynamic redrawing technology. Through pose estimation and a low-rank adapter mechanism, this technology automatically adjusts clothing details based on human pose, overcoming the limitations of traditional fixed-rank parameters in detail restoration. Furthermore, context-aware semantic segmentation technology is introduced to accurately locate clothing regions in complex backgrounds, further improving masking accuracy. Furthermore, to further optimize the restoration of clothing details, this method incorporates a visual language model (VL model). This model accurately generates clothing cue words, helping the segmentation model more accurately identify clothing regions, thereby improving generation results. Without requiring complex user input, the method automatically processes input images of various poses and backgrounds to produce high-quality costume replacement results.

[0012] Another innovative aspect of this digital human costume-changing method lies in its simple and easy-to-use workflow. Users no longer need to distinguish between top and bottom images or provide full-body model images; half-body or full-body model images can be automatically identified and processed. Furthermore, it supports a variety of application scenarios, including digital human costume-changing, virtual fitting, and social media content creation. By simplifying the workflow and improving the quality of the results, this method provides users with a brand new costume-changing experience while also providing strong support for the digital transformation of related industries.

[0013] Without requiring complex user input, the system automatically recognizes and processes input images of various poses and backgrounds, generating high-quality costume changes. This system is suitable for a variety of application scenarios, including but not limited to digital human costume changes, virtual fitting, and social media content creation. By streamlining operational processes and improving quality, it provides users with a brand new costume change experience while also providing strong support for the digital transformation of related industries. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 1 is a flow chart of a method for changing the appearance of a digital human portrait provided by the first embodiment of the present invention; Figure 2 1 is a schematic diagram of a process for segmenting a model image redrawing region according to an embodiment of the present invention; Figure 3 1 is a schematic diagram showing the effect of the method for changing the costume of a digital human portrait provided by an embodiment of the present invention; Figure 4 2 is a module diagram of a digital human portrait dressing device provided by a second embodiment of the present invention. DETAILED DESCRIPTION

[0015] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0016] refer to Figure 1 As shown, the first embodiment of the present invention discloses a method for changing the costume of a digital human portrait, which can be executed by a digital human portrait changing device (hereinafter referred to as the changing device), and in particular, is executed by one or more processors in the changing device to implement the following method: S1, obtaining a clothing image, using a VL visual model to identify the clothing image, and generating a clothing type recognition result; Preferably, the VL visual model is a Qwen-VL multimodal visual language model, and the clothing type recognition result includes tops and bottoms or suits.

[0017] Specifically, in this embodiment, first, clothing images provided by the user are obtained. These images can be single clothing pictures, pictures of models wearing clothing, or any images containing clothing uploaded by the user. In order to ensure the accuracy and consistency of the dressing effect, an advanced multimodal visual language model, namely the Qwen-VL model, is used to identify and classify clothing images. By combining the visual encoder and the language model, the Qwen-VL model can accurately identify the type of clothing and generate detailed clothing type recognition results, including information such as tops and bottoms or suits. This high-precision recognition capability provides a reliable foundation for subsequent dressing processing. Users do not need to manually specify the clothing type, and it is automatically identified and processed, which greatly simplifies the operation process and improves the user experience.

[0018] S2, based on the clothing type recognition result, extracting and processing the clothing image to generate a cross-modal feature vector and a high-dimensional feature vector, and introducing semantic constraints corresponding to the clothing image to obtain clothing features; Specifically, step S2 further includes: extracting the clothing image using an image block embedding module of a visual encoder with a ViT architecture and a 12-layer Transformer encoder to obtain global semantic features; According to the global semantic features, the redux model is used to construct and generate cross-modal feature vectors, and the CLIP visual encoding based on clothing images and text descriptions are semantically aligned to generate high-dimensional feature vectors.

[0019] Obtaining the prompt words of the original generated image, and performing tuning processing on the prompt words based on the LLM large language model, wherein the tuning processing includes using the joy caption two model to reversely infer the prompt words of the image; Clothing features are obtained based on cross-modal feature vectors, high-dimensional feature vectors and optimized prompt words.

[0020] In this embodiment, after obtaining the clothing type recognition results, the key clothing feature extraction stage begins. The core of this stage is to perform in-depth processing on the clothing image, generate cross-modal feature vectors and high-dimensional feature vectors, and introduce semantic constraints corresponding to the clothing image to obtain accurate clothing features. Specifically, a visual encoder based on the ViT (Vision Transformer) architecture is used to extract clothing images through its image block embedding module and 12-layer Transformer encoder to obtain global semantic features. The ViT architecture visual encoder can efficiently capture detailed information such as clothing texture, folds, and material, generate high-precision global semantic features, and provide a solid foundation for subsequent feature processing.

[0021] Based on the generated global semantic features, a redux-based approach is used to construct a cross-modal feature vector. The CLIP visual encoding of the garment image is semantically aligned with the text description to generate a high-dimensional feature vector. This process ensures the accuracy and consistency of the garment features by introducing semantic constraints. Semantic alignment of the CLIP visual encoding with the text description enables a better understanding of the garment's visual features and semantic information, resulting in the generation of dressing effects that better meet user needs.

[0022] To further optimize the outfit-changing effect, we obtain the prompt words from the original generated image (for example, "On the left is a picture of the item, and on the right is a young Asian man holding it fully. Hands clear and intact, detailed fingers, high-quality hand anatomy, realistic hand pose, well-defined finger joints, and natural finger positions") and refine them using the Large Language Model (LLM). Specifically, we use the JoyCaption2 model to infer the prompt words from the image, generating more accurate descriptive text (for example, "On the left is a picture of a canned drink on a white background with a green exterior, with the speaker in a nodding pose, and on the right is a young Asian man holding the item intact. The hand is clear and intact, with detailed fingers, high-quality hand anatomy, realistic hand pose, well-defined finger joints, and natural finger positions"). Through LLM refinement, we can generate prompt words that better match the characteristics of the clothing, thereby improving the naturalness and accuracy of the outfit-changing effect. The tuned prompt words are combined with cross-modal feature vectors and high-dimensional feature vectors to finally obtain accurate clothing features, providing high-quality input for clothing change processing.

[0023] Based on this step, this method achieves efficient, flexible, and high-quality costume changes. The ViT architecture's visual encoder and Transformer encoder ensure high-precision extraction of clothing features. Semantic alignment of the CLIP visual encoding with text descriptions further improves feature accuracy and consistency. Cue word optimization within the Large Language Model (LLM) provides strong support for the naturalness and accuracy of costume changes.

[0024] See also Figure 2 ,S3, obtain the model image, combine the SAM segmentation model and the G-DINO semantic segmentation technology to segment the model image, and associate the model image with the human body posture during the segmentation process based on the CLIP model to obtain the model image redrawing area; Specifically, step S3 further includes: obtaining a model image, using the SAM segmentation model to calculate the semantic similarity between each mask of the model image and the text prompt through the CLIP model, and retaining its top-3 candidates; The top-3 candidates are modified based on the G-DINO semantic segmentation technology. The text encoder of the CLIP model generates a fine-grained semantic query vector, which is dynamically decoded and combined with the variable attention mechanism to locate the clothing parts in the model image. The formula is: ,in, is the query vector, is the reference point coordinate, is the input model image, is the number of sampling points, is the learnable offset, is the weight coefficient (learned through the attention mechanism); Performing pose estimation on the model image using the OpenPose algorithm to generate a pose estimation result, and performing context association on the model image based on the CLIP model to generate an association result; The association results and the clothing parts in the located model image are fused with posture guidance, and their edges are optimized to segment and generate the model image redrawing area.

[0025] In this example, a model image is first obtained. These images can be full-length or half-length images of a person, which are then used for garment redrawing. To accurately segment the garment regions within the model image, the SAM segmentation model is combined with G-DINO semantic segmentation technology. This process not only accurately identifies the garment regions but also dynamically adjusts them based on the person's posture, ensuring natural and accurate segmentation results.

[0026] Specifically, the model image is first segmented using the SAM segmentation model. The SAM model then uses the CLIP model to calculate the semantic similarity between each mask in the model image and the textual hint, retaining the top three candidates. This process leverages the powerful semantic understanding capabilities of the CLIP model to ensure a high degree of alignment between the segmentation mask and the clothing region. By retaining the top three candidates, the segmentation results can be further optimized, improving both accuracy and robustness.

[0027] Next, the top-3 candidates are modified based on the G-DINO semantic segmentation technology. The text encoder of the CLIP model generates a fine-grained semantic query vector, which is dynamically decoded and combined with the variable attention mechanism to locate the clothing parts in the model image. This process is done by formula The parameters are explained in Table 1.

[0028] Table 1

[0029] This combination of dynamic decoding and attention mechanism enables accurate localization of clothing parts and maintains high segmentation accuracy even in complex model poses and occlusions.

[0030] To further enhance the naturalness of the segmentation, the OpenPose algorithm is used to perform pose estimation on the model image, generating pose estimation results. The OpenPose algorithm accurately identifies the key points of the human body in the model image, providing a foundation for subsequent pose-guided fusion. Based on the CLIP model, the pose estimation results are contextualized to generate association results. This process, leveraging the CLIP model's contextual awareness, ensures overall semantic consistency between the pose estimation results and the model image.

[0031] Finally, a pose-guided fusion is performed on the association results and the located clothing parts in the model image, and their edges are optimized to segment and generate the model image redrawing region. This pose-guided fusion ensures a natural fit between the clothing region and the model's pose, while edge optimization further enhances the naturalness and aesthetics of the segmentation results. The resulting model image redrawing region provides precise input for subsequent clothing redrawing, ensuring high-quality and natural-looking costume changes.

[0032] Based on step S3, this method not only accurately segments the clothing area in the model image but also dynamically adjusts and optimizes it based on the human pose. The combination of the SAM segmentation model and G-DINO semantic segmentation technology ensures high segmentation accuracy and robustness. The pose estimation of the OpenPose algorithm and the contextual association of the CLIP model further enhance the naturalness and consistency of the segmentation results. In step S4, the clothing features and the redrawn areas of the model image are redrawn using the FLux Fill model to generate the final image of the changed outfit.

[0033] Specifically, step S4 further includes: performing clothing contour adaptation processing on the clothing features and the redrawn area of ​​the model image based on the Inpaint function of the FLux Fill model in combination with the posture estimation algorithm; The splicing features of the clothing plane map and the redrawn area constructed in the latent space by the In-Context LoRA technology are introduced. The formula of the splicing feature is: , is the potential feature of the clothing plan, To redraw the latent features of the region mask, is the clothing feature projection matrix, is the mask feature projection matrix, For splicing operation, is a d×d dimensional real matrix space (the vector space of projection matrices); Based on the splicing features, a low-rank adapter is used to perform dynamic low-rank adaptation on the image after clothing contour adaptation to achieve consistent control of image details and generate a clothing change image.

[0034] Dynamic low-rank adaptation processing includes dynamically adjusting the LoRA matrix rank of the image using the joint bending posture key points identified by the OpenPose algorithm; Among them, the rank adjustment coefficient is calculated according to the posture key points , and according to the formula , dynamically adjust the LoRA matrix rank of the image, where, is the adjusted rank, As the basic rank.

[0035] In this embodiment, the Inpaint function of the FLux Fill model is used to process clothing features and the redrawn area of ​​the model image. The core of this function is the ability to intelligently fill and adjust missing or modified parts based on existing image information. In a dress-up scene, this means that the outline of the clothing can be accurately adapted to the model's posture and body contours, so that it fits naturally with the model's body. For example, when the model bends their arm, the folds and shape of the clothing in that area can be automatically adjusted to ensure a natural transition of the clothing. This posture-based clothing contour adaptation process greatly enhances the realism and naturalness of the dress-up effect, making the generated image look as if the model is actually wearing the clothing.

[0036] To further improve the quality of the costume change effect, we introduced In-Context LoRA technology. This technology constructs the splicing features of the clothing plan image and the redrawn area in the latent space and fuses these features using a specific formula. Specifically, the parameters in the formula are explained in Table 2.

[0037] Table 2

[0038] This stitching operation allows the garment to maintain its original design and texture while perfectly matching the model's body shape and posture. This technology not only improves the consistency of detail in the costume change effect, but also provides higher-quality feature input for subsequent image generation.

[0039] Next, based on the splicing features, a low-rank adapter is used to perform dynamic low-rank adaptation on the adapted image. The core of this processing process is to dynamically adjust the LoRA matrix rank of the image based on the joint curvature posture key points identified by the OpenPose algorithm. Specifically, the rank adjustment coefficient is calculated based on the posture key points, and the LoRA matrix rank of the image is dynamically adjusted according to the formula. For example, in the cuff area, due to the larger curvature of the joint, a higher rank will be assigned to retain more wrinkle details; in relatively flat areas, a lower rank can be used, thereby ensuring details while optimizing the use of computing resources. This dynamic adjustment mechanism makes it possible to flexibly control the degree of detail retention in different parts according to actual needs, further enhancing the realism and naturalness of the dressing effect.

[0040] In addition, based on this embodiment, multiple experiments were conducted to obtain the empirical values ​​of the joint rank adjustment coefficients, as shown in Table 3.

[0041] Table 3

[0042] Through the above steps, this method can generate high-quality, natural, and realistic clothing change images. The FLux Fill model's Inpaint function, combined with a pose estimation algorithm, ensures a natural fit between the garment contour and the model's pose. The In-Context LoRA technology's feature splicing in latent space further improves detail consistency and image quality. Dynamic low-rank adaptation flexibly adjusts the degree of detail preservation based on pose keypoints, making the clothing change effect more realistic and natural. The combined application of these technologies not only simplifies user operations but also significantly improves the quality of clothing change effects. This brings new technological breakthroughs to related industries and meets the diverse needs of users in different scenarios, such as digital human clothing change, virtual fitting, and social media content creation.

[0043] See also Figure 3 In summary, the digital human costume-changing method aims to provide users with an efficient, flexible, and high-quality costume-changing experience through advanced image processing and artificial intelligence technologies. Through a series of innovative techniques, this method addresses the key issues of existing costume-changing technologies, such as operational complexity, low reproduction of clothing details, and the inability to maintain facial consistency.

[0044] Specifically, during the implementation process, the user-provided clothing image is first acquired and recognized using an advanced multimodal visual language model to generate clothing type recognition results. This process not only accurately identifies whether the garment is a top or bottom or a set, but also provides accurate clothing type information for subsequent outfit changes, ensuring the naturalness and consistency of the outfit changes. This high-precision recognition capability eliminates the need for users to manually specify the clothing type; it is automatically identified and processed, greatly simplifying the operation process and improving the user experience.

[0045] Subsequently, feature extraction is performed on the clothing images. An advanced visual encoder, combined with a multi-layer Transformer encoder, is used to extract global semantic features from the clothing images. Based on this, a cross-modal feature vector construction technique is combined with semantic alignment to generate a high-dimensional feature vector. This process ensures the accuracy and consistency of clothing features by introducing semantic constraints, resulting in more realistic clothing details and a high degree of restoration of features such as texture, wrinkles, and material. Furthermore, a large language model is used to fine-tune the prompt words, further enhancing the realism and naturalness of the costume change effects.

[0046] In processing the model image, an advanced segmentation model and semantic segmentation techniques are combined to segment the model image. A context-aware model is used to correlate the model image with the human pose during segmentation, resulting in the model image redrawing region. The segmentation results are further optimized by calculating the semantic similarity between each mask and the textual hint and retaining the best candidates. Semantic segmentation techniques then refine these candidates, locating clothing parts within the model image using a fine-grained semantic query vector and an attention mechanism. This process not only accurately segments the clothing regions but also dynamically adjusts based on human pose, ensuring the naturalness and accuracy of the segmentation results. Furthermore, a pose estimation algorithm is used to estimate the pose of the model image, generating association results based on contextual association. By performing pose-guided fusion on the association results and the located clothing parts, and optimizing their edges, the model image redrawing region is finally segmented. This series of operations ensures a natural fit between the clothing regions and the model's pose, further enhancing the realism of the costume change effect.

[0047] During the final image generation phase, the garment features and the redrawn regions of the model image are redrawn using an advanced image redrawing model. The model's fill function, combined with a pose estimation algorithm, is used to adapt the garment contours to the redrawn regions. Context-aware technology is also introduced to construct the splicing features of the garment plan and the redrawn regions in the latent space, and dynamic low-rank adaptation techniques are used to control detail consistency in the adapted images. This process dynamically adjusts parameters to flexibly control the degree of detail preservation based on pose key points, further enhancing the realism and naturalness of the costume change effects.

[0048] Overall, this method achieves efficient, flexible, and high-quality costume changes through a series of innovative technologies. From high-precision recognition and feature extraction of clothing images, to precise segmentation and pose-guided fusion of model images, to the final generation of high-quality costume change images, each step demonstrates the technological advancement and practicality. The integrated application of these technologies not only simplifies user operations but also significantly improves the quality of costume changes. This provides a novel solution for application scenarios such as digital human costume changes, virtual fitting, and social media content creation, driving technological advancement in related industries.

[0049] See also Figure 4 The second embodiment of the present invention provides a digital human portrait dressing device, which includes: The clothing type recognition unit 101 is used to obtain a clothing image, recognize the clothing image using a VL visual model, and generate a clothing type recognition result; A clothing feature extraction unit 102 is configured to extract and process the clothing image based on the clothing type recognition result, generate a cross-modal feature vector and a high-dimensional feature vector, and introduce semantic constraints corresponding to the clothing image to obtain clothing features; The redrawing region segmentation unit 103 is used to obtain a model image, perform segmentation processing on the model image by combining the SAM segmentation model and the G-DINO semantic segmentation technology, and associate the model image with the human body posture during the segmentation processing based on the CLIP model to obtain the model image redrawing region; The redrawing unit 104 is used to redraw the clothing features and the model image redrawing area according to the FLux Fill model to generate a final clothing change image.

[0050] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for changing the appearance of a digital human portrait, characterized in that: include: Acquire a clothing image, identify the clothing image using a VL visual model, and generate a clothing type recognition result; Based on the clothing type recognition result, the clothing image is extracted and processed to generate a cross-modal feature vector and a high-dimensional feature vector, and semantic constraints corresponding to the clothing image are introduced to obtain clothing features; Obtain a model image, perform segmentation processing on the model image by combining the SAM segmentation model and the G-DINO semantic segmentation technology, and associate the model image with the human body posture during the segmentation processing based on the CLIP model to obtain a redrawing area of ​​the model image; The clothing features and the model image redrawing area are redrawn according to the FLux Fill model to generate the final dressing image.

2. The method for changing the digital human portrait according to claim 1, characterized in that: The VL visual model is a Qwen-VL multimodal visual language model, and the clothing type recognition result includes tops and bottoms or suits.

3. The method for changing the digital human portrait according to claim 1, characterized in that: Based on the clothing type recognition result, the clothing image is extracted and processed to generate a cross-modal feature vector and a high-dimensional feature vector, specifically: The image block embedding module of the ViT architecture visual encoder and the 12-layer Transformer encoder are used to extract the clothing image to obtain global semantic features; According to the global semantic features, the redux model is used to construct and generate cross-modal feature vectors, and the CLIP visual encoding based on clothing images and text descriptions are semantically aligned to generate high-dimensional feature vectors.

4. The method for changing the digital human portrait according to claim 3, characterized in that: And introduce semantic constraints corresponding to the clothing image to obtain clothing features, specifically: Obtaining the prompt words of the original generated image, and performing tuning processing on the prompt words based on the LLM large language model, wherein the tuning processing includes using the joy caption two model to reversely infer the prompt words of the image; Clothing features are obtained based on cross-modal feature vectors, high-dimensional feature vectors and optimized prompt words.

5. The method for changing the digital human portrait according to claim 1, characterized in that: Obtain a model image, segment the model image using the SAM segmentation model and G-DINO semantic segmentation technology, and associate the model image with the human body posture during segmentation based on the CLIP model to obtain the model image redrawing area, specifically: Get the model image, use the SAM segmentation model through the CLIP model to calculate the semantic similarity between each mask of the model image and the text prompt, and retain its top-3 candidates; The top-3 candidates are modified based on the G-DINO semantic segmentation technology. The text encoder of the CLIP model generates a fine-grained semantic query vector, which is dynamically decoded and combined with the variable attention mechanism to locate the clothing parts in the model image. The formula is: ,in, is the query vector, is the reference point coordinate, is the input model image, is the number of sampling points, is the learnable offset, is the weight coefficient; Performing pose estimation on the model image using the OpenPose algorithm to generate a pose estimation result, and performing context association on the model image based on the CLIP model to generate an association result; The association results and the clothing parts in the located model image are fused with posture guidance, and their edges are optimized to segment and generate the model image redrawing area.

6. The method for changing the digital human portrait according to claim 1, characterized in that: The clothing features and the model image redrawing area are redrawn according to the FLux Fill model to generate the final dressing image, specifically: The Inpaint function based on the FLux Fill model combines the posture estimation algorithm to adapt the clothing features to the redrawn area of ​​the model image; The splicing features of the clothing plane map and the redrawn area constructed in the latent space by the In-Context LoRA technology are introduced. The formula of the splicing feature is: , is the potential feature of the clothing plan, To redraw the potential features of the region mask, is the clothing feature projection matrix, is the mask feature projection matrix, For splicing operations, for dimensional real matrix space; Based on the splicing features, a low-rank adapter is used to perform dynamic low-rank adaptation on the image after clothing contour adaptation to achieve consistent control of image details and generate a clothing change image.

7. The method for changing the digital human portrait according to claim 6, characterized in that: Dynamic low-rank adaptation processing includes dynamically adjusting the LoRA matrix rank of the image using the joint bending posture key points identified by the OpenPose algorithm; Among them, the rank adjustment coefficient is calculated according to the posture key points , and according to the formula , dynamically adjust the LoRA matrix rank of the image, where, is the adjusted rank, As the basic rank.

8. A digital human portrait dressing device, characterized in that: include: A clothing type recognition unit is used to obtain a clothing image, identify the clothing image using a VL visual model, and generate a clothing type recognition result; a clothing feature extraction unit, configured to extract and process the clothing image based on the clothing type recognition result, generate a cross-modal feature vector and a high-dimensional feature vector, and introduce semantic constraints corresponding to the clothing image to obtain clothing features; A redrawing region segmentation unit is used to obtain a model image, perform segmentation processing on the model image by combining the SAM segmentation model and the G-DINO semantic segmentation technology, and associate the model image with the human body posture during the segmentation processing based on the CLIP model to obtain a redrawing region of the model image; The redrawing unit is used to redraw the clothing features and the model image redrawing area according to the FLux Fill model to generate the final dressing image.

Citation Information

Cited By

  • AI model chart generation method, system and device and medium

    CN121685752A