Portrait image head changing method and device based on diffusion model, and medium

By constructing a feature extraction module and a ControlNet module, combined with a diffusion model, facial and hair features are extracted and injected, solving the problem of insufficient identity and hairstyle preservation in the head-changing task, and achieving high-similarity image generation.

CN120656215APending Publication Date: 2025-09-16BEIJING SHUNSHI BROTHERS TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510032571.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively preserving the identity features and hairstyle of the source image in head-changing tasks, resulting in insufficient identity similarity and hairstyle similarity in the generated image.

Method used

By constructing a feature extraction module and a ControlNet module, combined with a diffusion model, the feature information of the face and hair is extracted and injected, and the model is trained using a local redrawing method to generate the target image.

Benefits of technology

It achieves the goal of accurately preserving the identity and hairstyle features of the source image while keeping the background of the target image unchanged, thereby improving the identity and hairstyle similarity of the generated image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656215A_ABST
    Figure CN120656215A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a portrait image head changing method and device based on a diffusion model and a medium, and the method comprises the following steps: obtaining image information which comprises a source image and a template image; constructing a feature extraction module, training the feature extraction module to obtain a feature extractor, training a ControlNet module of a diffusion model, and combining the ControlNet module with the diffusion model to form a head change task model; and inputting the source image and the template image into the head changing task model to obtain a target image. According to the method, the features of the whole head are extracted during feature extraction, and meanwhile, the spatial control condition contains hair edge detection, so that the hair style similarity of the generated image can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosed embodiments relate to the field of image processing technology, and more particularly to a method, device, and medium for head replacement of a portrait image based on a diffusion model. Background Art

[0002] Image generation technology has made rapid progress. Advances in generative adversarial networks (GANs) and graph diffusion models such as GLIDE, DALL E 2, and Stable Diffusion (SD) have introduced numerous downstream tasks. Identity-preserving image generation involves generating images of a person in various scenarios, given a portrait image, while preserving their identity (i.e., maintaining a high degree of facial similarity). Face swapping involves taking a source image containing a portrait and a template image and aiming to naturally swap the source image's face onto the template image, while preserving the template image's hair, clothing, and background. Currently, most research on these two areas is based on the SD model, and the basic approach involves extracting facial features from a pre-trained model and then inserting them into the SD model's Unified Network (UNet).

[0003] Head swapping involves replacing the hairstyle of the source image with that of the template image. The difficulty lies in the diversity of hairstyles, unlike faces, which can be standardized. This makes it difficult to predict hair changes during pose changes. Current face swapping tasks only consider face replacement, and identity-preserving portrait generation also focuses on maintaining facial similarity, without specifically focusing on hairstyle preservation. This also poses challenges in maintaining identity similarity. Summary of the Invention

[0004] To this end, the embodiments of the present disclosure provide a method, device, and medium for head replacement of a portrait image based on a diffusion model to solve the problem in related technologies that it is difficult to maintain the identity characteristics of the source image when generating a head replacement image due to the inability to maintain and extract the identity of the portrait.

[0005] In order to achieve a method for preserving the identity features of a source image and replacing it with a target image while keeping the background of the target image unchanged, the embodiments of the present disclosure provide the following technical solutions:

[0006] In a first aspect of the embodiments of the present disclosure, a method for replacing a head of a portrait image based on a diffusion model is provided, comprising the following steps:

[0007] S1. Acquire image information, where the image information includes a source image and a template image;

[0008] S2. Build a feature extraction module and train it to obtain a feature extractor, and train the ControlNet module of the diffusion model, and combine it with the diffusion model to form a head-changing task model;

[0009] S3. Input the source image and template image into the head-changing task model to obtain the target image.

[0010] In one embodiment, the step of training the feature extractor includes

[0011] Construct a structural model that preserves identity and hairstyle information;

[0012] Extracting identity and hairstyle information from the training image, and injecting the extracted identity and hairstyle information into the structural model to obtain the extracted features;

[0013] Inject the extracted features into the ControlNet module.

[0014] The model is trained using a local redrawing method to obtain a feature extractor.

[0015] In one embodiment, the step of constructing the structural model includes:

[0016] The ControlNet module is used to extract the face and head features as the cross-attention input, and the landmark (facial key feature points) and hair canny (edge ​​detection) are used as spatial control conditions. The information is encoded and injected into the SD Unet.

[0017] The spatial control information of the ControlNet module, the landmark is extracted through the OpenPose (pose estimation library) pose estimation algorithm, and the hair retention is extracted through the canny edge detection algorithm.

[0018] The model is trained using a local redrawing method to obtain the ControlNet weights.

[0019] In one embodiment, the step of extracting identity and hairstyle information includes:

[0020] The source image is intercepted by the ArcFace (deep metric learning method for face recognition) face recognition model to extract the face area, and the head area is segmented by the FaRL model;

[0021] For the segmented image, the ArcFace face recognition model is used to extract facial embedding features from the face image, and the CLIP (multimodal model based on contrastive learning) image encoder is used to extract corresponding CLIP image features from the face region image and the segmented head image;

[0022] Project the image features to the same dimension using a linear layer;

[0023] The projected face and head features are merged respectively, and the merged features are injected into the structural model.

[0024] In one embodiment, the method of training the model by using local redrawing includes:

[0025] First, a portrait image of another identity is selected and pose-aligned with the training image using the LivePortrait face reenactment method.

[0026] The head mask extracted by the FaRL model is combined with the head mask of the training image as the final mask for local redrawing.

[0027] When training using the local redrawing method, the feature extractor and ControlNet module are trained simultaneously.

[0028] In one embodiment, a method for inputting a source image and a template image into a head-changing task model to obtain a target image includes:

[0029] Replay the source image and adjust the pose to align with the template image;

[0030] Construct a landmark predictor to obtain a landmark image that retains the face shape of the source image while having the pose of the template image;

[0031] The hair canny of the replayed source image is extracted by using the canny edge detection algorithm and merged with the landmark to obtain the spatial control image;

[0032] extracting face and hair features of the replayed source image using a feature extractor to obtain extracted face and hair features;

[0033] Segment and extract the head mask of the source image and the head mask of the template image after reenactment, and merge them as the local redrawn mask image;

[0034] The spatial control image and features obtained in the above steps are input into the ControlNet module to generate a fused feature, and then the fused feature is input into the Unet of the SD model that has received the template image and the redrawn mask to generate the target image.

[0035] In one embodiment, the method for generating the landmark image includes:

[0036] Use DECA (face reconstruction model) 3D reconstruction network to extract face shape, posture, and expression parameters from source and template images;

[0037] The facial parameters of the source image are combined with the expression and posture parameters of the template image, and the 3D face is reconstructed using the 3DMM method, and then the landmark image is extracted.

[0038] In a second aspect of the embodiments of the present disclosure, a head-changing device for portrait images based on a diffusion model is provided, comprising:

[0039] Acquisition module: used for acquiring image information, wherein the image information includes a source image and a template image;

[0040] Training module: used to construct a feature extraction module based on the source image and train a feature extractor;

[0041] Image synthesis module: used to use the information of the feature extractor to replay the template image and merge it to obtain the target image.

[0042] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, comprising an input device and an output device, and further comprising

[0043] a processor adapted to implement one or more instructions; and

[0044] Computer-readable storage medium, wherein at least one instruction, at least one program, code set or instruction set is stored on the computer-readable storage medium, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the steps of any of the above methods.

[0045] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, characterized in that at least one instruction, at least one program, a code set or an instruction set is stored on the computer-readable storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the steps of any of the above methods.

[0046] According to the embodiment of the present disclosure, the method has the following advantages: the method comprises the following steps: obtaining image information, wherein the image information includes a source image and a template image; constructing a feature extraction module based on the source image and training a feature extractor; and using the information of the feature extractor to replay the template image and merge them to obtain a target image.

[0047] 1. The present invention extracts the features of the entire human head when extracting features. At the same time, the spatial control conditions include hair canny, which can ensure the hairstyle similarity of the generated image.

[0048] 2. The present invention determines the local redrawing mask by using the face reenactment method, which can provide a more accurate mask and prevent the inaccurate mask from affecting the generation result.

[0049] 3. The landmark predictor in the present invention can generate a landmark with the facial features of the source image, can retain the facial shape of the source image while determining the posture, and can improve the facial similarity of the generated result.

[0050] 4. In the present invention, multiple facial and head features are extracted through multiple pre-trained models and then fused, which can improve the similarity of the generated results compared to using a single feature. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] To more clearly illustrate the embodiments of the present disclosure or the technical solutions in the related technologies, the following briefly introduces the drawings required for the embodiments or the related technical descriptions. Obviously, the drawings described below are merely exemplary, and those skilled in the art can, without inventive effort, derive other implementation drawings based on the provided drawings.

[0052] Figure 1 4 is a flowchart of a method for replacing a head of a portrait image based on a diffusion model according to an exemplary embodiment;

[0053] Figure 2 The overall structure and training flow chart of a model generated in a method according to an exemplary embodiment are shown;

[0054] Figure 3 is a structural diagram of a feature extractor in a method according to an exemplary embodiment;

[0055] Figure 4 1 is a Resampler projection layer structure in a feature extractor in a method according to an exemplary embodiment;

[0056] Figure 5 1 is a structural diagram of a head-changing task model in a method according to an exemplary embodiment;

[0057] Figure 6 is a structural diagram of a landmark predictor in a method according to an exemplary embodiment;

[0058] Figure 7 It is a structural diagram of a device according to an exemplary embodiment. DETAILED DESCRIPTION

[0059] The following describes the implementation methods of the present disclosure using specific embodiments. Those skilled in the art can readily understand the other advantages and benefits of the present disclosure from the contents disclosed herein. Obviously, the described embodiments are only a portion of the embodiments of the present disclosure, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without inventive effort are also within the scope of protection of the present disclosure.

[0060] The realization of the head-changing task mainly lies in maintaining the identity and hairstyle of the portrait. By maintaining the identity and hairstyle of the source image in the training model and determining the generation position of the template image, the portrait with the maintained identity and hairstyle is regenerated at the determined position to obtain the target image.

[0061] First, we build a feature extractor to extract the person's identity information from the image, ultimately integrating it into a facial feature (information-wise, it's called a facial feature, but in terms of the diffusion model input, it's called a token, which can be considered an alias). In addition to the identity information of the source image, we also use spatial control information to control the position of the face and hair in the generated image (the local redrawn result), namely the landmark and hair canny. The landmark determines the head pose in the generated image, and the canny determines the hair area.

[0062] The source image's identity information and spatial control information are obtained. Note that feature extraction is only one stage; once this information is obtained, it must be passed to the diffusion generative model. Regarding information injection into the diffusion generative model, this can be achieved by using the entire ControlNet model or by inserting an additional module into the diffusion generative model. Alternatively, it can be understood as ControlNet fusing these two types of information and then injecting them into the diffusion generative model in some way. In fact, there are many methods for injecting information into the diffusion generative model, each with different requirements, model structures, and final results. Here, ControlNet is used.

[0063] It's important to note that during training, noise is added to the input image and the original image is restored from the noisy image. This involves the principle of the diffusion generative model. Simply put, the diffusion generative model continuously removes noise from a pure noise image to obtain a single image. Therefore, theoretically, each training step requires only one image. Noise is first added to this image, and then, under certain conditions (in this head-swapping method, the two types of information mentioned above), the image is restored. In this process, the optimal weights for the module to be trained are found. Training using the local redrawing method also requires a redrawing mask, and only the area within the redrawing mask is targeted. The redrawing mask used in training is the head region of the training image. To align the training conditions with the inference conditions (the redrawing mask during inference is a combination of the head regions of the redrawn source image and the template image after redrawing; obviously, this region may not be entirely the head region and may include some background), a second image with a different identity is introduced. After redrawing, the head region is extracted and merged with the head region of the training image to form the training redrawing mask.

[0064] like Figure 1 As shown, it is a method for replacing a head of a portrait image based on a diffusion model according to an exemplary embodiment of the present invention. The method includes the following steps:

[0065] S1. Acquire image information, which includes a source image and a template image;

[0066] S2. Build a feature extraction module and train it to obtain a feature extractor, and train the ControlNet module of the diffusion model, and combine it with the diffusion model to form a head-changing task model;

[0067] S3. Input the source image and template image into the head-changing task model to obtain the target image.

[0068] In one embodiment, the step of training the feature extractor includes

[0069] Build and maintain a structural model of the identity and hairstyle information;

[0070] like Figure 2 The overall structure and training process of the generative model shown in the figure. The steps of constructing the structural model include

[0071] The ControlNet module is used to extract the face and head features as the cross-attention input, and the landmark (facial key feature points) and hair canny are used as spatial control conditions. The information is encoded and injected into the SD Unet.

[0072] And for the spatial control information of the ControlNet module, the landmark is extracted through the OpenPose (pose estimation library) pose estimation algorithm, the hair retention is extracted through the canny edge detection algorithm, and the hair canny is extracted through the canny edge detection algorithm.

[0073] The model is trained using a local redrawing method to obtain the ControlNet weights.

[0074] In one embodiment, for the extraction of identity and hairstyle information, the feature extractor designed by the present invention is as follows Figure 3 shown.

[0075] The steps of extracting identity and hairstyle information include

[0076] For an input face image, the source image is first intercepted by the ArcFace (deep metric learning method for face recognition) face recognition model to extract the face area, and the head area is segmented by the FaRL model;

[0077] For the segmented image, the ArcFace face recognition model is used to extract the face embedding features from the face image, and the CLIP (multimodal model based on contrastive learning) image encoder is used to extract the corresponding CLIP image features from the face region image and the segmented head image;

[0078] Use a linear layer to project the image features to the same dimension;

[0079] Specifically, we first use three linear layers π1, π2, and π3 to project them to the same dimension d:

[0080] F′ hc =π1(F hc )

[0081] F′ fc =π2(F fc )

[0082] F′ fe =π3(F fe )

[0083] Among them, F hc is the CLIP image feature of the human head, F he For face CLIP features, F fc Embedding features for faces.

[0084] For the CLIP image features of the human head, the learnable embedding F of 16×d dimension is used. l As query, human head CLIP image feature F hc'After the latent passes through the Resampler projection layer, we get the 16×d dimension head tokenE h , which is calculated as follows:

[0085] Q=F l W q , K=Cat[F l , F hc ′]W k , V=Cat[F l , F hc ′]W v

[0086]

[0087] Among them, W q , W k , W v They are the query, key, and value matrices in attention calculation respectively.

[0088] For the face embedding feature F fe ' and face CLIP image features F fc ', with face embedding feature F fe 'As query, face CLIP image feature F fc 'As latent, after the Resampler projection layer, a 4×d-dimensional face token is obtained, which is calculated as follows:

[0089] Q′=F fe 'W' q , K′=Cat[F fe ′,F fc ′]W′ k , V′=Cat[F fe ′,F fc ′]W′ v

[0090]

[0091] Among them, W' q’ , W' k’ , W' v They are the query, key, and value matrices in attention calculation respectively.

[0092] Finally, the face tokenE f and tokenE h After merging, we get tokenE of 20×d dimension id :

[0093] E id=Cat[E f , E h ]

[0094] Where d = 2048, which is the tensor dimension in the cross-attention calculation of the SDXL model. The final tokenE id Injected into the ControlNet module through cross-attention calculation.

[0095] Among them, such as Figure 4 As shown in the figure, the Resampler (resampling module) projection layer structure of the feature extractor is as follows: the input query is q, the input query and latent are combined as k, v, and the output result is obtained after attention calculation and projection.

[0096] Extracting identity and hairstyle information from the source image, and injecting the extracted identity and hairstyle information into the structural model to obtain projected features;

[0097] Among them, the method of training the model by local redrawing includes

[0098] First, a portrait image of another identity is selected and pose-aligned with the training image using the LivePortrait face reenactment method.

[0099] The head mask extracted by the FaRL model is combined with the head mask of the training image as the final mask for local redrawing.

[0100] When training using the local redrawing method, the feature extractor and ControlNet module are trained simultaneously.

[0101] In one embodiment, a method for inputting a source image and a template image into a head replacement task model to obtain a target image includes, given a source image and a template image, being able to generate an image that maintains the background of the template image and contains the head of the source image.

[0102] The specific process of the head replacement task model is as follows Figure 5 As shown:

[0103] Use the LivePortrait method to recreate the face of the source image so that its posture is aligned with the template image;

[0104] Construct a landmark predictor to obtain a landmark image that retains the face shape of the source image while retaining the pose of the template image.

[0105] Specifically, a landmark predictor is used to obtain the landmark used for inference. The structure of the landmark predictor is as follows: Figure 6As shown, the DECA (Deep Constructed Face Model) 3D reconstruction network is first used to extract facial shape, pose, and expression parameters from the source and template images. The source image's facial shape parameters are then combined with the template image's expression and pose parameters to reconstruct a 3D face using the 3DMM method. This method then extracts a predicted landmark image. This landmark retains the face shape of the source image while retaining the pose of the template image, improving identity similarity in the generated image.

[0106] The hair canny of the reconstructed source image is extracted by using the canny edge detection algorithm and then merged with the landmark prediction image to obtain the spatial control image.

[0107] Use Figure 3 The feature extractor shown extracts the face and hair features of the replayed source image to obtain extracted face and hair features;

[0108] Use the FaRL model to segment and extract the head mask of the source image and the template image after reenactment (including the neck area to eliminate the skin color difference after head replacement), and merge them as the local redrawing mask image;

[0109] The spatial control image and features obtained in the above steps are input into the ControlNet module to generate a fused feature, and then the fused feature is input into the Unet of the SD model that has received the template image and the redrawn mask to generate the target image.

[0110] like Figure 7 As shown, in the second aspect of the embodiment of the present disclosure, a head-changing device for a portrait image based on a diffusion model is provided, which is applied to a head-changing method for a portrait image based on a diffusion model. The device includes

[0111] Acquisition module: used to acquire image information, which includes source image and template image;

[0112] Training module: used to construct a feature extraction module based on the source image and train the feature extractor;

[0113] Image synthesis module: used to use the information of the feature extractor to replay the template image and merge it to obtain the target image.

[0114] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, comprising an input device and an output device, and further comprising

[0115] a processor adapted to implement one or more instructions; and

[0116] Computer-readable storage medium, the computer-readable storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the steps of any of the above methods.

[0117] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, characterized in that at least one instruction, at least one program, code set or instruction set is stored on the computer-readable storage medium, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the steps of any of the above methods.

[0118] The above-mentioned embodiment method can be implemented by means of software plus the necessary general hardware platform, or of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the relevant technology, can be embodied in the form of a software product, and the limitations of the hardware structure platform cannot be associated with the limitations of the implementation of the method of the present disclosure. Therefore, the embodiments of the present method are all applicable to the electronic device and storage medium, and can achieve the same or similar beneficial effects.

[0119] Although the present disclosure has been described in detail above using general descriptions and specific embodiments, it will be apparent to those skilled in the art that modifications or improvements may be made based on the present disclosure. Therefore, such modifications or improvements, which do not depart from the spirit of the present disclosure, are intended to be within the scope of protection claimed by the present disclosure.

Claims

1. A head-changing method for portrait images based on a diffusion model, characterized in that: Includes the following steps S1. Acquire image information, where the image information includes a source image and a template image; S2. Build a feature extraction module and train it to obtain a feature extractor, and train the ControlNet module of the diffusion model, and combine it with the diffusion model to form a head-changing task model; S3. Input the source image and template image into the head-changing task model to obtain the target image.

2. The head-changing method for portrait images based on a diffusion model according to claim 1, characterized in that: The steps of training the feature extractor include: Construct a structural model that preserves identity and hairstyle information; Extracting identity and hairstyle information from the training image, and injecting the extracted identity and hairstyle information into the structural model to obtain the extracted features; Inject the extracted features into the ControlNet module. The model is trained using a local redrawing method to obtain a feature extractor.

3. The head-changing method for portrait images based on a diffusion model according to claim 2, characterized in that: The steps of constructing the structural model include: The ControlNet module is used to extract the face and head features as the cross-attention input, and the landmark and hair edge information is used as the spatial control condition. The information is encoded and injected into the SD Unet. The spatial control information of the ControlNet module, landmarks are extracted using the OpenPose pose estimation algorithm, and hair retention is extracted using the Canny edge detection algorithm; The model is trained using a local redrawing method to obtain the ControlNet weights.

4. The head-changing method for portrait images based on a diffusion model according to claim 2, characterized in that: The steps of extracting identity and hairstyle information include: The source image is intercepted by the ArcFace face recognition model to extract the face area, and the FaRL model is used to segment the head area; For the segmented images, the ArcFace face recognition model is used to extract face embedding features from the face image, and the CLIP image encoder is used to extract corresponding CLIP image features from the face region image and the segmented head image. Project the image features to the same dimension using a linear layer; The projected face and head features are merged respectively, and the merged features are injected into the structural model.

5. The head-changing method for portrait images based on a diffusion model according to claim 2, characterized in that: The method of training the model by using local redrawing includes: First, a portrait image of another identity is selected and pose-aligned with the training image using the LivePortrait face reenactment method. The head mask extracted by the FaRL model is combined with the head mask of the training image as the final mask for local redrawing. When training using the local redrawing method, the feature extractor and ControlNet module are trained simultaneously.

6. The head-changing method for portrait images based on a diffusion model according to claim 1, characterized in that: The method of inputting the source image and the template image into the head-changing task model to obtain the target image includes replaying the source image and adjusting the posture to align with the template image; Construct a landmark predictor to obtain a landmark image that retains the face shape of the source image while having the pose of the template image; The hair edge information of the reconstructed source image is extracted by using the Canny edge detection algorithm and then merged with the landmark to obtain the spatial control image. extracting face and hair features of the replayed source image using a feature extractor to obtain extracted face and hair features; Segment and extract the head mask of the source image and the head mask of the template image after reenactment, and merge them as the local redrawn mask image; The spatial control image and features obtained in the above steps are input into the ControlNet module to generate a fused feature, and then the fused feature is input into the Unet of the SD model that has received the template image and the redrawn mask to generate the target image.

7. The head-changing method for portrait images based on a diffusion model according to claim 6, characterized in that: The method for generating the landmark image includes: Use DECA's 3D reconstruction network to extract facial shape, posture, and expression parameters from the source image and template image; The facial parameters of the source image are combined with the expression and posture parameters of the template image, and the 3D face is reconstructed using the 3DMM method, and then the landmark image is extracted.

8. A head-changing device for portrait images based on a diffusion model, characterized in that: include Acquisition module: used for acquiring image information, wherein the image information includes a source image and a template image; Training module: Build a feature extraction module and train the feature extractor, which is then combined with the diffusion model to form a head-swapping task model; Image synthesis module: The source image and template image are input into the head-changing task model to obtain the target image.

9. An electronic device comprising an input device and an output device, characterized in that: Also includes a processor adapted to implement one or more instructions; and A computer-readable storage medium having stored thereon at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the steps of any method of claims 1-7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by the processor to implement the steps of any method in claims 1-7.

Citation Information

Cited By

  • Head portrait splicing method and device based on extension model, and storage medium

    CN121353074A

  • Avatar splicing method and device based on extended model and storage medium

    CN121353074B