Method and apparatus for updating a person image
By iteratively updating the initial character image using a neural network model and utilizing the pre-existing visual and morphological information of the preset character, the problems of hand occlusion and posture generation in virtual try-on are solved, generating high-quality virtual try-on images and improving naturalness and realism.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2024-11-30
- Publication Date
- 2026-06-02
AI Technical Summary
Existing virtual try-on technologies are prone to clothing distortion when generating virtual object hand and arm poses, and have difficulty handling hand occlusion issues. Existing solutions such as GANs and diffusion models have limitations in applicability and accuracy.
By utilizing the visual and morphological prior information of the preset character's hands, the initial character image is iteratively updated to generate the final character image of the preset character wearing preset clothing with hands occluded in the clothing area. A neural network model including a hand feature separation and embedding module, a hand pose aggregation module, a try-on module, and a clothing module is used for training and updating.
The generated final character image has significantly higher naturalness and realism than traditional virtual try-on solutions. The interaction between the hands and clothing is natural and smooth, the clothing details are realistic, and the areas obscured by the hands are handled accurately.
Smart Images

Figure CN122134983A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence services, and more specifically to a method for updating human images, an apparatus for updating human images, an electronic device, and a computer-readable storage medium. Background Technology
[0002] In recent years, Virtual Try-On (VTON) technology has demonstrated enormous potential in the e-commerce sector. The core purpose of VTON technology is to reshape the consumer shopping experience while significantly reducing advertising costs for apparel retailers. The essence of VTON lies in creating highly realistic and natural virtual avatars wearing clothing, while ensuring that clothing details are clearly visible. One branch of VTON technology is image-based VTON. Leveraging the powerful support of augmented reality (AR), artificial intelligence (AI), machine learning (ML) algorithms, and 3D modeling technology, image-based VTON can seamlessly overlay virtual clothing onto a user's personal photo taken with their mobile phone camera, presenting the user with a virtual object simulating the user wearing the clothing. VTON technology not only greatly enhances the fun and personalization of online shopping but also effectively reduces returns due to size or style mismatches, significantly improving customer satisfaction and loyalty.
[0003] In the field of image-based virtual try-on, the goal is to generate a virtual object of a person wearing specific clothing based on an initial image of the person. However, virtual objects generated by traditional techniques have stiff movements, limited hand and arm poses, and are prone to clothing distortion when generating arbitrary hand and arm poses.
[0004] To address this issue, two approaches have been proposed: refining the hand (or arm) parsing model and integrating hand (or arm) reconstruction into the clothing reconstruction process within the target image. The former faces challenges in segmentation accuracy and image artifacts. The latter utilizes methods such as Generative Adversarial Networks (GANs) and diffusion models. While the GAN approach can minimize the differences between clothing and the virtual object's body, its applicability is limited. The diffusion model approach improves image quality by preserving clothing details and fine-tuning the repair model, but still faces challenges such as hand occlusion.
[0005] Therefore, existing virtual try-on technology needs further improvement. Summary of the Invention
[0006] This disclosure provides a method for updating a person's image, a method for training a neural network, an apparatus for updating a person's image, an electronic device, a computer-readable storage medium, and a computer program product.
[0007] This disclosure provides a method for updating a person image. The method includes: obtaining a clothing image and an initial person image, wherein the clothing image includes a preset clothing, and the initial person image includes a preset person wearing the current clothing, wherein the area where the current clothing is located in the initial person image is occluded by the preset person's hand; determining clothing prior information of the preset clothing based on the clothing image; determining reference features corresponding to other areas outside the area where the current clothing is located, as well as visual prior information and morphological prior information corresponding to the preset person's hand based on the initial person image; and iteratively updating the initial person image based on the clothing prior information, the reference features, the visual prior information, and the morphological prior information to generate a final person image, wherein the final clothing worn by the preset person in the final person image corresponds to the preset clothing, and the area where the final clothing is located in the final person image has a corresponding hand occlusion.
[0008] This disclosure provides a method for training a neural network model. The neural network model includes a hand feature separation and embedding module, a hand pose aggregation module, a try-on module, and a clothing module. The method includes: obtaining clothing image samples and initial person image samples. The clothing image samples include preset clothing, and the initial person image samples include a preset person wearing the current clothing. The area where the current clothing is located in the initial person image samples is occluded by the hands of the preset person. Based on the clothing image samples, using the clothing module, determine the predicted value of the clothing prior information of the preset clothing. Based on the initial person image samples, determine the predicted value of reference features corresponding to the area outside the area where the current clothing is located. Based on the initial person image samples, using the hand feature separation and embedding module and the hand pose aggregation module, determine the preset person... The system uses the try-on module to iteratively update the initial character image samples to generate a predicted image of the final character image. The predicted image is generated based on the predicted values of the clothing prior information, the reference features, and the visual and morphological prior information corresponding to the hands of the preset character. In the predicted image of the final character image, the final clothing worn by the preset character corresponds to the preset clothing, and the area where the final clothing is located in the predicted image of the final character image has corresponding hand occlusion. In each iteration, the hand feature separation and embedding module, the hand pose aggregation module, the try-on module, and the clothing module are trained based on the difference between the predicted image of the final character image determined in that iteration and the final character image samples.
[0009] This disclosure provides an apparatus for updating a person image, characterized in that the apparatus includes: an acquisition module, a masking module, a hand feature separation and embedding module, a hand pose aggregation module, a try-on module, and a clothing module. The acquisition module is configured to: acquire a clothing image and an initial person image, wherein the clothing image includes a preset clothing, and the initial person image includes a preset person wearing the current clothing, wherein the area where the current clothing is located in the initial person image is occluded by the hand of the preset person. The clothing module is configured to: determine prior information of the preset clothing based on the clothing image. The masking module is configured to: determine reference features corresponding to the area outside the area where the current clothing is located based on the initial person image. The hand feature separation and embedding module and the hand pose aggregation module are configured to: jointly determine the morphological prior information corresponding to the hands of the preset person based on the initial person image; the hand feature separation and embedding module is also configured to: determine the visual prior information corresponding to the hands of the preset person based on the initial person image; and the try-on module is configured to: iteratively update the initial person image based on the clothing prior information, the reference features, and the visual prior information and the morphological prior information to generate a final person image, wherein in the final person image, the final clothing worn by the preset person corresponds to the preset clothing, and the area where the final clothing is located in the final person image has corresponding hand occlusion.
[0010] This disclosure provides an electronic device, including: one or more processors; and one or more memories, wherein the memories store a computer-executable program, and when the processor executes the computer-executable program, the above-described method is performed.
[0011] This disclosure provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the above-described method.
[0012] According to another aspect of this disclosure, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable medium and executes the computer instructions, causing the computer device to perform the methods provided in the foregoing aspects or various alternative implementations of the foregoing aspects.
[0013] Compared to traditional virtual try-on technology, the embodiments of this disclosure can utilize prior visual and morphological information corresponding to the hands of a preset person to generate a final image of the preset person wearing preset clothing with their hands partially obscuring the area of the clothing by iteratively updating the initial person image. The naturalness and realism of the final person image generated using the embodiments of this disclosure are significantly higher than those of traditional virtual try-on schemes. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. The accompanying drawings in the following description are merely exemplary embodiments of this disclosure.
[0015] Figure 1 This is an example schematic diagram illustrating a scenario according to an embodiment of the present disclosure.
[0016] Figure 2 A schematic diagram illustrating an application scenario according to an embodiment of the present disclosure is shown.
[0017] Figure 3 A flowchart of a method for updating a person's image according to an embodiment of the present disclosure is shown.
[0018] Figure 4 A schematic diagram of a try-on module according to an embodiment of the present disclosure is shown.
[0019] Figure 5 A schematic diagram illustrating the determination of prior information for clothing according to an embodiment of the present disclosure is shown.
[0020] Figure 6 A schematic diagram is shown illustrating the determination of visual prior information and morphological prior information corresponding to the hand of a preset character according to an embodiment of the present disclosure.
[0021] Figure 7 It shows Figure 6 A schematic diagram of the masked cross-attention model.
[0022] Figure 8 Another schematic diagram of an apparatus according to an embodiment of the present disclosure is shown.
[0023] Figure 9 A flowchart illustrating a method for training a neural network model according to an embodiment of the present disclosure is shown.
[0024] Figure 10 A schematic diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0025] Figure 11 A schematic diagram of the architecture of an exemplary computing device according to an embodiment of the present disclosure is shown.
[0026] Figure 12 A schematic diagram of a storage medium according to an embodiment of the present disclosure is shown. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this disclosure more apparent, exemplary embodiments according to this disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments of this disclosure. It should be understood that this disclosure is not limited to the exemplary embodiments described herein.
[0028] In this specification and accompanying drawings, operations and elements that are substantially the same or similar are indicated by the same or similar reference numerals, and repeated descriptions of these operations and elements are omitted. Furthermore, in the description of this disclosure, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance or order.
[0029] To facilitate the description of this disclosure, the following introduces concepts and general principles related to this disclosure.
[0030] Optionally, the models used in embodiments of this disclosure as described below can all be artificial intelligence models, especially artificial intelligence-based neural network models. Typically, artificial intelligence-based neural network models are implemented as acyclic graphs, where neurons are arranged in different layers. Generally, a neural network model includes an input layer and an output layer, separated by at least one hidden layer. The hidden layer transforms the input received by the input layer into a representation useful for generating the output in the output layer. Nodes are fully connected to nodes in adjacent layers via edges, and there are no edges between nodes within each layer. Data received at nodes in the input layer of the neural network is propagated to nodes in the output layer via any of the hidden layers, activation layers, pooling layers, convolutional layers, etc. The input and output of the neural network model can take various forms, and this disclosure does not limit this.
[0031] Diffusion models, especially stable diffusion models, are a special type of probabilistic model that aims to estimate the data distribution p(x) by gradually reducing noise in a normally distributed variable. Given an input image x0, the noise estimation process can be expressed as Equation (1).
[0032]
[0033] In this process, the variational autoencoder ε is responsible for compressing the input image x into a low-dimensional latent space z to obtain ε(x). Subsequently, the conditional U-Net denoiser ∈ θ(Conditional U-Netdenoiser) In the latent space, using the noisy latent representation z at time step t, ... t The noise θ is estimated using the text prompt condition C extracted by the text encoder. The conditional U-NET denoiser is a deep learning model based on the U-NET structure, consisting of a symmetrical encoder and decoder. It achieves the fusion of features at different resolutions through the text prompt condition C, thereby improving the ability to denoise images under specific conditions, effectively removing image noise while preserving details.
[0034] Based on this principle, diffusion models can generate high-quality, diverse images while maintaining accurate responses to input conditions, thus demonstrating great application potential in fields such as image generation. The embodiments disclosed herein will leverage this characteristic of diffusion models to improve traditional virtual try-on technology.
[0035] The solutions provided in this disclosure involve technologies such as artificial intelligence and / or machine learning, which are specifically illustrated through the following embodiments.
[0036] First, refer to Figure 1 The present disclosure describes the application scenarios of a method for updating human images and a corresponding apparatus according to embodiments thereof. Figure 1 A schematic diagram of an application scenario 100 according to an embodiment of the present disclosure is shown, wherein a server 110 and a plurality of terminals 120 are schematically illustrated.
[0037] The neural network model of this disclosure can be integrated into various electronic devices, for example. Figure 1 The neural network model can be integrated into any electronic device in server 110 and multiple terminals 120. For example, the neural network model can be integrated into terminal 120. Terminal 120 can be a mobile phone, tablet, laptop, desktop computer, personal computer (PC), smart speaker, or smartwatch, but is not limited to these. Alternatively, the neural network model can also be integrated into server 110. Server 110 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals and servers can be directly or indirectly connected via wired or wireless communication, and this disclosure does not impose any limitations.
[0038] It is understood that the apparatus for inference using the neural network model of this disclosure can be a terminal, a server, or a system composed of a terminal and a server. The method for generating the final image of a person according to this disclosure can be executed on a terminal, on a server, or jointly by a terminal and a server.
[0039] It is worth noting that both the terminal 120 and server 110 according to the embodiments of this disclosure adhere to data protection principles, respect users' data rights, and safeguard users' data security and privacy. The terminal 120 and server 110 according to the embodiments of this disclosure will clearly inform users of the purpose, method, and scope of collecting, using, storing, transmitting, and deleting user data, and obtain users' consent. The terminal 120 and server 110 according to the embodiments of this disclosure will take reasonable technical and management measures to prevent user data from being leaked, tampered with, damaged, or lost. The providers of the terminal 120 and server 110 according to the embodiments of this disclosure will regularly review and update user data, and promptly delete expired or useless data. Furthermore, cloud service providers using the embodiments of this disclosure respect users' rights to data access, correction, deletion, withdrawal of consent, complaints, and claims, and provide convenient channels and procedures to enable users to effectively exercise these rights.
[0040] Furthermore, the process of data analysis using artificial intelligence technology in terminal 120 or server 110 is conducted based on the principles of legality, rationality, and transparency. The data collected and processed by the artificial intelligence model according to embodiments of this disclosure is relevant, necessary, and appropriate for the predictive purpose, and does not contain any personally identifiable or sensitive information. The neural network model according to embodiments of this disclosure employs appropriate techniques and organizational measures to protect the security and integrity of data, preventing unauthorized access, use, or disclosure.
[0041] The artificial intelligence-based neural network model according to embodiments of this disclosure will comply with relevant data protection regulations and ethical principles. This neural network model is trained on a large amount of anonymized and de-identified data, and does not infringe on the privacy rights of any individual or group. The artificial intelligence model has also undergone rigorous testing and evaluation to ensure that its output results are accurate and reliable, and will not cause any misleading or discriminatory results. The artificial intelligence model is designed solely to improve service quality and customer satisfaction and will not be used for any illegal or unethical purposes. Furthermore, the neural network model will be regularly reviewed and updated to adapt to changes in the data environment and legal regulations.
[0042] Traditionally, in the context of image-based virtual try-on, the goal is to generate a final image of a person wearing a specific garment, given an initial image of the person. However, digital models in traditional image-based virtual try-on technologies often only show the effect of the clothing with their arms or hands close to their sides, which greatly limits the flexibility of movement of the digital models and makes their posture appear stiff.
[0043] Generating images of digital models wearing clothing with arbitrary arm and hand poses can lead to distortion of the clothing around the arms or hands. Two solutions have been proposed to address this issue: First, refine the analytical model of the arm or hand to accurately segment the tiny regions between the fingers. However, existing analytical models struggle to accurately segment small and dispersed finger regions, and adjusting the arm or hand using the analytical model may introduce residual artifacts and background from the model's image. Second, integrate the reconstruction of the hand region image into the fitting process, i.e., decompose the hand into appearance and structural knowledge during the fitting stage. This includes methods such as generative adversarial networks and diffusion models.
[0044] Schemes using generative adversarial networks typically involve initially deforming clothing to match the shape of the human body, and then using a generator to overlay the deformed clothing onto the human image. Although such schemes strive to reduce the differences between the deformed clothing and the human body, they are generally difficult to apply broadly to various human images, especially those with complex backgrounds or poses.
[0045] The diffusion model approach attempts to preserve clothing details and deform the garments to fit the fitting process. This approach treats virtual try-on as an image inpainting problem, fine-tuning the diffusion model using a virtual try-on dataset to generate higher-quality virtual try-on images. While the diffusion model approach addresses the naturalness issue of the synthesized images, it still faces challenges in handling the widespread hand-based occlusion problem in real-world virtual try-on scenarios.
[0046] Therefore, existing virtual try-on technology needs further improvement.
[0047] This disclosure provides a method for updating a person image, comprising: obtaining a clothing image and an initial person image, wherein the clothing image includes a preset clothing, and the initial person image includes a preset person wearing the current clothing, wherein the area where the current clothing is located in the initial person image is occluded by the hand of the preset person; determining clothing prior information of the preset clothing based on the clothing image; determining reference features corresponding to other areas outside the area where the current clothing is located, as well as visual prior information and morphological prior information corresponding to the hand of the preset person based on the initial person image; and iteratively updating the initial person image based on the clothing prior information, the reference features, and the visual and morphological prior information to generate a final person image, wherein the final clothing worn by the preset person in the final person image corresponds to the preset clothing, and the area where the final clothing is located in the final person image is occluded by a corresponding hand.
[0048] Compared to traditional virtual try-on technology, the embodiments of this disclosure can utilize prior visual and morphological information corresponding to the hands of a preset person to generate a final image of the preset person wearing preset clothing with hands occluding the clothing area by iteratively updating the initial person image. The naturalness and realism of the final person image generated using the embodiments of this disclosure are significantly higher than those of traditional virtual try-on schemes.
[0049] Furthermore, this disclosure also provides a method for training a neural network to train an improved neural network model. The neural network model includes a hand feature separation and embedding module, a hand pose aggregation module, a try-on module, and a clothing module. The method includes: obtaining clothing image samples and initial person image samples, wherein the clothing image samples include preset clothing, and the initial person image samples include a preset person wearing the current clothing, wherein the area where the current clothing is located in the initial person image samples is occluded by the hands of the preset person; based on the clothing image samples, using the clothing module, determining predicted values of the clothing prior information of the preset clothing; based on the initial person image samples, determining predicted values of reference features corresponding to areas outside the area where the current clothing is located; and based on the initial person image samples, using the hand feature separation and embedding module and the hand pose aggregation module, determining the visual prior information corresponding to the hands of the preset person. The system uses the try-on module to iteratively update the initial character image samples to generate a predicted image of the final character image. The predicted image is based on the predicted values of the clothing prior information, the reference features, and the visual and morphological prior information corresponding to the hands of the preset character. In the predicted image of the final character image, the final clothing worn by the preset character corresponds to the preset clothing, and the area where the final clothing is located in the predicted image of the final character image has corresponding hand occlusion. In each iteration, based on the difference between the predicted image of the final character image determined in that iteration and the final character image samples, the hand feature separation and embedding module, the hand pose aggregation module, the try-on module, and the clothing module are trained.
[0050] In the training process of the improved neural network model, the embodiments of this disclosure employ a two-stage training process, which can avoid overfitting during training and maintain the original pre-trained generative capabilities of each sub-module. Simultaneously, the hand edge loss function enables the hand feature separation and embedding module and the hand pose aggregation module to learn how to generate accurate prior information related to the hand, and enables the try-on module and the clothing module to learn how to better integrate this prior information into the final image generation process.
[0051] The following describes, with reference to the accompanying drawings, a method for generating a final human image, a method for training a neural network, and an apparatus for doing so according to embodiments of the present disclosure.
[0052] Figure 2 A schematic diagram of an application scenario according to an embodiment of this disclosure is shown.
[0053] This disclosed embodiment can be integrated into any virtual try-on app, allowing users to experience a one-click clothing try-on effect. For example... Figure 2 As shown, in actual operation of this virtual try-on app, users can upload a full-body photo of themselves wearing a simple white T-shirt as the initial person image, and select a T-shirt with a brand logo from the app's store. The flat lay image of this T-shirt with the brand logo will then serve as the garment image.
[0054] The virtual try-on app extracts prior information about the clothing from the clothing image, and then extracts reference features corresponding to the preset person from the initial person image, as well as visual and morphological prior information about the preset person's hands. Next, based on the clothing prior information, the reference features corresponding to the preset person, and the visual and morphological prior information about the preset person's hands, the virtual try-on app determines the final person image. In the final person image, the clothing worn by the preset person is a preset garment, and the area of the preset garment is partially obscured by hands.
[0055] The virtual try-on app will present a final image of the user, showcasing a virtual effect of the user wearing a T-shirt with a logo. The fabric material, color, and logo rendering are extremely realistic. The interaction between the hand and the clothing is natural and smooth, with realistic deformation of the clothing and soft shadows where fingers touch the fabric visible near the user's hands. This disclosed embodiment makes the entire virtual try-on process more convincing.
[0056] Figure 3 A flowchart of a method 30 for generating a final human image according to an embodiment of the present disclosure is shown.
[0057] like Figure 3 As shown, method 30 can be used on a terminal device or server (such as...) Figure 1 The method is executed at the terminal 120 or server 110. Method 30 includes the following operations S300 to S303. Of course, method 30 may also include more or fewer operations, and this disclosure is not limited thereto.
[0058] In operation S300, a clothing image and an initial character image are obtained. The clothing image includes a preset clothing, and the initial character image includes a preset character wearing the current clothing. In the initial character image, the area where the current clothing is located is obscured by the hand of the preset character.
[0059] Optionally, the clothing image can be as follows: Figure 2The clothing image shown depicts a preset outfit that a character in the final character image to be generated will wear. It is important to note that the current outfit worn by the initial character can be a single piece of clothing among multiple outfits worn by the initial character, and is not limited to all outfits worn by the initial character. For example, as shown in the attached image… Figure 2 As shown, the scope of the current clothing mainly includes tops but excludes trousers. In other cases, the current clothing may refer only to trousers, depending on the specific circumstances. Similarly, the preset clothing can also be a single piece of clothing; it is not required that the preset clothing be a complete set. Of course, this disclosure is not limited to this.
[0060] In operation S301, based on the clothing image, the prior information of the preset clothing is determined.
[0061] In the context of this disclosure, prior information refers to preliminary information that can be referenced based on existing experience, knowledge, or data in a specific task or domain. Optionally, clothing prior information can be a feature representation corresponding to a pre-defined garment extracted using any artificial intelligence model. Specifically, clothing prior information is a type of prior information related to the characteristics of a pre-defined garment, typically including characteristic parameters such as the garment's material, style, color, texture, and morphological structure, as well as constraints related to the garment's function or design. For example, the data that clothing prior information may include might involve: the fiber composition ratio of the fabric, fabric density, color space parameters, fabric reflectance distribution, texture directionality index, cut structural dimensions (such as sleeve length, neckline, etc.), dynamic wrinkle coefficient, and garment comfort index, etc. Of course, this disclosure is not limited to this.
[0062] In operation S302, based on the initial character image, reference features corresponding to other areas outside the area where the current clothing is located, as well as visual prior information and morphological prior information corresponding to the hand of the preset character are determined.
[0063] Optionally, the reference features corresponding to the area outside the current clothing location are independent of whether the preset character is wearing the current clothing or the preset clothing. These features indicate general characteristics of the area outside the current clothing location, such as the preset character's body shape, facial features, and background, and are universal in the latent space. That is, the reference features are feature vectors that remain consistent regardless of whether the target character is wearing different clothing. The visual prior information corresponding to the preset character's hands refers to information related to the appearance characteristics of the preset character's hands in the image, such as shape, structure, size, and color. The morphological prior information corresponding to the preset character's hands refers to information related to the morphological characteristics of the preset character's hands in the image, such as shape characteristics, posture characteristics, joint angles, and positional relationships. Of course, this disclosure is not limited to these.
[0064] In operation S303, the initial character image is iteratively updated based on the clothing prior information, the reference features, the visual prior information, and the morphological prior information to generate a final character image. In the final character image, the final clothing worn by the preset character corresponds to the preset clothing, and the area where the final clothing is located in the final character image has a corresponding hand occlusion.
[0065] Optionally, the iterative update process of the initial character image is a denoising process that utilizes "prior information about clothing, reference features corresponding to the preset character, and visual and morphological prior information about the preset character's hands." This process involves multiple iterations to determine the final character image. Specifically, the area where the original clothing was located in the initial character image can be considered a noise area. In each iterative update process, one or more of the aforementioned feature information can be used to correct the noise area in the image, gradually removing the influence of the original clothing (current clothing) and replacing it with the preset clothing.
[0066] Through multiple iterations, this denoising process is continuously refined under the control of "the prior information of the clothing, the reference features corresponding to the preset character, and the visual and morphological prior information corresponding to the hands of the preset character," until the area where the original clothing is located is completely replaced with the preset clothing, while keeping other features of the preset character (such as the face, posture, and hand areas of the target task) unchanged. Of course, this disclosure is not limited to this.
[0067] Optionally, such as Figure 2 As shown, the area obscured by the hand in the final character image is almost identical in position and detail to the area obscured by the hand in the current clothing area of the initial character image, particularly in the shape and detail of the hand. Within the final clothing area, the area outside the hand-obscured region is updated accordingly.
[0068] Therefore, according to the embodiments of this disclosure, a completely new final character image can be determined, which shows a preset character wearing preset clothing, and the hand features, posture, and details of the preset clothing are accurately preserved.
[0069] Next, refer to Figure 4 An example of iterative updates in operation S303 is further described. Figure 4 A schematic diagram of a try-on module according to an embodiment of the present disclosure is shown. Figure 4 This is merely one example of iterative updates, and this disclosure is not limited thereto.
[0070] Optionally, operation S303 includes: in the first iteration, performing the following steps: concatenating the initial character image with the reference features to obtain the input features of the first iteration; encoding the input features of the first iteration based on the visual prior information and the clothing prior information to obtain an encoded feature vector; and decoding the encoded feature vector of the first iteration based on the visual prior information, the morphological prior information, and the clothing prior information to obtain the reconstructed image of the first iteration.
[0071] Optionally, operation S303 includes: concatenating the reconstructed image from the previous iteration with the reference features to obtain input features; encoding the input features based on the visual prior information and the clothing prior information to obtain an encoded feature vector; and decoding the encoded feature vector based on the visual prior information, the morphological prior information, and the clothing prior information to obtain the reconstructed image for the current iteration.
[0072] In this process, the reconstructed image obtained in each iteration is the predicted image of the final person image corresponding to that iteration, and the reconstructed image obtained in the last iteration is the final person image. Of course, this disclosure is not limited to this.
[0073] refer to Figure 4 Let's assume we are currently in the t-th iteration. The reconstructed image from the previous iteration is shown as z. t-1 The reference feature is subsequently identified as ε(x) m ). Optionally, such as Figure 4 As shown, the masked image x can be obtained by masking the clothing area in the initial character image. m ; and extract the features ε(x) corresponding to the image after the mask is applied. m This refers to the reference features corresponding to the area outside the current clothing location. Specifically, the initial person image can be subdivided into different regions (such as head, arms, torso, etc.) using a Human Parsing model, and the clothing area can be clearly identified. Then, the key points of the preset person in the initial person image are located using a Human Pose estimation model to determine the preset person's posture and movement. Finally, based on the output of the above two models, a mask can be constructed to cover the clothing area, thereby generating an image x with clothing information removed. m Of course, this disclosure is not limited to this.
[0074] The reconstructed image z from the previous iteration t-1 With the reference feature ε(x) m The input features for the current iteration are obtained by concatenating these features. After this iteration, the output image z can be obtained. tIf this is the first iteration, the input features can be the initial image of the person and the reference features ε(x). m (The splicing of ) is not limited to this. Of course, this disclosure is not limited to this.
[0075] Optionally, such as Figure 4 As shown, in each iteration, the denoising of the reconstructed image generated in the previous iteration can be achieved by calling the TryonNET module. Specifically, the TryonNET module can be a neural network model with a U-NET structure. U-Net is an efficient convolutional neural network whose architecture resembles the English letter "U," hence the name U-NET. This architecture combines multiple stacked encoders (also known as downsampling paths) and corresponding stacked decoders (or upsampling paths).
[0076] During the encoding phase, stacked encoders progressively extract higher-dimensional features from the input while gradually reducing spatial dimensions, thereby capturing the global contextual information of the input image. Each encoder level increases the number of channels in the feature map through convolution operations, thus extracting more abstract and complex feature representations. Simultaneously, each encoder level can optionally use visual prior information and clothing prior information as control information to achieve a controlled downsampling process.
[0077] Optionally, the visual prior information can be multi-resolution information, thus adapting to different encoder levels. Similarly, the clothing prior information can also be multi-resolution information. It is worth noting that, in this disclosure, multi-resolution information or multi-resolution features refer to information with a specific resolution for different encoder levels, while single-resolution information or single-resolution features refer to information with the same resolution for different encoder levels. Of course, this disclosure is not limited thereto.
[0078] Furthermore, both visual and clothing prior information can be single-resolution control information, meaning the input dimension of the prior information for each encoder is the same. Simultaneously, both visual and clothing prior information can include both single-resolution and multi-resolution information, thus better adapting to the control downsampling process. Of course, this disclosure is not limited to these limitations.
[0079] During the decoding stage, the stacked decoders gradually restore the spatial resolution of the image, bringing it close to the size of the original input image. In this process, each decoder layer not only receives feature maps from the previous layer but also implements a controlled upsampling process using pre-defined visual prior information, morphological prior information, and clothing prior information corresponding to the person's hands. Similarly, the morphological prior information can be single-resolution information or multi-resolution information; or it can include a portion of single-resolution information as well as a portion of multi-resolution information. Of course, this disclosure is not limited to these limitations.
[0080] like Figure 4 As shown, after a single iteration, the details of the clothing are enhanced, noise is reduced, and the visual appearance and posture of the hands remain unchanged; in fact, the details of the hands become clearer. Therefore, after multiple iterations of updating the initial character image in operation S303, the task objective—determining the final character image—can be achieved. In the final character image, the clothing of the preset character is a preset garment, and the area of the preset garment contains hand occlusion. Of course, this disclosure is not limited to this.
[0081] Next, refer to Figure 5 Further examples describing prior information about clothing. Figure 5 A schematic diagram illustrating the determination of prior information for clothing according to an embodiment of the present disclosure is shown.
[0082] Optionally, such as Figure 5 As shown, the clothing image g is a flat lay image of the preset clothing, designed to showcase the core elements of the preset clothing, such as style, color, and texture, to facilitate the subsequent generation of the final character image. Of course, this disclosure is not limited to this.
[0083] Optionally, operation S301 may further include: determining at least one of single-resolution clothing features and multi-resolution clothing features based on the clothing image; and determining the clothing prior information based on at least one of the single-resolution clothing features and multi-resolution clothing features. The single-resolution clothing features are the same for each encoder and decoder layer in the try-on module, while the multi-resolution clothing features are different for each encoder and decoder layer in the try-on module. Of course, this disclosure is not limited thereto.
[0084] For example, through Figure 5 The illustrated scheme determines both single-resolution and multi-resolution clothing features as prior information. Those skilled in the art should understand that this disclosure is not limited thereto. For example, any embodiment of this disclosure may use only single-resolution clothing features as prior information, or it may use only multi-resolution clothing features. Furthermore, the following schemes for extracting single-resolution and multi-resolution clothing features are merely examples, and those skilled in the art should understand that other schemes can also be used to extract these two types of information.
[0085] Optionally, the CLIP (Contrastive Language–Image Pre-training) model can be invoked to determine the single-resolution clothing features corresponding to the clothing image g as part of the clothing prior information. Specifically, the CLIP model associates the clothing image g with its corresponding label—the preset clothing—through contrastive learning, thereby obtaining semantic information such as the style, color, and texture of the preset clothing, and binding this information with the task objective (i.e., generating the final image of the person whose clothing is the preset clothing). The single-resolution clothing features can be input into each encoding layer and each decoding layer in the try-on module. Although the CLIP model is used to describe the process of determining the single-resolution clothing features corresponding to the clothing image g, those skilled in the art should understand that this disclosure is not limited thereto; for example, traditional convolutional neural networks can also be used to extract single-resolution clothing features.
[0086] Alternatively, one can use, such as Figure 5 The illustrated clothing network model determines multi-resolution clothing features corresponding to a clothing image g as part of the clothing prior information. Similar to the try-on module, the clothing network model can also adopt a U-NET architecture, which combines encoder and decoder paths to achieve multi-scale feature extraction of the clothing image and inputs features corresponding to different scales to the try-on module, thereby achieving feature adaptation at different scales. In the encoder stage, the clothing network model progressively captures different levels of features of the clothing image g from low-level details to high-level semantics through stacked encoders. These features are effectively encoded at different resolutions and subsequently decoded by stacked decoders. The encoded feature vectors of different resolutions output by the encoder in the clothing network model are input to the encoder of the corresponding resolution of the try-on module. Similarly, the decoder takes the opposite operation to the encoder, progressively restoring the spatial resolution of the image through upsampling. This multi-resolution decoded information is also input to the corresponding resolution decoder of the try-on module, thereby providing multi-resolution prior information to the decoder of the try-on module. Multi-resolution clothing features can be input to the corresponding resolution encoding or decoding layer in the try-on module. Of course, this disclosure is not limited thereto.
[0087] like Figure 5 As shown, some embodiments of this disclosure determine prior information about clothing based on clothing images using CLIP models and clothing network models. This approach not only preserves the detailed information of the clothing images but also leverages the CLIP model's understanding and representation of clothing features. Especially when dealing with complex and varied clothing styles and textures, the combination of single-resolution and multi-resolution features enables the try-on module to obtain more accurate key attributes of the preset clothing, such as style, color, and material.
[0088] Next, refer to Figure 6 and Figure 7 Further examples describing prior information about clothing. Figure 6 A schematic diagram is shown illustrating the determination of visual prior information and morphological prior information corresponding to the hand of a preset character according to an embodiment of the present disclosure. Figure 7 It shows Figure 6 A schematic diagram of the masked cross-attention model.
[0089] First, the process of obtaining prior morphological information corresponding to the hand of a preset character is introduced. This process can be briefly described as follows: based on the initial character image, determine the hand pose features and hand structure features; and based on the hand pose features and hand structure features, determine the prior morphological information corresponding to the hand of the preset character. Determining the hand pose features includes: based on the initial character image, determining at least one of skeletal structure information, whole-body semantic information, and hand depth information; and based on at least one of skeletal structure information, whole-body semantic information, and hand depth information, determining the hand pose features. Thus, the hand pose features integrate at least one of skeletal structure information, whole-body semantic information, and hand depth information. Of course, this disclosure is not limited to this.
[0090] Optionally, the skeletal structure information indicates the skeletal information of the preset character, especially how the various joints of the preset character are connected. The whole-body semantic information indicates how to map the positions of each pixel in the initial character image onto the 3D surface of the human body. The depth information of the hand indicates the positional and structural information of the hand in the initial character image. Of course, this disclosure is not limited thereto.
[0091] Optionally, a subset of the morphological prior information for determining the preset person—the hand pose features F—can be provided by invoking the Hand-Pose Aggregation Net. hp Specifically, the hand pose aggregation module can be divided into three decoupled sub-modules: a body skeleton model, a body semantic model, and a pixel-level hand model. Of course, those skilled in the art should understand that the hand pose aggregation module of this disclosure may also include more or fewer sub-modules, and this disclosure is not limited thereto.
[0092] Optionally, the example architecture of the body skeleton model can be similar to the DWpose model to extract skeletal structure information F from the initial human image. dw (skeleton structure), thereby providing prior information on the shape of the hands of the pre-defined character to estimate the main joints of the pre-defined character, as well as the pose parameters of the feet, face, and hands. Of course, this disclosure is not limited thereto.
[0093] Optionally, an example architecture for the body semantic model can be similar to the Densepose model for extracting full-body semantic information F from an initial image of a person. dp (Body-semantic information) provides pixel-level pose parameters for the pre-defined shape of a person's hands. Specifically, body-semantic information can associate each pixel in a 2D image with its position on the 3D surface of the human body, thereby predicting the surface coordinates and body part classification of the pixels. Of course, this disclosure is not limited thereto.
[0094] Optionally, the example architecture of the pixel-level hand model can be similar to the Depth model to extract pixel-level hand position and structural information from the initial person image, thereby providing depth information F of the hand for the pre-defined morphological prior information of the person's hand. rh Of course, this disclosure is not limited to this.
[0095] Optionally, the outputs of the body skeleton model, body semantic model, and pixel-level hand model can be fed into the same zero-multilayer structure to facilitate overfitting during training and to optimize the skeletal structure information F. dw Whole-body semantic information F dp and the depth information of the hand F rh The zero-multilayer structure consists of seven stacked convolutional layers (conv) and the SiLU (Sigmoid Linear Unit) activation function. The SiLU activation function, also known as the Swish activation function, is a novel activation function defined as the input multiplied by its sigmoid function value. This activation function is smooth, non-linear, and can adaptively scale the input value, helping the gradient descent algorithm to stably update parameters during optimization. The last layer of the zero-multilayer structure is a zero-convolutional layer, a 1x1 convolution with its weights and biases initialized to zero.
[0096] By aggregating the three outputs of the body skeleton model, the body semantic model, and the pixel-level hand model, the output F of the hand pose aggregation module is obtained. hp It can be represented as formula (2).
[0097] F hp =F dw +F dp +F rh *w hand (2)
[0098] Among them, w hand This indicates the information F used to adjust the hand position and structure. rhThe weight.
[0099] like Figure 6 As shown, the output F of the hand posture aggregation module hp The hand structure information output from the Hand-feature Disentanglement Embedding model can be fused with a decoder network (which has multiple stacked decoding layers) to obtain multi-resolution morphological prior information corresponding to the hand of the preset character.
[0100] Optionally, the hand feature separation and embedding module can also consist of multiple decoupled sub-modules, namely: a 3D hand reconstruction model, a hand structure processing module, a visual feature extraction module, and a visual feature processing module. Of course, the hand feature separation and embedding module can also include more or fewer modules, and this disclosure is not limited thereto.
[0101] Optionally, the example architecture of the 3D hand reconstruction model can be similar to the Hand Mesh Recovery (HaMeR) model, which can be used to reconstruct a 3D model of a preset person's hand from an initial person image. For example, the 3D hand reconstruction model can extract a set of hand structure parameters from the initial person image to describe the hand's pose and shape. These parameters include joint angles, global position, and shape parameters, etc. For instance, this set of parameters can include parameters that can be extracted by the MANO model (hand model with articulated and non-rigid deformations), including but not limited to: 3 camera parameters (rotation angles with x, y, and z), 45 joint pose parameters (excluding joint 0 and the fingertips, the other 15 joints each have 3 parameters), and 10 shape parameters (used to determine the shape of the hand, such as finger length and palm width), etc. This set of parameters can also include parameters related to hand information, but this disclosure is not limited thereto.
[0102] Assume that the hand structure parameters extracted from the 3D hand reconstruction model include at least one of the following parameters: hand mask. Hand type Hand 3D Vertex Spatial joint position and joint rotation matrix Where n represents the number of hands.
[0103] Optionally, the initial character image may show only one hand or both hands. To handle different numbers of hand parameters, the hand parameters can be padded into a fixed length n=2. For hand type T hThe parameter uses 0 to represent the left hand, 1 to represent the right hand, and -1 as a padding value for images without a hand, to distinguish between different hand types. To ensure that other parameters are consistent with the hand type T... h The alignment and correspondence of parameters were determined, and a similar padding strategy was applied to other parameters as well.
[0104] Optionally, for the 3D vertices of the hand Spatial joint position and joint rotation matrix These three hand parameters can also be reduced in dimensionality using the Basic Point Set (BPS) scheme to reduce the data's dimensionality while retaining key information. For example, the joint rotation matrix... It can be encoded as a 6-dimensional vector.
[0105] Optionally, the hand structure parameters extracted from the 3D hand reconstruction model will be further processed by the hand structure processing module to generate hand structure features c. struc Optionally, the hand structure processing module consists of two parallel sub-modules: L r Modules and L n Modules. Of course, this disclosure is not limited to this.
[0106] Optionally, L r The full name of the module is Linear and ReLU activation module, which contains three stacked linear layers, each followed by a ReLU activation function. The ReLU activation function is used to introduce non-linear properties to accommodate complex hand structure features. r The module takes the hand vertices and joint positions as input and outputs a streamlined feature representation of the hand vertices and joint positions.
[0107] Optionally, L n The full name of the module is Linear and Layer Normalization, which contains two stacked linear layers, each followed by a normalization layer to improve the generalization ability of the output features. n The module takes joint rotation and hand type as input and outputs feature representations of joint rotation and hand type.
[0108] Optionally, from L r Modules and L n The module's output is aggregated, and the hand structure features c struc Therefore, the processing procedure of the hand structure processing module can be summarized by formula (3) as follows.
[0109] c struc =[Lr (Vh),L r (J 2d ), L n (T h ),L n (θ h (3)
[0110] Optionally, hand structural features c struc Optionally, the hand structure features can then be enhanced using a masked cross-attention module to obtain the enhanced hand structure features c. struc The following will be referenced. Figure 7 Further details on the mechanism of the masked cross-attention module will not be elaborated here. It is worth noting that the masked cross-attention module is optional; even without using it, the hand's structural features c... struc Enhancement can also be achieved by directly using the hand structure features generated by the hand structure processing module. struc As the output of the hand structure information and hand posture aggregation module, F hp The data is fused to obtain multi-resolution morphological prior information corresponding to the hand of the preset character. Of course, this disclosure is not limited to this.
[0111] Next, the process of obtaining the visual prior information corresponding to the hands of the preset character is described. This process can be summarized as follows: based on the initial character image, determine the hand image within the initial character image; based on the hand image within the initial character image, determine the initial visual features corresponding to the hands of the preset character; and perform linear transformation and normalization operations on the initial visual features corresponding to the hands of the preset character to determine the hand visual features, which will serve as the visual prior information corresponding to the hands of the preset character. Of course, this disclosure is not limited to this.
[0112] Optionally, the bounding box region of the hand can first be delineated in the initial image of the person using any object detection model, thereby extracting, for example, the bounding box region of the hand. Figure 6 The image shown is of the hand in the initial image of the person. Next, the visual feature extraction module can be used to extract initial visual features F∈R from the hand image. n×1536 The architecture of the visual feature extraction module can be similar to the DINOv2 (Dual-Stage Implicit Object-Oriented Network version 2) model, or a pre-trained, publicly available DINOv2 model can be directly adopted. Specifically, the DINOv2 model is a pre-trained large visual transformer (ViT) model that efficiently captures key information in hand images, enhances visual details of the hand, and generates a high-dimensional feature representation F∈R. n×1536 .
[0113] Optionally, the initial visual features F∈R n×1536 The hand will be further processed by the Hand-Appearprocessor (HA) module to generate hand visual features. appear The hand vision processing module optionally includes linear layers and normalization layers. Thus, through the hand vision processing module, the visual fidelity and detail of the hand are improved in the hand visual features c. appear The exact values are preserved. This process can be shown in Equation (4) as follows.
[0114]
[0115] Optionally, hand visual features c appear Optionally, the hand visual features c can then be enhanced using a masked cross-attention module to obtain the enhanced hand visual features. appear The following will be referenced. Figure 7 Further details on the mechanism of the masked cross-attention module will not be elaborated here. It is worth noting that the masked cross-attention module is optional; even if it is not used, it will not affect the hand's visual features c. appear Enhancement can also be achieved by directly using the hand visual features generated by the hand visual processing module. appear This serves as the visual prior information corresponding to the hands of the pre-defined character. Of course, this disclosure is not limited to this.
[0116] Next, through Figure 7 To further describe the working mechanism of the masked cross-attention module.
[0117] like Figure 7 As shown, mask-based region control can be achieved through a masked cross-attention module to enhance hand detail generation and maintain quality in virtual try-on. The working principle of the masked cross-attention module can be represented as follows:
[0118]
[0119] Where Q, K, and V are the query matrix, key matrix, and value matrix of the masked cross-attention operation, respectively, and M... h This represents the hand mask extracted from the preceding sequence. The hand mask M is then used as the parameter extracted from the preceding sequence. h It can adjust the cross-attention query vector to achieve hand visual feature c appear Or hand structural features c struc Adjusting attention.
[0120] For example, when generating morphological prior information, the hand structure feature c generated by the hand structure processing modulestruc The key and value matrices that can be used as masks for cross-attention operations, such as the hand mask M. h The output F of the hand posture aggregation module hp The product of the corresponding latent space vectors can be used as the query matrix for the cross-attention operation. Then, the cross-attention module integrates this query matrix, key matrix, and value matrix to obtain the enhanced hand structure features c. struc As prior information about form.
[0121] For example, when generating visual prior information, the hand visual features c generated by the hand visual processing module appear The key and value matrices that can be used as masks for cross-attention operations, such as the hand mask M. h The output z of the try-on module t The product of the corresponding latent space vectors can be used as the query matrix for the cross-attention operation. Then, the cross-attention module integrates this query matrix, key matrix, and value matrix to obtain the enhanced hand visual features c. appear As visual prior information.
[0122] It is worth noting that, although the output F of the hand gesture aggregation module is used here... hp The corresponding latent space vector and the output z of the try-on module t The corresponding latent space vector is used as an example to describe the generation process of the query matrix, but this disclosure is not limited thereto. In some cases, even the hand mask M can be used alone. h As a query matrix, or as a hand mask M h The output F of the hand posture aggregation module hp (or the output z of the try-on module) t The product of these products is used as the query matrix. This disclosure is not limited thereto.
[0123] Figure 8 A schematic diagram of an apparatus 80 for implementing method 30 according to an embodiment of the present disclosure is shown. Although for ease of description, as Figure 8 As shown, device 80 covers Figures 4 to 7 All modules mentioned above, but as stated above, may be optional, provided that device 80 can implement method 30.
[0124] like Figure 8As shown, the device 80 comprises at least four decoupled modules: a hand feature separation and embedding module, a hand pose aggregation module, a try-on module, and a clothing module. Furthermore, the device 80 may also include an acquisition module (not shown) and a masking module (not shown). The acquisition module is configured to acquire a clothing image and an initial person image, wherein the clothing image includes a preset clothing, and the initial person image includes a preset person wearing the current clothing, wherein the area where the current clothing is located in the initial person image is occluded by the hand of the preset person. The masking module is configured to determine reference features corresponding to the area outside the area where the current clothing is located, based on the initial person image.
[0125] The clothing module is configured to determine the prior information of the preset clothing based on the clothing image. The hand feature separation and embedding module and the hand pose aggregation module are configured to jointly determine the morphological prior information corresponding to the hands of the preset person based on the initial person image. In addition, the hand feature separation and embedding module is also used to determine the visual prior information corresponding to the hands of the preset person based on the initial person image.
[0126] The try-on module can iteratively update the initial character image based on the clothing prior information, the reference features, the visual prior information, and the morphological prior information to generate a final character image. In the final character image, the final clothing worn by the preset character corresponds to the preset clothing, and the area where the final clothing is located in the final character image has a corresponding hand occlusion.
[0127] Already referenced Figures 4 to 7 The working principles of these four modules have been described separately, and will not be repeated here.
[0128] Figure 9 A flowchart of a method 90 for training a neural network model according to an embodiment of the present disclosure is shown, which can train a neural network model. Figure 8 The neural network model shown is trained and includes: a hand feature separation and embedding module, a hand pose aggregation module, a try-on module, and a clothing module. Specifically, during inference, the neural network executes method 30. Method 90 can be used on a terminal device or server (such as...). Figure 1 The method is executed at the terminal 120 or server 110. Method 90 includes the following operations S901 to S905. Of course, method 90 may also include more or fewer operations, and this disclosure is not limited thereto.
[0129] Optionally, before performing the training process of method 90, Figure 8 Some sub-modules in the neural network model shown may have been pre-trained. For example, the try-on module and the clothing module could both have been pre-trained.
[0130] In an optional embodiment, the pre-training process of the try-on module and the clothing module can be summarized as follows: assuming that the visual and morphological prior information corresponding to the hands of the preset person does not contain arbitrary information (e.g., unit vectors or zero vectors), the try-on module and the clothing module are pre-trained using pairs of initial person image samples and clothing image samples. Specifically, the clothing region in the initial person image samples can be masked to obtain a masked image, and then the predicted value of the reference feature can be obtained. The masked image is then concatenated with the reference feature to obtain the initial values for training the try-on module and the clothing module. Next, the masked image is iteratively updated to determine the predicted value of the final person image sample. In each iteration, the predicted value of the final person image can be compared with the true value of the final person image (i.e., the initial person image sample), and the parameters of each neural network in the try-on module and the clothing module are adjusted based on the difference between the two.
[0131] In some alternative embodiments, pairs of initial person image samples, clothing image samples, and final person image samples can be used as training samples. In this case, the pre-training process of the try-on module and clothing module can be summarized as follows: assuming that the visual prior information and morphological prior information corresponding to the hands of the preset person do not contain arbitrary information (e.g., unit vectors or zero vectors), pairs of initial person image samples and clothing image samples are used to pre-train the try-on module and clothing module. Specifically, the clothing region in the initial person image sample can be masked to obtain a masked image, and then the predicted value of the reference feature can be obtained. Then, the initial person image sample and the reference feature are concatenated to obtain the initial values for training the try-on module and clothing module. Next, the masked image is iteratively updated to determine the predicted value of the final person image sample. In each iteration, the predicted value of the final person image can be compared with the true value of the final person image (i.e., the final person image sample), and the parameters of each neural network in the try-on module and clothing module are adjusted based on the difference between the two. Of course, other methods can also be used to train the try-on module and clothing module, and this disclosure is not limited thereto.
[0132] Furthermore, some sub-modules in the hand feature separation and embedding module and the hand pose aggregation module can also be pre-trained, such as the body skeleton model, body semantic model, and pixel-level hand model in the hand pose aggregation module. These models can directly use pre-trained, permitted open-source models. Similarly, the 3D hand reconstruction model and visual feature extraction module in the hand feature separation and embedding module can also be pre-trained. This disclosure is not limited thereto.
[0133] Optionally, the pre-trained fitting model method 90 can be used to perform a two-stage training process. In the first stage, the untrained sub-modules in the hand feature separation and embedding module and the hand pose aggregation module can be trained.
[0134] In the first stage, the neural network parameters in the try-on module and clothing module remain unchanged. Specifically, if the 3D hand reconstruction model and visual feature extraction module in the hand feature separation and embedding module have been pre-trained, the hand structure processing module, visual feature processing module, and mask cross-attention module in the hand feature separation and embedding module can be trained individually or jointly. Similarly, the decoder network and zero-multilayer structure in the hand pose aggregation module can also be trained individually or jointly in the first stage.
[0135] In the second stage, the neural network parameters of the hand feature separation and embedding module and the hand pose aggregation module remain unchanged; only the neural network parameters in the dressing module and the clothing module are fine-tuned. Method 90 can be used in both the first and second stages to... Figure 8 The neural network parameters in all or some modules of the software are trained. Of course, this disclosure is not limited to this.
[0136] Next, refer to Figure 9 Method 90 is described further.
[0137] In operation S900, a clothing image sample and an initial person image sample are obtained. The clothing image sample includes a preset clothing, and the initial person image sample includes a preset person wearing the current clothing. In the initial person image sample, the area where the current clothing is located is partially obscured by the preset person's hand. Optionally, the clothing image sample and a reference... Figure 3 The clothing images described are similar, but this disclosure is not limited thereto. Optionally, the initial human figure image sample is similar to the reference. Figure 3 The initial human figures described are similar, the only difference being that during training, the clothing image samples and the initial human figure image samples usually have a pre-defined matching relationship. Of course, this disclosure is not limited to this.
[0138] In operation S901, the clothing image sample is used to determine the predicted value of the clothing prior information of the preset clothing using the clothing module.
[0139] In operation S902, the initial human image sample determines the predicted value of the reference feature corresponding to the area outside the area where the current clothing is located.
[0140] In operation S903, based on the initial human image sample, the predicted values of the visual prior information and the morphological prior information corresponding to the hand of the preset human are determined by using the hand feature separation and embedding module and the hand pose aggregation module.
[0141] In operation S904, based on the predicted values of the clothing prior information, the predicted values of the reference features, and the predicted values of the visual prior information and morphological prior information corresponding to the hands of the preset person, the initial person image sample is iteratively updated using the try-on module to generate a predicted image of the final person image. In the predicted image of the final person image, the final clothing worn by the preset person corresponds to the preset clothing, and the area where the final clothing is located in the predicted image of the final person image has corresponding hand occlusion. Optionally, in each iteration, based on the difference between the predicted image of the final person image determined in that iteration and the final person image sample, the hand feature separation and embedding module, the hand pose aggregation module, the try-on module, and the clothing module are trained.
[0142] Optionally, in each iteration, the step of training the hand feature separation and embedding module, the hand pose aggregation module, the try-on module, and the clothing module by analyzing the difference between the predicted image of the final person image determined in the current iteration and the final person image sample includes: in a first stage, training the hand feature separation and embedding module and the hand pose aggregation module while keeping the neural network parameters of the try-on module and the clothing module unchanged; and in a second stage, fine-tuning the neural network parameters of the try-on module and the clothing module while keeping the neural network parameters of the hand feature separation and embedding module and the hand pose aggregation module unchanged.
[0143] Optionally, in each iteration, the step of training the hand feature separation and embedding module, the hand pose aggregation module, the try-on module, and the clothing module by determining the difference between the predicted image of the final person image determined in the current iteration and the predicted image of the final person image sample in the previous iteration on the hand edges; and training at least one of the hand feature separation and embedding module, the hand pose aggregation module, the try-on module, and the clothing module based on the difference. This disclosure is not limited thereto.
[0144] Optionally, in each iteration of the first stage, the hand edge loss function L can be used. handThe untrained sub-modules in the hand feature separation and embedding module and the hand pose aggregation module are trained to improve the clarity and accuracy of hand edges. The hand edge loss function L... hand This can be used to indicate the difference in hand edges between the predicted image corresponding to the final person image sample in the current iteration and the predicted image corresponding to the final person image sample in the previous iteration. Optionally, the hand edge loss function L... hand Based on the Canny edge detection operator, edge features of the hand region can be extracted and enhanced from the predicted image corresponding to the final human image sample generated in each iteration. The extracted edge features are then compared with the hand edge features in the real image using the MSE loss function, thereby achieving precise constraint and optimization of the hand shape and edges in the final human image sample.
[0145] For example, for a small value R that is less than a certain threshold T t This alternative embodiment directly uses a variational autoencoder (VAE) for decoding, from z t Before predicting z0, a noise reduction process is performed to sample x0. This process can be expressed as Equation (6).
[0146]
[0147] Subsequently, from the generated image z t The hand region is extracted through bounding box cropping, and then processed using the Canny edge detection operator. The processed hand region is then compared with the corresponding hand region in the ground truth image using the Mean Squared Error (MSE) loss function. Therefore, the hand edge loss function L... hand It can be represented as formula (7).
[0148]
[0149] The loss function L at the hand edge hand In this context, N represents the number of images. This represents the i-th generated image. Let represent the i-th real image. The function φ(·) represents applying the Canny edge detection operator to the image.
[0150] In each iteration of the second stage, the neural network parameters of the hand feature separation and embedding module and the hand pose aggregation module remain unchanged, while the neural network parameters of the try-on module and the clothing module are fine-tuned. In this stage, the try-on module and the clothing module already have relatively accurate visual and morphological prior information as references. The try-on module and the clothing module need to learn how to better generate the final portrait image based on this prior information. It is worth noting that the hand edge loss function L... hand It can also be applied in the second stage to measure whether the try-on module and the clothing module correctly apply visual prior information and morphological prior information, and whether they can accurately generate details related to the hand area. Of course, this disclosure is not limited thereto.
[0151] Furthermore, in both the first and second phases, the difference between the predicted image corresponding to the final person image sample and the final person image sample can be used during training to align the training objective with the task objective. When the difference between the predicted image corresponding to the final person image sample and the final person image sample is sufficiently small, the training phase can be considered complete.
[0152] This two-stage training process avoids overfitting during training while maintaining the original generative capabilities of each pre-trained sub-module. Simultaneously, the hand edge loss function L... hand It also enables the hand feature separation and embedding module and the hand pose aggregation module to learn how to generate accurate prior information related to the hands, and enables the try-on module and the clothing module to learn how to better integrate this prior information into the final image generation process. Of course, this disclosure is not limited thereto.
[0153] According to another aspect of this disclosure, an electronic device is also provided for implementing the methods according to embodiments of this disclosure. Figure 10 A schematic diagram of an electronic device 2000 according to an embodiment of the present disclosure is shown.
[0154] like Figure 10 As shown, the electronic device 2000 may include one or more processors 2010 and one or more memories 2020. The memories 2020 store computer-readable code that, when executed by the one or more processors 2010, can perform the methods described above.
[0155] The processor in this disclosure embodiment can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, operations, and logic block diagrams disclosed in this disclosure embodiment. The general-purpose processor can be a microprocessor or any conventional processor, and can be based on an x86 architecture or an ARM architecture.
[0156] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0157] For example, the method or apparatus according to embodiments of this disclosure can also be used by means of Figure 11 The architecture of the computing device 3000 shown is used for implementation. For example... Figure 11 As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage devices in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the methods provided in this disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 11 The architecture shown is merely exemplary and can be omitted as needed when implementing different devices. Figure 11 One or more components in the computing device shown.
[0158] According to another aspect of this disclosure, a computer-readable storage medium is also provided. Figure 12 A schematic diagram of a storage medium 4000 according to the present disclosure is shown.
[0159] like Figure 12As shown, the computer storage medium 4020 stores computer-readable instructions 4010. When the computer-readable instructions 4010 are executed by a processor, the methods according to embodiments of the present disclosure described with reference to the above figures can be performed. The computer-readable storage medium in the embodiments of the present disclosure may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory used in the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0160] This disclosure also provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a method according to an embodiment of this disclosure.
[0161] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0162] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0163] The exemplary embodiments of this disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art will understand that various modifications and combinations can be made to these embodiments or their features, and such modifications should fall within the scope of this disclosure. For example, the following provides an overview of some aspects of this disclosure, which can be combined with any other aspects.
[0164] Aspect 1: A method for updating a person image, characterized in that the method includes: obtaining a clothing image and an initial person image, the clothing image including a preset clothing, the initial person image including a preset person wearing the current clothing, wherein the area where the current clothing is located in the initial person image is occluded by the hand of the preset person; determining clothing prior information of the preset clothing based on the clothing image; determining reference features corresponding to other areas outside the area where the current clothing is located, as well as visual prior information and morphological prior information corresponding to the hand of the preset person based on the initial person image; and iteratively updating the initial person image based on the clothing prior information, the reference features, the visual prior information, and the morphological prior information to generate a final person image, wherein the final clothing worn by the preset person corresponds to the preset clothing in the final person image, and the area where the final clothing is located in the final person image is occluded by a corresponding hand.
[0165] Aspect 2: The iterative update of the initial character image includes: in the first iteration, performing the following steps: concatenating the initial character image with the reference features to obtain the input features of the first iteration; encoding the input features of the first iteration based on the visual prior information and the clothing prior information to obtain an encoded feature vector; and decoding the encoded feature vector of the first iteration based on the visual prior information, the morphological prior information, and the clothing prior information to obtain the reconstructed image of the first iteration; in each subsequent iteration, performing the following steps: concatenating the reconstructed image of the previous iteration with the reference features to obtain the input features; encoding the input features based on the visual prior information and the clothing prior information to obtain an encoded feature vector; and decoding the encoded feature vector based on the visual prior information, the morphological prior information, and the clothing prior information to obtain the reconstructed image of the current iteration. Wherein, the reconstructed image obtained in each iteration is the predicted image of the final character image corresponding to that iteration, and the reconstructed image obtained in the last iteration is the final character image.
[0166] Aspect 3: Determining the reference features based on the initial character image includes: masking the area where the current clothing is located in the initial character image to obtain a masked image; and determining the reference features corresponding to the area outside the area where the current clothing is located based on the masked image.
[0167] Aspect 4: The determination of prior information about clothing based on clothing images includes: determining at least one of single-resolution clothing features and multi-resolution clothing features based on the clothing images; and determining the prior information about clothing based on at least one of single-resolution clothing features and multi-resolution clothing features.
[0168] Aspect 5: Determining the morphological prior information corresponding to the hand of the preset character includes: determining hand posture features and hand structure features based on the initial character image; and determining the morphological prior information corresponding to the hand of the preset character based on the hand posture features and hand structure features. Optionally, determining the hand posture features includes: determining at least one of skeletal structure information, whole-body semantic information, and hand depth information based on the initial character image; and determining the hand posture features based on at least one of skeletal structure information, whole-body semantic information, and hand depth information.
[0169] Aspect 6: The determination of the visual prior information corresponding to the hand of the preset character includes: determining the hand image in the initial character image based on the initial character image; determining the initial visual features corresponding to the hand of the preset character based on the hand image in the initial character image; and performing linear transformation and normalization operations on the initial visual features corresponding to the hand of the preset character to determine the hand visual features, which will be used as the visual prior information corresponding to the hand of the preset character.
[0170] Aspect 7: The determination of hand structure features further includes: determining a hand mask in the initial person image based on the initial person image; determining a query matrix for enhancing the hand structure features based on the hand mask and the hand pose features; determining a key matrix and a value matrix for enhancing the hand structure features based on the hand structure features; and determining the enhanced hand structure features using a cross-attention operation with the query matrix, key matrix, and value matrix for enhancing the hand structure features.
[0171] Aspect 8: The determination of hand visual features further includes: determining a hand mask in the initial person image based on the initial person image; determining a query matrix for enhancing the hand visual features based on the hand mask and the final person image of the current iteration; determining a key matrix and a value matrix for enhancing the hand visual features based on the hand visual features; and determining the enhanced hand visual features using a cross-attention operation with the query matrix, key matrix, and value matrix for enhancing the hand visual features.
[0172] Aspect 9: A method for training a neural network model, characterized in that the neural network model includes a hand feature separation and embedding module, a hand pose aggregation module, a try-on module, and a clothing module, wherein the method includes: obtaining clothing image samples and initial person image samples, the clothing image samples including preset clothing, the initial person image samples including a preset person wearing the current clothing, wherein the area where the current clothing is located in the initial person image samples has hand occlusion based on the preset person; based on the clothing image samples, using the clothing module, determining the predicted value of the clothing prior information of the preset clothing; based on the initial person image samples, determining the predicted value of the reference features corresponding to the area outside the area where the current clothing is located; based on the initial person image samples, using the hand feature separation and embedding module and the hand pose aggregation module, determining the preset person... The system uses the try-on module to iteratively update the initial person image samples to generate a predicted image of the final person image, based on the predicted values of the visual prior information and morphological prior information corresponding to the hands of the object; and based on the predicted values of the clothing prior information, the predicted values of the reference features, and the predicted values of the visual prior information and morphological prior information corresponding to the hands of the preset person. In the predicted image of the final person image, the final clothing worn by the preset person corresponds to the preset clothing, and the area where the final clothing is located in the predicted image of the final person image has corresponding hand occlusion. In each iteration, based on the difference between the predicted image of the final person image determined in that iteration and the final person image samples, the hand feature separation and embedding module, the hand pose aggregation module, the try-on module, and the clothing module are trained.
[0173] Aspect 10: In each iteration, training the hand feature separation and embedding module, the hand pose aggregation module, the try-on module, and the clothing module based on the difference between the predicted image corresponding to the final person image sample of the current iteration and the final person image sample includes: in a first stage, training the hand feature separation and embedding module and the hand pose aggregation module while keeping the neural network parameters of the try-on module and the clothing module unchanged; and in a second stage, fine-tuning the neural network parameters of the try-on module and the clothing module while keeping the neural network parameters of the hand feature separation and embedding module and the hand pose aggregation module unchanged.
[0174] Aspect 11: In each iteration, training the hand feature separation and embedding module, the hand pose aggregation module, the try-on module, and the clothing module based on the difference between the predicted image corresponding to the final person image sample of the current iteration and the final person image sample includes: in each iteration, determining the difference on the hand edge between the predicted image corresponding to the final person image sample of the current iteration and the predicted image corresponding to the final person image sample of the previous iteration; and training at least one of the hand feature separation and embedding module, the hand pose aggregation module, the try-on module, and the clothing module based on the difference.
[0175] Aspect 12: An apparatus for generating a final person image, characterized in that the apparatus comprises: an acquisition module, a masking module, a hand feature separation and embedding module, a hand pose aggregation module, a try-on module, and a clothing module, wherein the acquisition module is configured to: acquire a clothing image and an initial person image, the clothing image including a preset clothing, the initial person image including a preset person wearing the current clothing, wherein the area where the current clothing is located in the initial person image has hand occlusion based on the preset person; the clothing module is configured to: determine the clothing prior information of the preset clothing based on the clothing image; the masking module is configured to: determine the reference features corresponding to the area outside the area where the current clothing is located based on the initial person image; and the hand... The feature separation and embedding module and the hand pose aggregation module are configured to: jointly determine the morphological prior information corresponding to the hand of the preset person based on the initial person image; the hand feature separation and embedding module is also configured to: determine the visual prior information corresponding to the hand of the preset person based on the initial person image; and the try-on module is configured to: iteratively update the initial person image based on the clothing prior information, the reference features, and the visual prior information and the morphological prior information to generate a final person image, wherein in the final person image, the final clothing worn by the preset person corresponds to the preset clothing, and the area where the final clothing is located in the final person image has corresponding hand occlusion.
[0176] Aspect 13: An electronic device comprising: one or more processors; and one or more memories, wherein the memories store a computer-executable program that, when executed by the processor, performs the method described in any aspect.
[0177] Aspect 14: A computer-readable storage medium having stored thereon computer instructions which, when executed by a processor, implement the method described in any aspect.
[0178] Aspect 15: A computer program product comprising computer instructions stored in a computer-readable storage medium, wherein a processor of a computer device reads from and executes the computer instructions, causing the computer device to perform the method described in any aspect.
Claims
1. A method for updating a person's image, characterized in that, The method includes: Obtain a clothing image and an initial character image, wherein the clothing image includes a preset clothing, and the initial character image includes a preset character wearing the current clothing, wherein the area where the current clothing is located in the initial character image is obscured by the hands of the preset character; Based on the clothing image, determine the prior information of the preset clothing; Based on the initial character image, determine the reference features corresponding to areas other than the area where the current clothing is located, as well as the visual prior information and morphological prior information corresponding to the hands of the preset character; and Based on the clothing prior information, the reference features, the visual prior information, and the morphological prior information, the initial character image is iteratively updated to generate a final character image. In the final character image, the final clothing worn by the preset character corresponds to the preset clothing, and the area where the final clothing is located in the final character image has a corresponding hand occlusion.
2. The method according to claim 1, wherein, The iterative update of the initial character image includes: In the first iteration, the following steps are performed: The initial character image is concatenated with the reference features to obtain the input features for the first iteration; Based on the visual prior information and the clothing prior information, the input features of the first iteration are encoded to obtain an encoded feature vector; and Based on the visual prior information, the morphological prior information, and the clothing prior information, the encoded feature vector of the first iteration is decoded to obtain the reconstructed image of the first iteration; In subsequent iterations following the first iteration, the following steps are performed: The input features are obtained by concatenating the reconstructed image from the previous iteration with the reference features. Based on the visual prior information and the clothing prior information, the input features are encoded to obtain an encoded feature vector; and Based on the visual prior information, the morphological prior information, and the clothing prior information, the encoded feature vector is decoded to obtain the reconstructed image for the current iteration; In this process, the reconstructed image obtained in each iteration is the predicted image of the final person image corresponding to that iteration, and the reconstructed image obtained in the last iteration is the final person image.
3. The method according to any one of claims 1 to 2, wherein, The process of determining the reference features based on the initial image of the person includes: In the initial character image, the area where the current clothing is located is masked to obtain a masked image; and Based on the masked image, reference features corresponding to the area outside the area where the current clothing is located are determined.
4. The method according to any one of claims 1 to 3, wherein, The determination of prior information about clothing based on clothing images includes: Based on the clothing image, determine at least one of single-resolution clothing features and multi-resolution clothing features; and The prior information of the clothing is determined based on at least one of single-resolution clothing features and multi-resolution clothing features.
5. The method according to any one of claims 1 to 4, wherein, The determination of the prior morphological information corresponding to the hand of the preset character includes: Based on the initial image of the person, determine the hand posture features and hand structural features; and Based on hand posture features and hand structure features, the morphological prior information corresponding to the hand of the preset character is determined.
6. The method according to claim 5, wherein, Determining the hand posture features includes: Based on the initial human image, determine at least one of the following: skeletal structure information, whole-body semantic information, and hand depth information; and The hand pose features are determined based on at least one of skeletal structure information, whole-body semantic information, and hand depth information.
7. The method according to any one of claims 1 to 6, wherein, The determination of the visual prior information corresponding to the hand of the preset character includes: Based on the initial person image, determine the hand image within the initial person image; Based on the hand image in the initial character image, determine the initial visual features corresponding to the hands of the preset character; and A linear transformation and normalization operation is performed on the initial visual features corresponding to the hands of the preset character to determine the visual features of the hands. The visual features of the hands will be used as the visual prior information corresponding to the hands of the preset character.
8. The method according to any one of claims 5 to 7, wherein, The determination of hand structural features includes: Based on the initial image of the person, determine the hand mask in the initial image of the person; Based on the hand mask and the hand pose features, a query matrix is determined to enhance the hand structural features; Based on the aforementioned hand structural features, a key matrix and a value matrix are determined to enhance the hand structural features; and Using a query matrix, key matrix, and value matrix to enhance the hand structural features, the enhanced hand structural features are determined by cross-attention operation.
9. The method according to any one of claims 7 to 8, wherein, The determination of hand visual features includes: Based on the initial image of the person, determine the hand mask in the initial image of the person; Based on the hand mask and the reconstructed image of the current iteration, a query matrix is determined to enhance the visual features of the hand; Based on the aforementioned hand visual features, a key matrix and a value matrix are determined to enhance the hand visual features; and Using a query matrix, key matrix, and value matrix for enhancing the visual features of the hand, the enhanced visual features of the hand are determined by employing a cross-attention operation.
10. A method for training a neural network model, characterized in that, The neural network model includes a hand feature separation and embedding module, a hand pose aggregation module, a try-on module, and a clothing module, wherein the method includes: Obtain clothing image samples and initial character image samples. The clothing image samples include preset clothing, and the initial character image samples include preset characters wearing the current clothing. In the initial character image samples, the area where the current clothing is located is obscured by the hands of the preset characters. Based on the clothing image samples, the clothing module is used to determine the predicted value of the clothing prior information of the preset clothing. Based on the initial human image sample, the predicted value of the reference feature corresponding to the area outside the area where the current clothing is located is determined; Based on initial human image samples, the hand feature separation and embedding module and the hand pose aggregation module are used to determine the predicted values of the visual prior information and the morphological prior information corresponding to the hands of the preset human figure; and Based on the predicted values of the clothing prior information, the predicted values of the reference features, and the predicted values of the visual and morphological prior information corresponding to the hands of the preset person, the initial person image samples are iteratively updated using the try-on module to generate a predicted image of the final person image. In the predicted image of the final person image, the final clothing worn by the preset person corresponds to the preset clothing, and the area where the final clothing is located in the predicted image of the final person image has a corresponding hand occlusion. In each iteration, the hand feature separation and embedding module, the hand pose aggregation module, the try-on module, and the clothing module are trained based on the difference between the predicted image of the final person image determined in that iteration and the final person image sample.
11. The method of claim 9, wherein, In each iteration, the difference between the predicted image of the final person image determined in that iteration and the final person image sample is used to train the hand feature separation and embedding module, the hand pose aggregation module, the try-on module, and the clothing module, including: In the first stage, while keeping the neural network parameters of the try-on module and the clothing module unchanged, the hand feature separation and embedding module and the hand pose aggregation module are trained; and In the second stage, the neural network parameters of the hand feature separation and embedding module and the hand pose aggregation module are kept unchanged, while the neural network parameters of the try-on module and the clothing module are fine-tuned.
12. The method according to any one of claims 9-10, wherein, In each iteration, the difference between the predicted image of the final person image determined in that iteration and the final person image sample is used to train the hand feature separation and embedding module, the hand pose aggregation module, the try-on module, and the clothing module, including: In each iteration, the difference between the predicted image corresponding to the final person image sample in the current iteration and the predicted image corresponding to the final person image sample in the previous iteration at the hand edge is determined; and Based on the differences, at least one of the hand feature separation and embedding module, the hand pose aggregation module, the try-on module, and the clothing module is trained.
13. A device for updating a person's image, characterized in that, The device includes: an acquisition module, a mask module, a hand feature separation and embedding module, a hand pose aggregation module, a try-on module, and a clothing module, wherein... The acquisition module is configured to: acquire a clothing image and an initial character image, wherein the clothing image includes a preset clothing, and the initial character image includes a preset character wearing the current clothing, wherein the area where the current clothing is located in the initial character image is obscured by the hand of the preset character; The clothing module is configured to: determine the prior information of the preset clothing based on the clothing image; The masking module is configured to: determine reference features corresponding to the area outside the area where the current clothing is located, based on the initial image of the person; The hand feature separation and embedding module and the hand pose aggregation module are configured to: jointly determine the morphological prior information corresponding to the hand of the preset person based on the initial person image; The hand feature separation and embedding module is further configured to: determine the visual prior information corresponding to the hands of the preset person based on the initial person image; and The try-on module is configured to iteratively update the initial character image based on the clothing prior information, the reference features, the visual prior information, and the morphological prior information to generate a final character image. In the final character image, the final clothing worn by the preset character corresponds to the preset clothing, and the area where the final clothing is located in the final character image has a corresponding hand occlusion.
14. An electronic device comprising: One or more processors; and One or more memories, wherein the memories store a computer-executable program that, when executed by the processor, performs the method of any one of claims 1-12.
15. A computer-readable storage medium having stored thereon computer instructions that, when executed by a processor, implement the method as claimed in any one of claims 1-12.