Method for acquiring wearing image and related apparatus

By acquiring real-life images of users and 3D models of target items, and combining 3D modeling and fusion technologies, virtual wearable images are generated, solving the problem of existing technologies being unable to display the right size and achieving a more realistic virtual try-on effect.

WO2025261119A1PCT designated stage Publication Date: 2025-12-26HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/097949
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-30
Filing Date
2025-05-29
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing virtual try-on technology cannot effectively demonstrate whether the size of the target item is suitable when the user wears it, resulting in the user not being able to truly experience the virtual wearing effect.

Method used

By acquiring real-life images of users and 3D models of target items, and combining 3D modeling and fusion technologies, virtual wearable images are generated to demonstrate whether the size of the target item is suitable for the user. Furthermore, semantic information and image enhancement technologies are used to improve the realism of the images and the user experience.

Benefits of technology

This technology enables virtual wearable images to accurately reflect whether the size of the target item is suitable for the user, thus improving the user experience and the realism of virtual try-on.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025097949_26122025_PF_FP_ABST
    Figure CN2025097949_26122025_PF_FP_ABST
Patent Text Reader

Abstract

A method for acquiring a wearing image and a related apparatus, which can be used in the field of virtual fitting. The method comprises: acquiring a real-shot image and a first image, wherein the real-shot image is a two-dimensional image captured by a camera, the real-shot image comprises a plurality of human body parts of a user, the first image is a two-dimensional image obtained by rendering a wearing three-dimensional model, and the wearing three-dimensional model is obtained by combining a three-dimensional model of at least one of the plurality of human body parts with a three-dimensional model of a target item; and then performing a fusion operation on the real-shot image and the first image to obtain a second image, wherein the second image comprises an image of a plurality of human body parts wearing the target item. The first image can reflect information about whether the size of the target item is appropriate when the user wears the target item, and the real-shot image can carry more realistic user information. Therefore, the second image can not only show whether the size of the target item is appropriate when the user wears the target item, but also can restore the real situation of the user.
Need to check novelty before this filing date? Find Prior Art

Description

A method and related apparatus for acquiring wearable images

[0001] This application claims priority to Chinese Patent Application No. 202410804373.8, filed on June 20, 2024, entitled "A Method and Apparatus for Trying on Clothes", and to Chinese Patent Application No. 202411388117.1, filed on September 30, 2024, entitled "A Method and Apparatus for Acquiring Wearing Images", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to computer technology, and more particularly to a method and apparatus for acquiring wearable images. Background Technology

[0003] With the development of technology, people's lives are becoming more and more convenient. With the rapid development of e-commerce, purchasing wearable items through online platforms has become a major choice for users. However, when purchasing wearable items through online platforms, users cannot try them on like in physical stores, and therefore cannot see the effect of wearing the item.

[0004] In recent years, virtual try-on technology has become the main way to solve the above problems. For example, it can acquire a two-dimensional image of the user and a real-life image of the target item to be worn, and synthesize the user's real-life image and the target item's real-life image to obtain the user's virtual wear image. However, the virtual wear image obtained by the above method cannot show whether the size of the target item to be tried on is suitable. Summary of the Invention

[0005] This application provides a method and related apparatus for acquiring wearable images. The obtained wearable images can not only show whether the size is suitable when the user wears the target item, but also better reproduce the user's actual situation.

[0006] This application provides the following technical solution:

[0007] Firstly, this application provides a method for acquiring wearable images, which can be used in the field of virtual try-on. In this method, a first device acquires at least one real-life image of a user. Each user's real-life image is an image of the user taken by a camera; in other words, each user's real-life image is a real image, including multiple body parts of the user. The first device can also acquire a first image corresponding to each user's real-life image. The first image includes a two-dimensional image obtained by rendering a three-dimensional (3D) wearable model to two dimensions. The first device can perform a fusion operation on each real-life image and the first image corresponding to each real-life image to obtain a second image corresponding to each real-life image. The second image includes images of the target item worn on multiple body parts of the user, that is, a second image combining the real and virtual elements is obtained based on the real image (i.e., the user's real-life image) and the virtual image (i.e., the first image).

[0008] Optionally, before acquiring the first image, the first device may also receive a target item selection request, which instructs the first device to determine the target item from multiple items.

[0009] For example, if the user's real-life image is a head image, multiple body parts of the user are included within the head, which could include the skull, hair, and neck. As another example, if the user's real-life image is an upper-body image, multiple body parts of the user are included within the head and torso, which could include the head, torso, arms, and hands.

[0010] For example, if the user's real-life image is a full-body image, multiple body parts can include the head, torso, arms, hands, legs, and feet. Or, if the user's real-life image is an image of their hand, multiple body parts are contained within the torso and can include the wrist, palm, and fingers.

[0011] It should be noted that the granularity of human body parts division can vary in different application scenarios. For example, in some scenarios, the head is divided into three human body parts: skull, hair, and neck, while in other scenarios, the head is considered as a single human body part. The granularity of human body parts division in this application can be determined in conjunction with the actual application scenario.

[0012] For example, the real-shot image of at least one user can be at least one video frame in the video; or, the real-shot image of each user can be an independently captured real-shot image. In other words, the real-shot image of each user may not be a video frame in the video, but an independently captured real-shot image.

[0013] For example, the above fusion operation may specifically include at least one of the following: cutting out the real-shot image and the first image respectively and then stitching them together; cutting out the real-shot image and the first image respectively and then stitching them together and filling in the remaining parts; directly superimposing the real-shot image and the first image; using the real-shot image to enhance the first image; or, using the real-shot image and the first image to regenerate the second image, etc.

[0014] In this implementation, since the dimensions of the target item and at least one human body part of the user can be fully reflected in three-dimensional space, a wearable three-dimensional model is obtained by combining the three-dimensional model of at least one human body part and the three-dimensional model of the target item. This wearable three-dimensional model can reflect whether the size of the target item is suitable for the user. The wearable three-dimensional model is rendered into two dimensions to obtain the first image. The first image retains the information on whether the size of the target item is suitable for the user. The real-shot image contains more realistic human body parts. The second image is obtained by performing a fusion operation on the real-shot image and the first image. Therefore, the second image can not only show whether the size of the target item is suitable for the user, but also better restore the user's real situation.

[0015] In one possible implementation, the second image is obtained by stitching together a first image region from the captured image and a second image region from the first image. The first image region includes at least one human body part not covered by an object similar to the target object, and the second image region includes the target object. For example, the first device performs a fusion operation on the captured image and the first image to obtain the second image, including: stitching together the first image region from the captured image and the second image region from the first image to obtain the second image.

[0016] This implementation provides a method to obtain a second image based on a real-shot image and a first image. The second image is obtained by stitching together the first image region in the real-shot image and the second image region in the first image. The second image region in the first image can show whether the size of the target item is suitable when the user wears it. Thus, the second image can reflect whether the size of the target item is suitable. Since the first image region in the real-shot image can reflect the user's real situation, the realism of the second image is greatly improved, which enhances the user's feeling of trying on the item in real life and is conducive to improving the user experience of this solution.

[0017] In one possible implementation, the second image is obtained based on the real-shot image, the first image, first semantic information of the real-shot image, and second semantic information of the first image. For example, before stitching together the first image region in the real-shot image with the second image region in the first image to obtain the second image, the first device may also perform semantic recognition on the real-shot image to obtain first semantic information of the real-shot image, which indicates the category of pixels in the real-shot image; determine the first image region in the real-shot image based on the first semantic information, where the category of pixels in the first image region is different from the category of the target item; perform semantic recognition on the first image to obtain second semantic information of the first image, which indicates the category of pixels in the first image; and determine the second image region in the first image based on the second semantic information, where the category of pixels in the second image region is consistent with the category of the target item.

[0018] In this implementation, first semantic information and second semantic information are also obtained. The first semantic information indicates the category of pixels in the real-shot image. The first image region in the real-shot image can be determined based on the first semantic information. The second semantic information indicates the category of pixels in the first image. The second image region in the first image can be determined based on the second semantic information. This reduces the difficulty of the stitching step and helps to improve the quality when performing the stitching step, so as to obtain a second image of better quality.

[0019] In one possible implementation, the first semantic information further indicates pixels in the user's real-world image categorized as torso, arm, hand, leg, and / or foot. If the first image region in the user's real-world image is larger than the second image region in the first image, image enhancement can also be performed on blank areas in the user's real-world image based on the first semantic information. The blank areas in the user's real-world image are the first image region minus the second image region. For example, if pixels within the first image region of the user's real-world image are removed, and the real-world image with the first image region removed is then stitched together with the second image region, blank areas (i.e., the first image region minus the second image region) still exist in the user's real-world image. In this case, image enhancement can be performed on the pixels in the blank areas based on their semantic categories to further improve the realism of the second image.

[0020] In one possible implementation, the second image is obtained by enhancing the exposed areas in the first image using the user's face area from the captured image. The exposed areas in the first image can also be understood as areas not covered by clothing, including the face. Optionally, the exposed areas in the first image may also include exposed skin, exposed hair, or other types of exposed areas.

[0021] This implementation provides an alternative approach to obtaining a second image based on the user's real-shot image and the first image, improving the flexibility of the solution. Furthermore, by using the facial region in the user's real-shot image to enhance the first image and obtain the second image, the second image becomes more similar to the real user, and the second image can more accurately reflect whether the size of the target item is appropriate.

[0022] In one possible implementation, the second image is obtained based on the user's real-shot image, the first image, third semantic information of the user's real-shot image, and fourth semantic information of the first image. For example, before the first device performs a fusion operation with the first image region in the real-shot image and the first image to obtain the second image, the method further includes: the first device performing semantic recognition on the real-shot image to obtain third semantic information of the real-shot image, the third semantic information indicating the category of pixels in the real-shot image; determining the user's face region in the real-shot image based on the third semantic information, the category of pixels in the user's face region in the real-shot image being "face"; performing semantic recognition on the first image to obtain fourth semantic information of the first image, the fourth semantic information indicating the category of pixels in the first image; and determining the exposed area in the first image based on the fourth semantic information, the category of pixels in the exposed area including "face".

[0023] In this implementation, the third semantic information can be used to quickly and accurately locate the facial region in the user's real-shot image, and the fourth semantic information can be used to quickly and accurately locate the user's exposed area in the first image. Thus, by using the third and fourth semantic information, the first image can be enhanced using the facial region in the user's real-shot image, which helps to reduce the difficulty of the aforementioned image enhancement part and helps to obtain a second image of better quality.

[0024] In one possible implementation, the third semantic information further indicates pixels in the background region of the captured image, and the fourth semantic information further indicates pixels in the background region of the first image. Each second image is obtained by enhancing the exposed area of ​​the user in the first image using the user's face region in the captured image, and enhancing the background region in the first image using the background region of the captured image.

[0025] In this implementation, third and fourth semantic information can be used to enhance the background area in the first image based on the background area in the real-shot image, thereby further improving the realism of the second image and creating a more realistic feeling for the user wearing the target item, thus improving the user experience of this solution.

[0026] In one possible implementation, the first device acquires a first image, which includes an image obtained by rendering a wearable 3D model to 2D. The process includes: acquiring a 3D model of at least one human body part of a user, obtained by performing 3D modeling based on real-life images of the user; acquiring a 3D model of a target item; and then generating at least one wearable 3D model corresponding one-to-one with the real-life images of the user based on the 3D model of the user's at least one human body part and the 3D model of the target item. Each wearable 3D model can be understood as a model of the user's at least one human body part wearing the 3D model of the target item in the posture shown in the real-life image. Each first image includes an image obtained by rendering each wearable 3D model to 2D. In this implementation, the 3D model of at least one human body part of the user is obtained by performing 3D modeling based on real-life images. Therefore, the 3D model of the user's at least one human body part is closer to the user's actual appearance, and the wearable 3D model can more accurately reflect whether the size of the target item is suitable for the user. Consequently, the first and second images can more accurately reflect whether the target item fits the user, thus improving the user experience.

[0027] In one possible implementation, when the acquired at least one real-shot image is at least one video frame from a captured video, the first device generates a first image based on a 3D model of at least one human body part of the user and a 3D model of the target object. This may include: the first device generating at least one first image corresponding one-to-one with the at least one video frame based on the 3D model of the user's at least one human body part, the 3D model of the target object, and the at least one video frame, wherein the user's 3D model is obtained by performing 3D modeling on one of the at least one video frame. The first device obtains a second image corresponding to each user's real-shot image based on each real-shot image and the first image corresponding to each real-shot image. Optionally, the aforementioned video frame among the at least one video frame is the first video frame in the at least one video frame.

[0028] In this implementation, when at least one real-shot image of the user is at least one video frame in the video, since the user's three-dimensional dimensions are the same in different video frames, three-dimensional modeling can be performed based on only one video frame to obtain a three-dimensional model of at least one human body part of the user. Then, using the three-dimensional model of at least one human body part of the user corresponding to the aforementioned video frame, the three-dimensional model of the target object, and at least one video frame, at least one first image corresponding to at least one video frame is generated, without repeating three-dimensional modeling for each video frame, thereby reducing the consumption of computer resources.

[0029] In one possible implementation, multiple body parts are included in the torso and / or head. For example, if the user's real-life image is a head image, multiple body parts are included in the head. Or, if the user's real-life image is an upper-body image, multiple body parts are included in the head and torso. Or, if the user's real-life image is a full-body image, multiple body parts are included in the head and torso, etc. The various representations of multiple body parts provided in this application embodiment improve the implementation flexibility of this solution and also facilitate the expansion of its application scenarios.

[0030] In one possible implementation, the camera position set in the rendering is obtained based on the camera position of the user's actual captured image, and / or, the shooting angle set in the rendering is obtained based on the shooting angle of the user's actual captured image.

[0031] In this implementation, it is beneficial to ensure that the camera position and / or shooting angle used during rendering are consistent with the camera position and / or shooting angle when capturing the user's actual image. This makes the camera position and / or shooting angle corresponding to the first image consistent with the camera position and / or shooting angle of the user's actual image. In other words, alignment between the first image and the user's actual image is achieved in the dimension of camera position and / or shooting angle, which is beneficial to further improve the realism of the updated virtual wearable target item.

[0032] Optionally, when performing 3D modeling operations, the first device can obtain the camera position and shooting angle when capturing the user's actual images. For example, when performing 3D modeling based on the first video frame in at least one video frame, the first device can obtain the camera position and shooting angle when capturing that first video frame; correspondingly, when performing 3D modeling based on each independent actual image, the first device can obtain the camera position and shooting angle when capturing each independent actual image.

[0033] Secondly, this application provides a device for acquiring wearable images, which can be used in the field of virtual try-on. The device includes: an acquisition module for acquiring a two-dimensional real-shot image of a user, wherein the real-shot image of the user is a photographed image; the acquisition module is further configured to acquire a first image, wherein the first image is a first image of the user, and the first image is an image obtained by rendering a three-dimensional target object worn by the user in a three-dimensional space to two dimensions; and a processing module for obtaining a second image based on the real-shot image of the user and the first image, wherein the second image is a second image of the user.

[0034] In the second aspect of this application, the wearable image acquisition device may also perform the steps performed by the first device in the first aspect and various possible implementations of the first aspect. The specific implementation of the steps in the second aspect and various possible implementations of the second aspect, the meaning of the terms, and the beneficial effects brought about by each possible implementation can all be referred to the description in the first aspect and various possible implementations of the first aspect, and will not be repeated here.

[0035] Thirdly, this application provides a computing device cluster, which includes at least one computing device, each computing device including a processor and a memory: the memory is used to store instructions; the processor is used to cause the computing device cluster to perform the method described in the first aspect or any possible implementation of the first aspect according to the instructions.

[0036] Fourthly, this application provides a computer storage medium storing one or more instructions, which, when executed by one or more computers, cause the one or more computers to perform the method described in the first aspect or any possible implementation of the first aspect.

[0037] Fifthly, this application provides a computer program product that stores instructions that, when executed by a computer, cause the computer to perform the method described in the first aspect or any of the possible implementations of the first aspect.

[0038] The beneficial effects brought about by the second to fifth aspects of this application can be referred to the descriptions in the first aspect and the various possible implementations of the first aspect, and will not be repeated here. Attached Figure Description

[0039] Figure 1 is a system architecture diagram of a wearable image acquisition system provided in an embodiment of this application;

[0040] Figure 2 is a schematic flowchart of a method for acquiring wearable images provided in an embodiment of this application;

[0041] Figure 3 is a schematic diagram of a three-dimensional model of clothing and a mesh that makes up the three-dimensional model of clothing provided in an embodiment of this application;

[0042] Figure 4 is a schematic diagram of a three-dimensional model of at least one human body part of a user provided in an embodiment of this application;

[0043] Figure 5 is a schematic diagram of a user's real-shot image, a first image, and a second image provided in an embodiment of this application;

[0044] Figure 6 is a schematic diagram of a user's real-shot image provided in an embodiment of this application and a first image region removed from the real-shot image;

[0045] Figure 7 is a schematic diagram of obtaining a second image region from a first image according to an embodiment of this application;

[0046] Figure 8 is a schematic diagram of the face region, the first image, and the second image in a real-shot image provided in an embodiment of this application;

[0047] Figure 9 is a schematic diagram of obtaining an image of a user's exposed area from a first image according to an embodiment of this application;

[0048] Figure 10 is a schematic diagram of a wearable image acquisition device provided in an embodiment of this application;

[0049] Figure 11 is a schematic diagram of a wearable image acquisition system provided in an embodiment of this application;

[0050] Figure 12 is a schematic diagram of a computing device provided in an embodiment of this application;

[0051] Figure 13 is a schematic diagram of a computing device cluster provided in an embodiment of this application;

[0052] Figure 14 is a schematic diagram of computer devices in a computer cluster connected via a network according to an embodiment of this application. Detailed Implementation

[0053] The embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are merely some, and not all, of the embodiments of this application. Those skilled in the art will recognize that, with the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0054] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the description of embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0055] In the embodiments of this application, "instruction" can include direct and indirect instructions, as well as explicit and implicit instructions. The information indicated by a certain piece of information (hereinafter referred to as instruction information) is called the information to be instructed. In specific implementation, there are many ways to indicate the information to be instructed, such as, but not limited to, directly indicating the information to be instructed, such as the information to be instructed itself or its index. It can also indirectly indicate the information to be instructed by indicating other information, where there is an association between the other information and the information to be instructed; or it can indicate only a part of the information to be instructed, while the other parts are known or pre-agreed upon. For example, the instruction can be implemented by using a pre-agreed (e.g., protocol predefined) arrangement of various information, thereby reducing the instruction overhead to a certain extent. This application does not limit the specific method of instruction. It is understood that for the sender of the instruction information, the instruction information can be used to indicate the information to be instructed; for the receiver of the instruction information, the instruction information can be used to determine the information to be instructed.

[0056] The method provided in this application can be applied to various virtual try-on scenarios; for example, when a user purchases a target item through an online platform, virtual try-on technology can be used to obtain a virtual image of the user wearing the target item; another example is when a user wants to try on a target item in a game, virtual try-on technology can be used to obtain a virtual image of the user wearing the target item, and so on, without exhaustive list.

[0057] Current virtual try-on technology can include: acquiring a two-dimensional (2D) image of the user and a real-life image of the target item to be worn, and then synthesizing the user's real-life image and the target item's real-life image to obtain a virtual wear image of the user. However, the virtual wear image obtained by the aforementioned method cannot show whether the size of the target item is suitable when trying it on. To solve the aforementioned problem, this application discloses that: the first device not only acquires a real-life image of the user, which is a two-dimensional image captured by a camera and includes multiple body parts of the user; the first device also acquires a first image, which includes a two-dimensional image obtained by rendering a three-dimensional (3D) model of the wearer onto a two-dimensional plane. The three-dimensional model of the wearer is obtained by combining a three-dimensional model of at least one of the aforementioned multiple body parts with a three-dimensional model of the target item; then the first device can combine the aforementioned real-life image and the first image to obtain a second image, which includes an image of the user's multiple body parts wearing the target item. In other words, the second image can be understood as a virtual wear image of the user's multiple body parts wearing the target item.

[0058] Since three-dimensional space can fully reflect the size of the target item and at least one human body part of the user, the wearable three-dimensional model obtained by combining the three-dimensional model of at least one human body part and the three-dimensional model of the target item can reflect whether the size of the target item is suitable for the user. The wearable three-dimensional model is rendered into two dimensions to obtain the first image. The first image retains the information on whether the size of the target item is suitable for the user. The real-shot image contains more realistic human body parts. The second image is obtained by performing a fusion operation on the real-shot image and the first image. Therefore, the second image can not only show whether the size of the target item is suitable for the user, but also better restore the user's real situation.

[0059] Optionally, a machine learning model is used in the process of generating the second image. Before describing the specific implementation process of the method provided in this application in detail, the virtual wearable image acquisition system provided in the embodiment of this application will be introduced with reference to Figure 1. Please refer to Figure 1 first. Figure 1 is a system architecture diagram of the virtual wearable image acquisition system provided in the embodiment of this application. In Figure 1, the virtual wearable image acquisition system includes a training device 110, a database 120, a cloud server 130, a data storage system 140, and a client device 150. The cloud server 130 includes a computing module 131.

[0060] The database 120 stores a training data set. The training device 110 uses the training data set to iteratively train the machine learning model 101 to obtain the machine learning model 101 that has undergone training operations. The machine learning model 101 can be specifically represented as a neural network or as a non-neural network model.

[0061] The machine learning model 101, trained by the training device 110, can be deployed to the computing module 131 of the cloud server 130. The cloud server 130 can access data, code, etc., from the data storage system 140, and can also store data, instructions, etc., in the data storage system 140. The data storage system 140 can be located within the cloud server 130, or it can be an external storage device relative to the cloud server 130. The cloud server 130 can generate a second image using the machine learning model 101 in the computing module 131.

[0062] In some embodiments of this application, please refer to FIG1. ​​The cloud server 130 and the client device 150 can be separate independent devices. The cloud server 130 is configured with an input / output (I / O) interface to interact with the client device 150. The client device sends the data to be processed to the cloud server 130 through the I / O interface. After the cloud server 130 generates the second image through the machine learning model 101 in the computing module 131, it can return the aforementioned second image to the client device 150 through the I / O interface.

[0063] It is worth noting that Figure 1 is merely a schematic diagram of one architecture of the virtual wearable image acquisition system provided in an embodiment of the present invention, and the positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in some other embodiments of this application, the cloud server 130 and the client device 150 may be integrated into the same device, so the user can directly interact with the cloud server 130. Exemplarily, the cloud server 130 may be a module in the host CPU of the client device 150 that uses a machine learning model to process data. The cloud server 130 may also be a graphics processing unit (GPU) or a neural network processor (NPU) in the client device, with the GPU or NPU acting as a coprocessor mounted on the host processor, and the host processor allocating tasks. As another example, in some other embodiments of this application, the training device 110 and the cloud server 130 may be integrated into the same device.

[0064] Based on the above description, the specific implementation flow of the method provided in the embodiments of this application will be described below. Specifically, please refer to Figure 2, which is a schematic flowchart of a method for acquiring wearable images provided in the embodiments of this application. The method for acquiring wearable images provided in the embodiments of this application may include:

[0065] 201. Acquire real-shot images. Real-shot images are two-dimensional images captured by the camera and include multiple body parts of the user.

[0066] For example, the first device can acquire at least one real-shot image of the user, wherein the real-shot image may also be referred to as a real image or other names, etc. The real-shot image of the user may include multiple body parts of the user.

[0067] Optionally, multiple body parts of the user can be included in the head and / or torso. For example, if the user's real-life image is a head image, multiple body parts of the user can be included in the head, such as the skull, hair, and neck. As another example, if the user's real-life image is an upper-body image, multiple body parts of the user can be included in the head and torso, such as the head, torso, arms, and hands.

[0068] For example, if the user's real-life image is a full-body image, multiple body parts of the user are included within the head and torso. These multiple body parts could include the head, torso, arms, hands, legs, and feet. Similarly, if the user's real-life image is an image of the user's hand, multiple body parts of the user are included within the torso. These multiple body parts could include the wrist, palm, and fingers, etc.

[0069] It should be noted that the granularity of human body parts division can vary in different application scenarios. For example, in some scenarios, the head is divided into three human body parts: skull, hair, and neck, while in other scenarios, the head is considered as a single human body part. The granularity of human body parts division in this application can be determined in conjunction with the actual application scenario.

[0070] The various representations of multiple body parts provided in this application embodiment improve the implementation flexibility of this solution and also help to expand the application scenarios of this solution.

[0071] In one scenario, the at least one captured image can be at least one video frame from a video, meaning each user's captured image is a single video frame. It should be noted that if the first device acquires a video, the video can be segmented into at least one video frame. In another scenario, the captured image for each user can be an independently captured image; in other words, the captured image for each user may not be a video frame from a video, but rather an independently captured image.

[0072] For example, in one scenario, the first device is a cloud server, and step 201 may include: the first device receiving at least one real-shot image sent by the client device. In another scenario, the first device is a client device that directly interacts with the user, and step 201 may include: the first device obtaining at least one real-shot image from a local photo album, or step 201 may include: the first device capturing at least one real-shot image using a camera.

[0073] In step 201, optionally, after acquiring at least one real-shot image, the first device may also perform image enhancement on each real-shot image to improve the clarity of each real-shot image.

[0074] For example, a large model for image processing can be used to enhance each real-world image. For instance, a first prompt and each real-world image are input into the large model for image processing to obtain an updated real-world image output by the large model. The first prompt can be used to prompt the large model to improve the clarity of the input image. The large model for image processing is a machine learning model. Optionally, the aforementioned large model for image processing can be a machine learning model based on an attention mechanism.

[0075] Alternatively, a machine learning model specifically designed to improve image clarity can be used to process each real-world image to obtain an updated real-world image. In other words, the function of the aforementioned machine learning model is to improve image clarity. For example, the aforementioned machine learning models can be convolutional neural networks, fully connected neural networks, residual neural networks, attention-based neural networks, support vector machines, or other types of machine learning models.

[0076] 202. Obtain a first image, which includes a two-dimensional image obtained by rendering the wearable three-dimensional model. The wearable three-dimensional model is obtained by combining a three-dimensional model of at least one human body part from multiple human body parts with a three-dimensional model of the target object.

[0077] Optionally, before performing step 202, the first device may also receive a target item selection request. The target item selection request instructs the first device to determine the target item from multiple items. Optionally, the target item selection request may include a unique identifier for the target item. For example, in one scenario, the first device is a cloud server, and step 202 may include: the first device receiving a target item selection request sent by a client device. In another scenario, the first device is a client device that directly interacts with the user, and step 202 may include: the first device receiving a target item selection request input by the user. For example, when the user clicks on a target item, the first device may receive the user-input target item selection request; or, for example, when the user drags the target item to a preset area, the first device may receive the user-input target item selection request, etc. The specific details can be determined based on the actual application scenario.

[0078] For example, the category of the target items mentioned above (hereinafter referred to as "Category 1" for ease of description) can be wigs, tops, skirts, headphones, glasses, necklaces, bracelets, watches, pants, shoes, or other types of target items, etc., which can be determined in combination with the actual application scenario.

[0079] At least one of the aforementioned multiple human body parts includes at least the human body part wearing the target item; optionally, the aforementioned at least one human body part includes all of the aforementioned multiple human body parts.

[0080] For example, if the target item is a wig, at least one of the multiple human body parts may include the head; optionally, at least one of the multiple human body parts may also include the neck or other human body parts. As another example, if the target item is a top, at least one of the multiple human body parts may include the torso and arms; optionally, the at least one human body part may also include the head and neck. As another example, if the target item is a long skirt, the at least one human body part may include the head, torso, arms, hands, legs, and feet. As another example, if the target item is headphones, the at least one human body part may include the head; optionally, the at least one human body part may also include hair and the neck. The specific parts included in the at least one human body part can be determined based on the actual application scenario.

[0081] For example, the first device may store first information representing a three-dimensional model of the target item. This first information may include first position information of each of a plurality of first vertices constituting the three-dimensional model of the target item. For instance, the first position information of each first vertex may be the coordinates of each first vertex in three-dimensional space, such as coordinates on the x-axis, y-axis, and z-axis. The first information also includes first connection relationship information between different first vertices among the plurality of first vertices. This first connection relationship information indicates which of the plurality of first vertices are connected, and the connected first vertices can be connected to form multiple meshes constituting the three-dimensional model of the target item. Optionally, the first information may also include at least one of the following information of the three-dimensional model of the target item: texture information, material information, color information, weight, or other physical parameter information, etc.

[0082] To understand this solution more intuitively, please refer to Figure 3. Figure 3 is a schematic diagram of a three-dimensional model of clothing and the mesh that makes up the three-dimensional model of clothing provided in the embodiment of this application. Figure 3 includes a left sub-schematic diagram and a right sub-schematic diagram. The left sub-schematic diagram of Figure 3 shows the three-dimensional model of clothing. The right sub-schematic diagram of Figure 3 takes the collar part of the aforementioned three-dimensional model of clothing as an example to show the vertices and mesh of the collar part that makes up the three-dimensional model of clothing. It should be noted that Figure 3 is only a schematic diagram obtained after visualizing the three-dimensional model of clothing in this solution for easy understanding. The first device can be deployed as the first information used to represent the three-dimensional model of the target item. In addition, the example in Figure 3 is only for easy understanding of the concept of the three-dimensional model of the target item and is not intended to limit this solution.

[0083] For example, step 202 may include: the first device acquiring at least one first image corresponding one-to-one with the at least one real-shot image; in one case, the first device is a cloud server, and step 201 may include: the first device receiving at least one first image sent by a client device corresponding one-to-one with the at least one real-shot image acquired in step 201; or, step 201 may include: the first device generating the at least one first image based on the at least one real-shot image acquired in step 201. In another case, the first device is a client device that directly interacts with the user, and step 201 may include: the first device generating the at least one first image based on the at least one real-shot image acquired in step 201.

[0084] Regarding the specific implementation of the first device generating at least one first image based on at least one acquired real-shot image, exemplarily, the first device can acquire a 3D model of at least one human body part of the user, acquire a 3D model of the target item, and then obtain a wearable 3D model based on the 3D model of the user's at least one human body part and the 3D model of the target item. The first image includes an image obtained by rendering the wearable 3D model to 2D. Optionally, the 3D model of the user's at least one human body part is obtained by performing 3D modeling based on the user's real-shot image. In this embodiment of the application, the user's 3D model is obtained by performing 3D modeling based on the user's real-shot image, so the aforementioned 3D model of at least one human body part can be closer to the user's actual situation. Therefore, the wearable 3D model can more accurately reflect whether the size of the user's at least one human body part is suitable when wearing the target item. Consequently, the first image and the second image can also more accurately reflect whether the target item fits the user, which is beneficial to improving the user experience.

[0085] Regarding the specific implementation of the first device acquiring a 3D model of at least one human body part of a user, in one case, if the at least one real-shot image obtained in step 201 is at least one video frame in a video, the first device can use a 3D modeling algorithm to perform 3D modeling based on only one of the aforementioned at least one video frame to obtain a 3D model of a user corresponding to that one video frame; optionally, the aforementioned one video frame can be the first video frame in the at least one video frame, or the aforementioned one video frame can be any one of the at least one video frames. Alternatively, the first device can also use a 3D modeling algorithm to perform 3D modeling based on each of the aforementioned at least one video frame to obtain a 3D model of at least one user corresponding to each of the at least one video frame. In another case, if the at least one real-shot image obtained in step 201 is at least one independently captured real-shot image, the first device can use a 3D modeling algorithm to perform 3D modeling based on each of the aforementioned at least one independently captured real-shot images to obtain a 3D model of at least one user corresponding to each of the at least one independently captured real-shot images.

[0086] For example, the 3D modeling algorithm can be the 3D human pose and shape regression with pyramidal mesh aligns the feedback loop (PyMAF) algorithm, the 4DHumans algorithm, or other 3D modeling algorithms, etc. The specific algorithm can be determined based on the actual application scenario.

[0087] For example, the 3D model of at least one human body part of the user can be a model of the user's entire body, or it can include a model of the user's upper body; or, the 3D model of at least one human body part of the user can be a model of the user's head, etc. The specific representation of the 3D model of at least one human body part of the user can be determined in combination with the actual application scenario. For example, if the target item to be tried on is a wig, headphones, glasses, or necklace, then the 3D model of at least one human body part of the user can be a model of the user's head; as another example, if the target item to be tried on is a top, bracelet, or watch, then the 3D model of at least one human body part of the user can be a model of the user's upper body; as another example, if the target item to be tried on is pants or shoes, then the 3D model of at least one human body part of the user can be a model of the user's entire body, etc., and so on.

[0088] For example, after performing the aforementioned 3D modeling operation, the first device can obtain second information for a 3D model representing at least one human body part of the user. Each piece of second information may include second position information of each of the plurality of second vertices constituting the 3D model of the at least one human body part of the user. For example, the second position information of each second vertex may be the coordinates of each second vertex in 3D space. The second information may also include second connection relationship information between different second vertices among the aforementioned plurality of second vertices. This second connection relationship information indicates which of the aforementioned plurality of second vertices are connected. Second vertices with connection relationships can be connected to obtain multiple meshes constituting the user's 3D model. Optionally, the second information may also include skin color or other information, etc., which are not exhaustively listed here.

[0089] To understand this solution more intuitively, please refer to Figure 4. Figure 4 is a schematic diagram of a three-dimensional model of at least one human body part of a user provided in an embodiment of this application. In Figure 4, the three-dimensional model of at least one human body part of a user is taken as an example of the model of the user's entire human body. It should be noted that Figure 4 is only a schematic diagram obtained after visualizing the three-dimensional model of at least one human body part of the user in this solution. The second information deployed in the first device can be used to represent the three-dimensional model of at least one human body part of the user. In addition, the example in Figure 4 is only for the convenience of understanding the concept of the three-dimensional model of at least one human body part of the user and is not intended to limit this solution.

[0090] Alternatively, the three-dimensional model of at least one human body part of the user can be a three-dimensional model pre-stored in the first device. Optionally, the aforementioned three-dimensional model can be obtained by three-dimensional modeling based on the size parameters of at least one human body part of the user. The aforementioned size parameters may include at least one of the following: height, weight, head circumference, neck circumference, shoulder width, arm length, wrist circumference, leg length, ankle circumference, foot length, or other size parameters, etc. The specific parameters can be determined in combination with the actual application scenario. The example here is only for the convenience of understanding this solution.

[0091] Alternatively, the 3D model of at least one human body part of the user can be obtained by 3D modeling based on real-shot images other than those in step 201. Optionally, before performing step 201, the first device can pre-acquire multiple real-shot images of at least one human body part of the user. These multiple real-shot images can be 2D images obtained by a camera capturing at least one human body part of the user from multiple shooting angles. 3D modeling can be performed using these multiple real-shot images of at least one human body part of the user to obtain the 3D model of at least one human body part of the user. It should be noted that the first device can also use other methods to obtain the 3D model of at least one human body part of the user. The example here is only to demonstrate the feasibility of this solution. The specific method of obtaining the 3D model of at least one human body part of the user can be determined in combination with the actual application scenario.

[0092] Regarding the specific implementation of the first device generating at least one first image corresponding one-to-one with the aforementioned at least one real-shot image based on a 3D model of at least one human body part of the user and a 3D model of the target object, optionally, the first device can determine correspondence information, which indicates which of the multiple second vertices of the multiple first vertices constituting the 3D model of the target object corresponds to each of the multiple second vertices constituting the 3D model of at least one human body part of the user; furthermore, the first device can determine offset information corresponding to the 3D model of the target object, which can be offset information in 3D space, for example, the offset information in 3D space can include offsets on the x-axis, y-axis and z-axis; the first device updates the first position information of each first vertex according to the offset information corresponding to the 3D model of the target object, to obtain the updated first position information of each first vertex. The aforementioned steps simulate moving the 3D model of the target object onto the 3D model of at least one human body part of the user.

[0093] The first device uses a physics simulation engine to obtain at least one wearable 3D model corresponding to the target item worn by the user in 3D space, based on the first position information of each first vertex (or the updated first position information of each first vertex) and the second position information of each second vertex. Each wearable 3D model may include: a 3D model of at least one human body part of the user and third information.

[0094] For example, at least one third piece of information corresponds one-to-one with at least one real-shot image obtained in step 201. The third piece of information corresponding to each real-shot image may include: the third position information of each first vertex of the three-dimensional model that makes up the target item when the user wears the target item in the form shown in the real-shot image in three-dimensional space.

[0095] The physical simulation in this application can also be referred to as three-dimensional simulation. For example, the physical simulation engine may use the finite element method (FEM), finite integral in time domain (FDID) method, or other types of physical simulation algorithms.

[0096] Furthermore, in one case, if the at least one real-shot image obtained in step 201 is at least one video frame in a video, the first device can generate at least one first image corresponding to at least one video frame based on a three-dimensional model of at least one human body part of the user corresponding to one of the at least one video frame, a three-dimensional model of the target object, and at least one video frame.

[0097] For example, the first device inputs the first position information of each first vertex (or the updated first position information of each first vertex), the second position information of all second vertices of the 3D model of at least one human body part of the user corresponding to one of the at least one video frame, and all video frames in the video into a physics simulation engine to obtain at least one third information output by the physics simulation engine that corresponds one-to-one with at least one video frame. In this embodiment, when the acquired at least one real-shot image of the user is at least one video frame in the video, since the 3D dimensions of the user are the same in different video frames, a 3D model of at least one human body part of the user can be obtained by performing 3D modeling based on only one of the at least one video frame. Then, using the 3D model of at least one human body part of the user corresponding to the aforementioned video frame, the 3D model of the target object, and at least one video frame, at least one first image corresponding one-to-one with at least one video frame is generated, without repeating 3D modeling for each video frame, thereby reducing the consumption of computer resources.

[0098] Alternatively, the first device can input the first position information (or the updated first position information of each first vertex) and the second position information of all second vertices included in the 3D model of at least one human body part of at least one user corresponding to at least one video frame into the physics simulation engine to obtain at least one third piece of information output by the physics simulation engine corresponding to at least one video frame. For example, the physics simulation engine can not only simulate the changes in the position of the vertices of the 3D model of the target object under the influence of gravity and human body size, but also simulate the dynamic effects of the 3D model of the target object across multiple video frames.

[0099] In another scenario, if the real-shot image of at least one user obtained in step 201 is at least one independently captured real-shot image, the first device can input the first position information of each first vertex (or the updated first position information of each first vertex) and the second position information of all second vertices included in the three-dimensional model of at least one human body part of at least one user corresponding to at least one independently captured real-shot image into the physics simulation engine to obtain at least one third information corresponding to at least one independently captured real-shot image. For example, the physics simulation engine can simulate the changes in the position of the vertices of the three-dimensional model of the target object under the influence of gravity and human body size.

[0100] After acquiring at least one third piece of information corresponding to at least one real-shot image of at least one user, the first device can perform a rendering operation using a rendering engine based on the second information of the three-dimensional model of at least one human body part of the user corresponding to each real-shot image and the third information corresponding to each real-shot image, to obtain a first image corresponding to each user's real-shot image; the first device repeats the aforementioned steps at least once to obtain at least one first image corresponding to at least one real-shot image of at least one user obtained in step 201.

[0101] The phrase "rendering a 2D image from a 3D model of a wearable device" can be understood as projecting the 3D model onto a 2D screen. To put it more vividly, it can be understood as taking a photograph of the 3D model in 3D space. In other words, it is the process of projecting the 3D model onto a real-world image according to a pre-set camera position and shooting angle.

[0102] Optionally, the texture, material, or color of the target object in the first image can be obtained based on first information used to represent the three-dimensional model of the target object; the skin color of the user in the first image can be obtained based on second information used to represent at least one part of the user's body.

[0103] Optionally, the camera position set in the rendering is based on the camera position of the user's real-shot image obtained in step 201, and / or the shooting angle set in the rendering is based on the shooting angle of the user's real-shot image obtained in step 201. This helps to ensure that the camera position and / or shooting angle used during rendering are consistent with the camera position and / or shooting angle when capturing the user's real-shot image, thereby making the camera position and / or shooting angle corresponding to the first image consistent with the camera position and / or shooting angle of the user's real-shot image. That is, alignment between the first image and the user's real-shot image is achieved in the dimension of camera position and / or shooting angle, which helps to further improve the realism of the updated virtual wearable target item.

[0104] In one scenario, if the at least one real-shot image obtained in step 201 is at least one video frame in a video, the first device may optionally obtain the camera position and / or the shooting angle of the first video frame, regard the camera position of the first video frame as the camera position corresponding to each video frame, regard the shooting angle of the first video frame as the shooting angle of each video frame, or replace the first video frame with other video frames in at least one video frame, etc.

[0105] To further understand this solution, for example, the at least one video frame obtained in step 201 includes 220 video frames, wherein the camera position when the first video frame is captured is position 0, and the shooting angle when the first video frame is captured is angle 0. Therefore, the camera position used when performing the rendering operation to obtain the first image corresponding to each of the aforementioned 220 video frames is position 0, and the shooting angle used is angle 0. It should be understood that the example here is only for the convenience of understanding this solution and is not intended to limit this solution.

[0106] In another scenario, if at least one real-shot image obtained in step 201 is at least one independently captured real-shot image, the first device can determine the camera position when capturing each real-shot image as the camera position corresponding to that real-shot image when performing the rendering operation; and determine the shooting angle when capturing each real-shot image as the shooting angle corresponding to that real-shot image when performing the rendering operation.

[0107] To further understand this solution, exemplarily, the at least one independently captured real-shot image obtained in step 201 includes real-shot image 1, real-shot image 2, and real-shot image 3. Wherein, the camera position when capturing real-shot image 1 is position 1, and the shooting angle when capturing real-shot image 1 is angle 1; the camera position when capturing real-shot image 2 is position 2, and the shooting angle when capturing real-shot image 2 is angle 2; the camera position when capturing real-shot image 3 is position 3, and the shooting angle when capturing real-shot image 3 is angle 3. Therefore, the camera position used when performing the rendering operation to obtain the first image corresponding to real-shot image 1 can be position 1, and the shooting angle used can be angle 1; the camera position used when performing the rendering operation to obtain the first image corresponding to real-shot image 2 can be position 2, and the shooting angle used can be angle 2; the camera position used when performing the rendering operation to obtain the first image corresponding to real-shot image 3 can be position 3, and the shooting angle used can be angle 3. It should be understood that this example is only for the convenience of understanding this solution and is not intended to limit this solution.

[0108] Optionally, when performing the above-mentioned 3D modeling operation, the first device can obtain the camera position and shooting angle when capturing the actual image. For example, when performing 3D modeling based on the first video frame in at least one video frame, the first device can obtain the camera position and shooting angle when capturing the first video frame; correspondingly, when performing 3D modeling based on each independently captured actual image, the first device can obtain the camera position and shooting angle when capturing each actual image.

[0109] 203. Perform a fusion operation on the real-shot image and the first image to obtain a second image, which includes images of multiple human body parts wearing the target item.

[0110] For example, the first device obtains a second image corresponding to each user's real-shot image based on the real-shot image of each user in at least one user's real-shot image and a first image corresponding to each user's real-shot image. That is, each second image is a virtual wearable image obtained by fusing the real image (i.e., the user's real-shot image) and the virtual image (i.e., the first image).

[0111] For example, the above fusion operation may specifically include at least one of the following: cutting out and stitching the real-shot image and the first image respectively; cutting out and stitching the real-shot image and the first image respectively and then filling in the remaining parts; directly superimposing the real-shot image and the first image; using the real-shot image to enhance the first image; regenerating the second image using the real-shot image and the first image; or other fusion methods, etc., which will not be exhaustive here.

[0112] Further, if the aforementioned at least one user's real-shot image is at least one independently captured real-shot image, step 203 may include: the first device performing a fusion operation on each independently captured real-shot image and the first image corresponding to each independently captured real-shot image to obtain a second image corresponding to each independently captured real-shot image. Alternatively, if the aforementioned at least one user's real-shot image is at least one video frame, step 203 may include: the first device performing a fusion operation on each video frame and the first image corresponding to each video frame to obtain a second image corresponding to each video frame.

[0113] For example, in one scenario, step 203 may include: a first device sending each captured image and a first image corresponding to each captured image, and receiving a second image corresponding to each captured image, wherein the second image corresponding to each captured image is obtained by performing a fusion operation on each captured image and the first image corresponding to each captured image. For instance, the first device is a client device, which may send each captured image and the first image corresponding to each captured image to a cloud server and receive the second image corresponding to each captured image sent by the cloud server.

[0114] In another scenario, step 203 may include: the first device performing a fusion operation on each real-shot image and the first image corresponding to each real-shot image to generate a second image corresponding to each user's real-shot image. For example, if the first device is a client device or a cloud server, it may generate a second image. The specific method can be determined based on the actual application scenario.

[0115] The following describes the specific process of obtaining a second image corresponding to each real-shot image based on each real-shot image and a corresponding first image. Optionally, in one case, each second image can be obtained by stitching together a first image region in the real-shot image and a second image region in the first image. In other words, each second image is obtained by stitching together a first image region in each real-shot image and a corresponding second image region in the first image. The first image region includes at least one human body part not covered by an object of the same type as the target object, and the second image region includes the target object.

[0116] To more intuitively understand this solution, please refer to Figure 5. Figure 5 is a schematic diagram of a user's real-shot image, a first image, and a second image provided in an embodiment of this application. As shown in Figure 5, the user's real-shot image is a true image of the user wearing their own top. The first image is an image obtained by rendering the 3D model of the clothing to 2D. The second image is obtained by stitching the top area in the first image with the body parts not covered by the top in the real-shot image. It should be understood that the example in Figure 5 is only for the convenience of understanding this solution and is not intended to limit this solution. It should be noted that the user's face is obscured in Figure 5 and several subsequent figures in this embodiment of the application. The actual face will be displayed when shown to the user.

[0117] In this embodiment of the application, a method is provided to obtain a second image based on a real-shot image and a first image. The second image is obtained by stitching the first image area in the real-shot image with the second image area in the first image. The second image area in the first image can show whether the size of the target item is suitable when the user wears it. Thus, the second image can reflect whether the size of the target item is suitable. Since the first image area in the real-shot image can reflect the user's real situation, the realism of the second image is greatly improved, which improves the user's feeling of trying on the item in real life and is conducive to improving the user experience of this solution.

[0118] For example, in one implementation, the second image is obtained based on the user's actual photograph, the first image, first semantic information of the user's actual photograph, and second semantic information of the first image. The category of the target item the user is trying on is the first category. The first semantic information indicates the category of pixels in the user's actual photograph, thus the first semantic information can indicate pixels of the first category in the actual photograph. The second semantic information indicates the category of pixels in the first image, thus the second semantic information can indicate pixels of the first category in the first image. In other words, the first semantic information indicates which pixels in the user's actual photograph belong to the first category, and the second semantic information indicates which pixels in the first image belong to the first category.

[0119] Furthermore, the first device obtains a second image based on each captured image, the corresponding first image, the first semantic information of the captured image, and the second semantic information of the first image. This can include: the first device sending each captured image, the corresponding first image, the first semantic information of the captured image, and the second semantic information of the first image to other devices to obtain a second image corresponding to each user's captured image sent by the other devices. For example, when the first device is a client device, it can send the user's captured image, the first image, the first semantic information, and the second semantic information to a cloud server. The cloud server generates a second image based on the received information using a machine learning model, and the client device receives the second image sent by the cloud server.

[0120] Alternatively, the first device may obtain a second image based on each captured image, the corresponding first image, the first semantic information of the captured image, and the second semantic information of the first image. This may include: the first device inputting each captured image, the corresponding first image, the first semantic information of the captured image, and the second semantic information of the first image into a machine learning model to generate a second image corresponding to each captured image. For example, when the first device is a client device or a cloud server, a locally deployed machine learning model can also be used to generate the second image.

[0121] In one scenario, a large model for image processing can be used to generate a second image. For example, a second prompt message, the user's real-life image, and the first image are input into the large model for image processing to obtain a second image output by the large model. The second prompt message can be used to prompt the fusion of the area of ​​the target item included in the first image into the user's real-life image. The second prompt message may also include first semantic information and second semantic information.

[0122] Alternatively, a machine learning model specifically designed for image enhancement can be used to process the user's real-shot image to obtain the second image. The function of the aforementioned machine learning model is to fuse the region of the target item included in the reference image (e.g., the first image) into the image to be processed (e.g., the user's real-shot image). The aforementioned machine learning model can be a convolutional neural network, a fully connected neural network, a residual neural network, or other types of neural networks, etc.

[0123] Alternatively, before stitching the first image region in each captured image with the corresponding second image region in the first image, the first device can perform semantic recognition on each captured image to obtain first semantic information for each captured image, which indicates the category of pixels in the captured image; based on the first semantic information, determine the first image region in each captured image, where the category of pixels in the first image region is different from the category of the target item; perform semantic recognition on each first image to obtain second semantic information for each first image, which indicates the category of pixels in the first image; based on the second semantic information, determine the second image region in each first image, where the category of pixels in the second image region is consistent with the category of the target item; then the first device can stitch the first image region in each captured image with the corresponding second image region in the first image to obtain the second image.

[0124] In this embodiment, first semantic information and second semantic information are also obtained. The first semantic information indicates the category of pixels in the real-shot image. The first image region in the real-shot image can be determined based on the first semantic information. The second semantic information indicates the category of pixels in the first image. The second image region in the first image can be determined based on the second semantic information. This reduces the difficulty of the stitching step and helps to improve the quality when performing the stitching step, so as to obtain a second image of better quality.

[0125] For example, the first semantic information can be obtained by performing semantic recognition on the real-shot image. For instance, the first semantic information of the real-shot image can be identified using human parsing technology. Correspondingly, the second semantic information can be obtained by performing semantic recognition on the first image.

[0126] The following description, in conjunction with illustrations, explains the principle of stitching together a first image region and a second image region from a real-shot image. Since the executing entity can be either a first device or a cloud server connected to the first device, the executing entity is omitted in the subsequent description. For example, pixels belonging to the first category in the user's real-shot image can be determined based on first semantic information, and items of the first category can be removed from the user's real-shot image. For instance, the color of pixels wearing items of the first category in the user's real-shot image can be set as the background color to remove items of the first category from the user's real-shot image. For a more intuitive understanding of this solution, please refer to Figure 6. Figure 6 is a real-shot image provided in an embodiment of this application and a schematic diagram showing the removal of the first image region from the real-shot image. In Figure 6, taking a top as an example of the target item category, in the right sub-schematic diagram of Figure 6, a portion of the user's real-shot image wearing a top (i.e., an example of the first image region) has been removed. It should be understood that the example in Figure 6 is only for ease of understanding and is not intended to limit this solution.

[0127] Pixels belonging to the first category in the first image can be determined based on the second semantic information. A partial image including pixels belonging to the first category can be obtained from the first image. Please refer to Figure 7. Figure 7 is a schematic diagram of obtaining a second image region from the first image according to an embodiment of this application. The left sub-schematic diagram of Figure 7 shows the first image. The black area in the middle sub-schematic diagram of Figure 7 indicates pixels belonging to the category of "shirt". The right sub-schematic diagram of Figure 7 shows a partial image including pixels belonging to the category of "shirt" (i.e., an example of the second image region). It should be understood that the example in Figure 7 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0128] Then, the real-life image of the items in the first category removed and the local image including the pixels of the first category can be combined to obtain the second image. A schematic diagram of the second image can be seen in the second image shown in Figure 5 above, which will not be described again here.

[0129] In another implementation, the second image can also be generated based on the third prompt information, the real-shot image, and the first image, using a large model for image processing. For example, the third prompt information can prompt the fusion of the target object region included in the first image with at least one human body part in the real-shot image that is not covered by an object of the same type as the target object, and the third prompt information does not include the first semantic information and the second semantic information.

[0130] Optionally, the first semantic information may also indicate pixels in the real-world image categorized as torso, arms, hands, legs, and / or feet. For example, the first semantic information of the real-world image can be identified using human parsing and dense pose recognition techniques. If the first image region of pixels of the first category (i.e., the category of the target object) in the real-world image is larger than the second image region of pixels of the first category in the first image, image enhancement can also be performed on blank areas in the real-world image based on the first semantic information. The blank areas in the real-world image are the first image region minus the second image region. In other words, after stitching the first image region in the real-world image with the second image region in the first image, image enhancement can also be performed on blank areas in the real-world image based on the first semantic information.

[0131] For example, if pixels in the first image region of a real-shot image are removed, and the user's real-shot image after removing the first image region is stitched together with a partial image including the second image region, there will still be blank areas in the user's real-shot image (that is, the area of ​​the first image region minus the second image region). Then, based on the semantic category of the pixels in the blank area, image enhancement can be performed on the pixels in the blank area to further improve the realism of the second image.

[0132] For example, if the first category is a top, and the user is wearing long sleeves in the real-life image, while the top in the first image is short sleeves, then after compositing the real-life image of the user with the long sleeves removed and the local image including the pixels of the short sleeves, a blank area will appear in the user's real-life image. Then, the semantic category of the pixels in the blank area (i.e., the arm) can be used to enhance the pixels in the blank area, making the blank area look more like an arm. This example is only for the convenience of understanding this solution and is not intended to limit this solution.

[0133] Optionally, in another scenario, the first image region in the captured image includes the user's face region, and each second image is obtained by enhancing the exposed area of ​​the user in the first image using the face region from the corresponding captured image. The exposed area in the first image includes the face region from the first image. Optionally, the exposed area in the first image may also include exposed skin areas, exposed hair areas, or other types of exposed areas, which can be determined based on the actual application scenario.

[0134] To more intuitively understand this solution, please refer to Figure 8. Figure 8 is another schematic diagram of the face region, the first image, and the second image in the real-shot image provided in the embodiment of this application. The left sub-schematic diagram in Figure 8 shows the real image of the user's head included in the real-shot image. The first image is the image obtained by rendering the wearable 3D model to 2D. The wearable 3D model is obtained by combining the 3D model of the user's entire body and the 3D model of the long skirt. The second image is the second image obtained by enhancing the exposed area in the first image based on the face region in the real-shot image. The face region in the second image is more similar to the user's real face, and the exposed skin color in the second image is also more similar to the user's real skin tone. The hair color in the second image is also more similar to the user's real hair color. The user's face is occluded in Figure 8, but the face will be shown when the second image is actually displayed to the user. It should be understood that the example in Figure 8 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0135] In this embodiment of the application, an alternative approach is provided to obtain a second image based on the user's real-shot image and a first image, which improves the implementation flexibility of this solution. In addition, the first image is enhanced by using the facial region in the user's real-shot image to obtain the second image, thereby making the second image more similar to the real user, and the second image can more accurately reflect whether the size of the target item is appropriate.

[0136] For example, in one implementation, the second image is obtained based on the user's real-shot image, the first image, third semantic information of the user's real-shot image, and fourth semantic information of the first image. The third semantic information indicates the category of pixels in the user's real-shot image, and the fourth semantic information indicates the category of pixels in the first image. The third semantic information is used to indicate pixels in the facial region of the real-shot image, and the fourth semantic information is used to indicate pixels in the user's exposed area of ​​the first image. In other words, the third semantic information indicates which pixels in the user's real-shot image are facial regions, and the fourth semantic information indicates which pixels in the first image are the user's exposed areas. The first device obtains the second image corresponding to each user's real-shot image based on each user's real-shot image and the corresponding first image, including: the first device obtaining the second image corresponding to each user's real-shot image based on each user's real-shot image, the corresponding first image, the third semantic information of the user's real-shot image, and the fourth semantic information of the first image.

[0137] Similar to the description in the previous case, the first device can send the user's real-shot image, the first image, the third semantic information, and the fourth semantic information to other devices, and receive the second image sent by other devices; or, the first device can also generate the second image locally. The above two cases will not be described in detail here.

[0138] In one scenario, a large image processing model can be used to generate the second image. For example, the fourth prompt information, the user's actual photo, and the first image are input into the large image processing model to obtain the second image output by the large model. The fourth prompt information can be used to prompt the user to enhance the user's exposed area in the first image using the facial region in the user's actual photo. To further prompt the user about the facial region in the actual photo and the user's exposed area in the first image, the fourth prompt information may also include third semantic information and fourth semantic information, and / or, the fourth prompt information may also include a partial image of the facial region in the user's actual photo obtained based on the third semantic information and an image of the user's exposed area in the second image obtained based on the fourth semantic information.

[0139] Alternatively, the second image can be obtained by processing the user's real-life image using a machine learning model specifically designed for image enhancement. The function of the aforementioned machine learning model is to enhance the image to be processed (e.g., the first image) by using the facial region included in the reference image (e.g., the user's real-life image) to include the user's exposed areas.

[0140] Optionally, before using the user's face region in the real-shot image to enhance the exposed area in the first image, the first device may also perform semantic recognition on the real-shot image to obtain third semantic information of the real-shot image, the third semantic information indicating the category of pixels in the real-shot image; determine the user's face region in the real-shot image based on the third semantic information, the category of pixels in the user's face region in the real-shot image being "face"; perform semantic recognition on the first image to obtain fourth semantic information of the first image, the fourth semantic information indicating the category of pixels in the first image; determine the exposed area in the first image based on the fourth semantic information, the category of pixels in the exposed area including "face". Therefore, step 203 may include: the first device using the user's face region in each real-shot image to enhance the exposed area in the corresponding first image to obtain a second image corresponding to each real-shot image.

[0141] The following description, in conjunction with illustrations, explains the principle of enhancing the exposed area in the first image using the user's face region in the real-shot image. Since the executing entity can be either the first device or a cloud server connected to the first device, the executing entity is omitted in the following description. For example, pixels classified as "face" in the user's real-shot image can be determined based on third semantic information, and a local image including pixels classified as "face" can be obtained from the user's real-shot image.

[0142] The pixels included in the user's exposed area in the first image can also be determined based on the fourth semantic information. The user's exposed area can also be understood as the area not covered by the target item. An image of the user's exposed area can be obtained from the first image. For example, the pixels of the user's exposed area in the first image can be set to white, and the other pixels in the first image can be set to black, so as to obtain an image of the user's exposed area from the first image. To understand this solution more intuitively, please refer to Figure 9. Figure 9 is a schematic diagram of obtaining an image of the user's exposed area from the first image according to an embodiment of this application. The left sub-schematic diagram of Figure 9 shows the first image. As shown in the left sub-schematic diagram of Figure 9, the target item that the user is trying on is a skirt. The user's exposed area in the first image (that is, the area not covered by the skirt) includes the head, neck, arms, ankles, and feet. The right sub-schematic diagram of Figure 9 shows an image including the user's exposed area in the first image. It should be understood that the example in Figure 9 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0143] The machine learning model takes a local image including pixels categorized as face, an image of the user's exposed area obtained from the first image, and the first image as input. The machine learning model then performs image enhancement on the user's exposed area in the first image to obtain the second image. A schematic diagram of the aforementioned second image can be found in Figure 8 above, and will not be elaborated here.

[0144] Optionally, the third semantic information further indicates the pixels of the background region in the user's real-shot image, and the fourth semantic information further indicates the pixels of the background region in the first image. Each second image is obtained by enhancing the first image using the face region in the real-shot image and enhancing the background region of the first image using the background region in the real-shot image. In this embodiment, the third and fourth semantic information can also be used to enhance the background region in the first image based on the background region in the real-shot image, thereby further improving the realism of the second image and creating a more realistic feeling for the user wearing the target item, thus improving the user experience of this solution.

[0145] For example, a large image processing model can be used to generate the second image. For example, the fifth prompt message, the user's actual photograph, and the first image are input into the large image processing model to obtain the second image output by the large model. The fifth prompt message can be used to prompt image enhancement of the user's exposed area in the first image using the facial region in the user's actual photograph, and to prompt image enhancement of the exposed background region in the first image using the background region in the user's actual photograph. The fifth prompt message may also include third and fourth semantic information. It should be noted that other methods can also be used to implement the aforementioned functions, which are not exhaustively described in this application.

[0146] In this embodiment, the third semantic information can be used to quickly and accurately locate the facial region in the user's real-shot image, and the fourth semantic information can be used to quickly and accurately locate the user's exposed area in the first image. Thus, by using the third and fourth semantic information, the first image can be enhanced using the facial region in the user's real-shot image, which helps to reduce the difficulty of the aforementioned image enhancement part and helps to obtain a second image of better quality.

[0147] In another implementation, the second image may also be generated based on the sixth prompt information, the user's real-shot image, and the first image, using a large model for image processing. For example, the sixth prompt information may include a text description prompting the use of the user's face region in the real-shot image to enhance the user's exposed area in the first image (optionally, it may also include a text description prompting the use of the background region in the real-shot image to enhance the exposed background region in the first image), and the sixth prompt information does not include the first semantic information and the second semantic information.

[0148] Optionally, if the video obtained in step 201 is at least one video frame, then after the first device obtains the second image corresponding to each video frame, it can compose the user's virtual wearable video from the second images corresponding to all video frames in the at least one video frame; it can also perform inter-frame consistency processing on the user's virtual wearable video to obtain an updated user's virtual wearable video.

[0149] For example, the large model used for image processing can be used to perform inter-frame consistency processing on the aforementioned user's virtual wearable video. For instance, the seventh prompt information and the aforementioned user's virtual wearable video are input into the large model used for image processing to obtain the updated user's virtual wearable video output by the large model. The seventh prompt information can be used to prompt for improved inter-frame consistency between different video frames. The role of performing inter-frame consistency processing can be understood as improving the similarity between adjacent video frames, so that the transition between different video frames in the updated user's virtual wearable video frames is smoother, and after being combined into a virtual wearable video, it can also display the dynamic effect of the user wearing the target item.

[0150] Alternatively, a machine learning model specifically designed to improve the inter-frame consistency of video can be used to process the aforementioned user's virtual wearable video to obtain an updated version of the user's virtual wearable video. In other words, the function of the aforementioned machine learning model only includes improving the inter-frame consistency between different video frames.

[0151] In this implementation, since the dimensions of the target item and at least one human body part of the user can be fully reflected in three-dimensional space, a wearable three-dimensional model is obtained by combining the three-dimensional model of at least one human body part and the three-dimensional model of the target item. This wearable three-dimensional model can reflect whether the size of the target item is suitable for the user. The wearable three-dimensional model is rendered into two dimensions to obtain the first image. The first image retains the information on whether the size of the target item is suitable for the user. The real-shot image contains more realistic human body parts. The second image is obtained by performing a fusion operation on the real-shot image and the first image. Therefore, the second image can not only show whether the size of the target item is suitable for the user, but also better restore the user's real situation.

[0152] Based on the embodiments corresponding to Figures 1 to 9, in order to better implement the above-mentioned solutions of the embodiments of this application, related equipment for implementing the above solutions is also provided below. Specifically, referring to Figure 10, Figure 10 is a structural schematic diagram of a wearable image acquisition device provided in an embodiment of this application. The wearable image acquisition device 1000 includes: an acquisition module 1001, used to acquire a real-shot image, the real-shot image being a two-dimensional image captured by a camera, the real-shot image including multiple human body parts of the user; the acquisition module 1001 is also used to acquire a first image, the first image including a two-dimensional image obtained by rendering a wearable three-dimensional model, the wearable three-dimensional model being obtained by combining a three-dimensional model of at least one human body part among multiple human body parts with a three-dimensional model of a target item; and a fusion module 1002, used to perform a fusion operation on the real-shot image and the first image to obtain a second image, the second image including an image of multiple human body parts wearing the target item.

[0153] Optionally, the fusion module 1002 is specifically used to stitch together a first image region in the real-shot image with a second image region in the first image to obtain a second image, wherein the first image region includes at least one human body part not covered by an object of the same type as the target object, and the second image region includes the target object.

[0154] Optionally, the wearable image acquisition device 1000 further includes: a recognition module 1003, used to perform semantic recognition on the real-shot image to obtain first semantic information of the real-shot image, the first semantic information indicating the category of pixels in the real-shot image; a determination module 1004, used to determine a first image region in the real-shot image based on the first semantic information, the category of pixels in the first image region being different from the category of the target item; the recognition module 1003 is also used to perform semantic recognition on the first image to obtain second semantic information of the first image, the second semantic information indicating the category of pixels in the first image; the determination module 1004 is also used to determine a second image region in the first image based on the second semantic information, the category of pixels in the second image region being consistent with the category of the target item.

[0155] Optionally, the first image region in the real-shot image includes the user's face region, and the second image is obtained by enhancing the exposed area in the first image using the user's face region in the real-shot image, and the exposed area includes the face region.

[0156] Optionally, the wearable image acquisition device 1000 further includes: a recognition module 1003, used to perform semantic recognition on the real-shot image to obtain third semantic information of the real-shot image, the third semantic information indicating the category of pixels in the real-shot image; a determination module 1004, used to determine the user's face region in the real-shot image based on the third semantic information, the category of pixels in the user's face region in the real-shot image being face; the recognition module 1003 is also used to perform semantic recognition on the first image to obtain fourth semantic information of the first image, the fourth semantic information indicating the category of pixels in the first image; the determination module 1004 is also used to determine the exposed area in the first image based on the fourth semantic information, the category of pixels in the exposed area including face.

[0157] Optionally, the acquisition module 1001 is specifically used for: acquiring a 3D model of at least one human body part of the user, the 3D model of at least one human body part of the user is obtained by 3D modeling based on real-shot images; acquiring a 3D model of the target item; and obtaining a wearable 3D model based on the 3D model of at least one human body part of the user and the 3D model of the target item, the first image including an image obtained by rendering the wearable 3D model to 2D.

[0158] Optionally, the real-shot image is at least one video frame in the real-shot video. The acquisition module is specifically used to generate at least one first image corresponding to at least one video frame based on a three-dimensional model of at least one human body part, a three-dimensional model of the target object, and at least one video frame. The three-dimensional model of at least one human body part is obtained by performing three-dimensional modeling on one of the video frames in the at least one video frame.

[0159] Optionally, multiple body parts are included in the torso and / or head of the human body.

[0160] Optionally, the camera position set in the rendering is obtained based on the camera position of the actual captured image, and / or, the shooting angle set in the rendering is obtained based on the shooting angle of the actual captured image.

[0161] The acquisition module 1001 and the fusion module 1002 (optionally, also including the identification module 1003 and the determination module 1004) can be implemented in software or in hardware. For example, the implementation of the acquisition module 1001 will be described below. Similarly, the implementation of the fusion module 1002, the identification module 1003, and the determination module 1004 can refer to the implementation of the acquisition module 1001.

[0162] As an example of a software functional unit, module 1001 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, module 1001 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed within the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0163] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0164] As an example of a hardware functional unit, the acquisition module 1001 may include at least one computing device, such as a server. Alternatively, the acquisition module 1001 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0165] The multiple computing devices included in the acquisition module 1001 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the acquisition module 1001 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the acquisition module 1001 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0166] It should be noted that, in other embodiments, the acquisition module 1001 can be used to execute any step in the method for acquiring wearable images, the fusion module 1002 can be used to execute any step in the method for acquiring wearable images, the recognition module 1003 can be used to execute any step in the method for acquiring wearable images, and the determination module 1004 can be used to execute any step in the method for acquiring wearable images. The steps implemented by the acquisition module 1001 and the fusion module 1002 (optionally, also including the recognition module 1003 and the determination module 1004) can be specified as needed. The acquisition module 1001 and the fusion module 1002 (optionally, also including the recognition module 1003 and the determination module 1004) respectively implement different steps in the method for acquiring wearable images to realize all the functions of the wearable image acquisition device 1000.

[0167] This application also provides a system for acquiring wearable images. Figure 11 is a schematic diagram of a system for acquiring wearable images provided in an embodiment of this application. As shown in Figure 11, the system for acquiring wearable images includes: a device for acquiring wearable images 1000 and a device for acquiring wearable images 1100.

[0168] The wearable image acquisition device 1100 is used to send real-shot images to the wearable image acquisition device 1000;

[0169] The wearable image acquisition device 1000 is used to acquire a real-shot image, which is a two-dimensional image captured by a camera and includes multiple body parts of the user; acquire a first image, which includes a two-dimensional image obtained by rendering a wearable three-dimensional model, which is obtained by combining a three-dimensional model of at least one body part among multiple body parts with a three-dimensional model of a target item; and perform a fusion operation on the real-shot image and the first image to obtain a second image, which includes an image of multiple body parts wearing the target item.

[0170] The wearable image acquisition device 1100 is used to receive a second image sent by the wearable image acquisition device 1000.

[0171] Both the wearable image acquisition device 1000 and the wearable image acquisition device 1100 can be implemented in software or in hardware. For example, the implementation of the wearable image acquisition device 1000 will be described below. Similarly, the implementation of the wearable image acquisition device 1100 can be referred to the implementation of the wearable image acquisition device 1000.

[0172] As an example of a software functional unit, the wearable image acquisition device 1000 may include code running on a computing instance. The computing instance may be at least one of a physical host (computing device), a virtual machine, a container, or other computing devices. Further, the aforementioned computing device may be one or more. For example, the wearable image acquisition device 1000 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the application may be distributed in the same region or in different regions. The multiple hosts / virtual machines / containers used to run the code may be distributed in the same Availability Zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0173] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same VPC or across multiple VPCs. Typically, a VPC is set up within a single region. Communication between two VPCs within the same region, and between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0174] As an example of a hardware functional unit, the wearable image acquisition device 1000 may include at least one computing device, such as a server. Alternatively, the wearable image acquisition device 1000 may also be a device implemented using an ASIC or a PLD. The PLD may be implemented using a CPLD, FPGA, GAL, or any combination thereof.

[0175] The wearable image acquisition device 1000 includes multiple computing devices that can be distributed in the same region or in different regions. Similarly, the multiple computing devices can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0176] This application also provides a computing device 100. As shown in FIG12, FIG12 is a schematic diagram of the structure of a computing device provided in an embodiment of this application. The computing device 100 includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other through the bus 102. The computing device 100 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 100.

[0177] Bus 102 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 12, but this does not imply that there is only one bus or one type of bus. Bus 104 can include pathways for transmitting information between various components of computing device 100 (e.g., memory 106, processor 104, communication interface 108).

[0178] The processor 104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0179] Memory 106 may include volatile memory, such as random access memory (RAM). Processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0180] The memory 106 stores executable program code, and the processor 104 executes the executable program code to implement the functions of the aforementioned acquisition module and fusion module, respectively. Optionally, the processor 104 executes the executable program code to implement the functions of the aforementioned recognition module and determination module, thereby realizing the method for acquiring wearable images. That is, the memory 106 stores instructions for executing the method for acquiring wearable images.

[0181] Alternatively, the memory 106 stores executable code, which the processor 104 executes to implement the functions of the aforementioned wearable image acquisition device 1000 and wearable image acquisition device 1100, thereby realizing the wearable image acquisition method. That is, the memory 106 stores instructions for executing the wearable image acquisition method.

[0182] The communication interface 103 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 100 and other devices or communication networks.

[0183] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0184] As shown in Figure 13, which is a schematic diagram of a computing device cluster provided in an embodiment of this application, the computing device cluster includes at least one computing device 100. The memory 106 in one or more computing devices 100 in the computing device cluster may store the same instructions for executing a method for acquiring wearable images.

[0185] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the method of acquiring wearable images. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for executing the method of acquiring wearable images.

[0186] It should be noted that the memory 106 in different computing devices 100 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the wearable image acquisition device 1000. That is, the instructions stored in the memory 106 of different computing devices 100 can implement the functions of one or more modules among the acquisition module and fusion module (optionally, also including an identification module and a determination module).

[0187] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 14 illustrates one possible implementation, and is a schematic diagram of computer devices in a computer cluster provided in this application embodiment being connected via a network. As shown in Figure 14, two computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 106 in computing device 100A stores instructions for executing the functions of the acquisition module and the fusion module. Simultaneously, the memory 106 in computing device 100B stores instructions for executing the functions of the identification module and the determination module.

[0188] The connection method between the computing device clusters shown in Figure 14 can be considered as follows: taking into account that the wearable image acquisition method provided in this application needs to independently identify the semantic information in the image and manage the two functions of image recognition and image fusion separately, it is considered that the functions implemented by the recognition module and the determination module are executed by the computing device 100B.

[0189] It should be understood that the functions of computing device 100A shown in Figure 14 can also be performed by multiple computing devices 100. Similarly, the functions of computing device 100B can also be performed by multiple computing devices 100.

[0190] This application embodiment also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection method of the computing device cluster described in Figures 13 and 14. The difference is that the memory 106 in one or more computing devices 100 in this computing device cluster can store the same instructions for executing the method of acquiring wearable images.

[0191] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the method of acquiring wearable images. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for executing the method of acquiring wearable images.

[0192] It should be noted that the memory 106 in different computing devices 100 within the computing device cluster can store different instructions for executing some functions of the wearable image acquisition system. That is, the instructions stored in the memory 106 of different computing devices 100 can implement the functions of one or more of the wearable image acquisition devices 1000 and 1100.

[0193] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform a method for acquiring wearable images.

[0194] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to perform a method for acquiring wearable images.

[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for acquiring wearable images, characterized in that, The method includes: Acquire real-shot images, which are two-dimensional images captured by a camera, and the real-shot images include multiple body parts of the user; Acquire a first image, the first image including a two-dimensional image obtained by rendering a wearable three-dimensional model, wherein the wearable three-dimensional model is obtained by combining a three-dimensional model of at least one human body part among the plurality of human body parts with a three-dimensional model of the target item; The real-shot image and the first image are fused to obtain a second image, which includes images of the target item worn on the multiple human body parts.

2. The method according to claim 1, characterized in that, The step of fusing the real-shot image and the first image to obtain the second image includes: The second image is obtained by stitching together a first image region from the real-shot image and a second image region from the first image. The first image region includes at least one human body part that is not covered by an object of the same type as the target object, and the second image region includes the target object.

3. The method according to claim 2, characterized in that, Before stitching the first image region in the real-shot image with the second image region in the first image to obtain the second image, the method further includes: Semantic recognition is performed on the real-shot image to obtain the first semantic information of the real-shot image, and the first semantic information indicates the category of the pixel in the real-shot image; Based on the first semantic information, the first image region in the real-shot image is determined, and the category of the pixels in the first image region is different from the category of the target item; Semantic recognition is performed on the first image to obtain second semantic information of the first image, and the second semantic information indicates the category of the pixels in the first image; The second image region in the first image is determined based on the second semantic information, and the category of the pixels in the second image region is consistent with the category of the target item.

4. The method according to claim 1, characterized in that, The first image region in the real-shot image includes the user's face region, and the second image is obtained by enhancing the exposed area in the first image using the user's face region in the real-shot image, the exposed area including the face region.

5. The method according to claim 4, characterized in that, Before performing a fusion operation between the first image region in the real-shot image and the first image to obtain the second image, the method further includes: Semantic recognition is performed on the real-shot image to obtain the third semantic information of the real-shot image, and the third semantic information indicates the category of the pixels in the real-shot image; Based on the third semantic information, the user's facial region in the real-shot image is determined, and the pixel type of the user's facial region in the real-shot image is face; Semantic recognition is performed on the first image to obtain fourth semantic information of the first image, wherein the fourth semantic information indicates the category of the pixels in the first image; The exposed area in the first image is determined based on the fourth semantic information, and the category of the pixels in the exposed area includes face.

6. The method according to any one of claims 1 to 5, characterized in that, The acquisition of the first image, which includes an image obtained by rendering the wearable 3D model to 2D, includes: A three-dimensional model of at least one human body part of the user is obtained, and the three-dimensional model of at least one human body part of the user is obtained by three-dimensional modeling based on the real-shot image; Obtain a 3D model of the target item; Based on the three-dimensional model of at least one human body part of the user and the three-dimensional model of the target item, the wearable three-dimensional model is obtained, and the first image includes an image obtained by rendering the wearable three-dimensional model into two dimensions.

7. The method according to any one of claims 1 to 6, characterized in that, The real-shot image is at least one video frame from the real-shot video, and the acquisition of the first image includes: Based on the three-dimensional model of the at least one human body part, the three-dimensional model of the target object, and the at least one video frame, at least one first image corresponding to the at least one video frame is generated, wherein the three-dimensional model of the at least one human body part is obtained by performing three-dimensional modeling on one of the at least one video frames.

8. The method according to any one of claims 1 to 7, characterized in that, The plurality of human body parts are included in the torso and / or head of the human body.

9. The method according to any one of claims 1 to 8, characterized in that, The camera position set in the rendering is obtained based on the camera position of the actual captured image, and / or the shooting angle set in the rendering is obtained based on the shooting angle of the actual captured image.

10. A device for acquiring wearable images, characterized in that, The device includes: The acquisition module is used to acquire real-shot images, which are two-dimensional images captured by a camera and include multiple body parts of the user. The acquisition module is further configured to acquire a first image, the first image including a two-dimensional image obtained by rendering the wearable three-dimensional model, wherein the wearable three-dimensional model is obtained by combining a three-dimensional model of at least one of the plurality of human body parts with a three-dimensional model of the target item. The fusion module is used to perform a fusion operation on the real-shot image and the first image to obtain a second image, the second image including images of the multiple human body parts wearing the target item.

11. The apparatus according to claim 10, characterized in that, The fusion module is specifically used to stitch together a first image region in the real-shot image with a second image region in the first image to obtain a second image, wherein the first image region includes at least one human body part not covered by an object of the same type as the target object, and the second image region includes the target object.

12. The apparatus according to claim 11, characterized in that, The device further includes: The recognition module is used to perform semantic recognition on the real-shot image to obtain first semantic information of the real-shot image, wherein the first semantic information indicates the category of pixels in the real-shot image. The determining module is used to determine the first image region in the real-shot image based on the first semantic information, wherein the category of the pixels in the first image region is different from the category of the target item; The recognition module is further configured to perform semantic recognition on the first image to obtain second semantic information of the first image, wherein the second semantic information indicates the category of the pixels in the first image; The determining module is further configured to determine the second image region in the first image based on the second semantic information, wherein the category of the pixels in the second image region is consistent with the category of the target item.

13. The apparatus according to claim 10, characterized in that, The first image region in the real-shot image includes the user's face region, and the second image is obtained by enhancing the exposed area in the first image using the user's face region in the real-shot image, the exposed area including the face region.

14. The apparatus according to claim 13, characterized in that, The device further includes: The recognition module is used to perform semantic recognition on the real-shot image to obtain third semantic information of the real-shot image, wherein the third semantic information indicates the category of the pixels in the real-shot image; The determination module is used to determine the user's facial region in the real-shot image based on the third semantic information, wherein the pixel points in the user's facial region in the real-shot image are classified as faces; The recognition module is further configured to perform semantic recognition on the first image to obtain fourth semantic information of the first image, wherein the fourth semantic information indicates the category of the pixels in the first image; The determining module is further configured to determine the exposed area in the first image based on the fourth semantic information, wherein the category of the pixels in the exposed area includes face.

15. The apparatus according to any one of claims 10 to 14, characterized in that, The acquisition module is specifically used for: A three-dimensional model of at least one human body part of the user is obtained, and the three-dimensional model of at least one human body part of the user is obtained by three-dimensional modeling based on the real-shot image; Obtain a 3D model of the target item; Based on the three-dimensional model of at least one human body part of the user and the three-dimensional model of the target item, the wearable three-dimensional model is obtained, and the first image includes an image obtained by rendering the wearable three-dimensional model into two dimensions.

16. The apparatus according to any one of claims 10 to 15, characterized in that, The real-shot image is at least one video frame in the real-shot video. The acquisition module is specifically used to generate at least one first image corresponding to the at least one video frame based on the three-dimensional model of the at least one human body part, the three-dimensional model of the target object and the at least one video frame. The three-dimensional model of the at least one human body part is obtained by performing three-dimensional modeling on one of the at least one video frame.

17. The apparatus according to any one of claims 10 to 16, characterized in that, The plurality of human body parts are included in the torso and / or head of the human body.

18. The apparatus according to any one of claims 10 to 17, characterized in that, The camera position set in the rendering is obtained based on the camera position of the actual captured image, and / or the shooting angle set in the rendering is obtained based on the shooting angle of the actual captured image.

19. A computing device cluster, characterized in that, The computing device cluster includes at least one computing device, each computing device including a processor and memory: The memory is used to store instructions; The processor is configured to, according to the instructions, cause the computing device cluster to perform the method of any one of claims 1 to 9.

20. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions that, when executed by one or more computers, cause the one or more computers to perform the method of any one of claims 1 to 9.

21. A computer program product, characterized in that, The computer program product stores instructions that, when executed by a computer, cause the computer to perform the method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Virtual fitting device and virtual fitting method thereof

    CN105528799A

  • Virtual fitting method and device, electronic equipment and storage medium

    CN109829785A

  • Semantic generation method and device, aircraft and storage medium

    CN110832494A

  • Virtual garment wearing method and device, live broadcast system, electronic equipment and medium

    CN117409141A

  • Virtual clothing wearing method, virtual live broadcast method, virtual clothing wearing device, virtual live broadcast equipment and medium

    CN117611640A