Virtual fitting method and related device

By generating 3D models of human body parts and objects, and displaying stress and deformation, the problem of not being able to try on clothes on online platforms is solved, thus improving the user experience.

WO2025261116A1PCT designated stage Publication Date: 2025-12-26HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/097887
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-30
Filing Date
2025-05-29
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Unlike in physical stores, users cannot try on clothing items when purchasing them through online platforms, making it impossible to determine if the items are suitable.

Method used

By acquiring images of user body parts, and using 3D models and target object 3D models, status information is generated to display the stress or deformation of the body parts when wearing the target object from the target perspective, including stress cloud diagrams, deformation diagrams, and displacement information.

Benefits of technology

Users can intuitively understand the tightness of the item when wearing it, improving the fitting experience and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025097887_26122025_PF_FP_ABST
    Figure CN2025097887_26122025_PF_FP_ABST
Patent Text Reader

Abstract

A virtual fitting method and a related device, applicable in the field of virtual fitting. The method comprises: acquiring a first image frame, the first image frame comprising an image of a human body part at a target viewing angle; receiving a fitting article determination request, the fitting article determination request being used to indicate simulation of the human body part wearing a target article; providing a display page, the display page being used to display state information of the human body part wearing the target article at the target viewing angle, the state information being obtained by processing on the basis of a three-dimensional model of the human body part, a three-dimensional model of the target article, and the target viewing angle. Due to the fact that the three-dimensional model of the human body part can reflect the size of the human body part, and the three-dimensional model of the target article can reflect the size of the target article, the state information can accurately indicate a stress condition or a deformation condition of the human body part wearing the target article at the target viewing angle. By means of the state information, a loose or tight condition of a user wearing the target article can be known, as well as whether a size is appropriate.
Need to check novelty before this filing date? Find Prior Art

Description

A virtual wearable method and related devices

[0001] This application claims priority to Chinese Patent Application No. 202410814451.2, filed on June 20, 2024, entitled “A Wearable Method and Device”, the entire contents of which are incorporated herein by reference.

[0002] This application claims priority to Chinese Patent Application No. 202411389771.4, filed with the State Intellectual Property Office of China on September 30, 2024, entitled "A Virtual Wearing Method and Related Device", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to computer technology, and more particularly to a virtual wearable method and related equipment. Background Technology

[0004] With the development of technology, people's lives are becoming more and more convenient. With the rapid development of e-commerce, purchasing wearable items through online platforms has become a major choice for users. However, when purchasing wearable items through online platforms, users cannot try them on like in physical stores, and therefore cannot see the effect of wearing the item.

[0005] To solve the above problem, a virtual image of the user wearing the item can be obtained. However, this virtual image is merely a superposition of the item to be worn and the user's body, and the user cannot know whether the size of the item being tried on is suitable. Summary of the Invention

[0006] This application provides a virtual wearable method and related equipment. By using state information to indicate the force or deformation of the human body parts when wearing the target item from the target perspective, the tightness of the target item when the user wears it can be known through the state information, thereby determining whether the size of the target item is suitable.

[0007] This application provides the following technical solution:

[0008] Firstly, this application provides a virtual wearing method that can be used in the field of virtual try-on. In this method, a first device can acquire a first image frame, each of which includes an image of at least one human body part from a target perspective. The first device can also receive a wearing item confirmation request, which is used to instruct the simulated human body part to wear the target item. Furthermore, the first device can provide a display page, which is used to display the status information of the aforementioned at least one human body part wearing the target item from the target perspective. The status information is obtained based on the three-dimensional model of the aforementioned at least one human body part, the three-dimensional model of the target item, and the target perspective. The status information indicates the force or deformation of the human body part wearing the target item from the target perspective.

[0009] In this implementation, a first image frame is acquired, which includes an image frame of a human body part from the target viewpoint; a confirmation request for wearing an item is received, which is used to instruct the simulated human body part to wear the target item; then a display page is provided, which is used to display the status information of the human body part wearing the target item from the target viewpoint. The status information is obtained based on the three-dimensional model of the human body part, the three-dimensional model of the target item, and the target viewpoint. Since the three-dimensional model of the human body part can reflect the size of the human body part, and the three-dimensional model of the target item can reflect the size of the target item, the status information can indicate the force or deformation of the human body part wearing the target item from the target viewpoint. Therefore, the tightness of the target item when the user wears it can be known through the status information, thereby determining whether the size of the target item being tried on is suitable.

[0010] In one possible implementation, the state information includes at least one of the following: a stress cloud map indicating the force on at least one wearable part of a human body when wearing a target item from a target perspective; stress values ​​at at least one location of the wearable part when wearing the target item from a target perspective, wherein a larger stress value at a location indicates a greater force at that location, and a smaller stress value at a location indicates a smaller force at that location; a deformation map of the wearable part when wearing the target item from a target perspective, which reflects the deformation of at least one location of the wearable part when the first user's wearable part wears the target item; and displacement information of at least one location of the wearable part when wearing the target item from a target perspective, wherein the displacement information indicates the deformation of the wearable part when wearing the target item from a target perspective, and the aforementioned displacement information may include displacement distance.

[0011] In this implementation, at least one piece of information that the state information may include is clearly defined, which reduces the implementation difficulty of this solution. The stress cloud diagram can more intuitively show the stress situation of the wearing part when wearing the target item, and the stress value can more accurately reflect the stress situation at at least one position of the wearing part when wearing the target item. The deformation diagram can more intuitively reflect the deformation caused to the wearing part when wearing the target item, and the displacement information can more accurately reflect the deformation caused to the wearing part when wearing the target item. The state information can be selected according to the specific application scenario, which not only improves the implementation flexibility of this solution, but also helps to expand the application scenarios of this solution.

[0012] In one possible implementation, the display page is further used to display a second image frame, which includes an image of the wearable part of the human body in the first image frame superimposed with an image of the target item. The image of the target item is an image of the target item from the target's perspective. In this implementation, directly superimposing the image of the target item onto the wearable part of the first image frame helps to improve the realism of the second image frame, thereby enhancing the user's try-on experience.

[0013] In one possible implementation, the state information includes a stress cloud map, which is superimposed on the second image frame. The image of the target object and the stress cloud map corresponding to the first image frame can be superimposed on at least one human body part, including a wearable part, in the first image frame, thereby enabling the display of the second image frame and the state information.

[0014] In this implementation, since the stress cloud map can more intuitively and vividly display the stress on the body parts of the first user, the user can not only understand the effect of trying on the target item with the help of the second image frame and the stress cloud map, but also intuitively understand the stress on the body parts when wearing the target item. Thus, the user can have a more comprehensive understanding of the situation when wearing the target item, which is conducive to improving the user experience. Directly superimposing the image frame of the target item and the stress cloud map onto the body parts of the first image frame can more comprehensively restore the situation when the user wears the target item. The user can more clearly understand the stress on the body parts when wearing the target item, which is conducive to further improving the user experience.

[0015] In one possible implementation, the stress cloud map has a display state and a non-display state. The switching between these states is triggered by a user-inputted switching operation. The non-display state of the stress cloud map can be understood as not displaying the stress cloud map. For example, an icon for receiving a second operation can be deployed on the display device of the stress cloud map. This second operation could be a switching operation between the display and non-display states of the stress cloud map. Alternatively, receiving a click operation from the user within a first area of ​​the stress cloud map is considered a received switching operation that triggers the switching between the display and non-display states of the stress cloud map. Another example is that the aforementioned switching operation can be input via voice.

[0016] In this implementation, users can control the stress cloud map to switch between displayed and undisplayed states by inputting a switching operation. Users can choose whether to display the stress cloud map according to their actual needs, which improves the flexibility of users when trying on target items and helps to further improve the user experience of this solution.

[0017] In one possible implementation, the method further includes: a first device acquiring multiple third image frames, each of the multiple third image frames including images of at least one human body part from multiple perspectives; and then the first device performing three-dimensional reconstruction based on the multiple third image frames to obtain a three-dimensional model of the aforementioned at least one human body part.

[0018] For example, the 3D modeling algorithm can be the 3D human pose and shape regression with pyramidal mesh aligns the feedback loop (PyMAF) algorithm, the 4D Humans algorithm, or other 3D modeling algorithms, etc. The specific algorithm can be determined based on the actual application scenario.

[0019] For example, the aforementioned multiple perspectives may include multiple image frames obtained by shooting the body parts of the first user from multiple different angles. The aforementioned "perspective" may also be replaced by shooting angle, shooting perspective, or angle.

[0020] In this implementation, since the three-dimensional model of a human body part is constructed based on image frames from multiple perspectives of that human body part, the size of that human body part in the three-dimensional model is more similar to the actual size of that human body part. In other words, the three-dimensional model of that human body part can more accurately reflect the actual size of that human body part.

[0021] In one possible implementation, the method further includes: a first device adjusting the posture of a three-dimensional model of at least one human body part based on a first image frame to obtain a processed three-dimensional model of the at least one human body part, wherein the posture of at least one human body part in the processed three-dimensional model of the at least one human body part is consistent with the posture of at least one human body part in the first image frame, wherein the state information corresponding to the first image frame is obtained based on the processed three-dimensional model of at least one human body part, the three-dimensional model of the target object, and the target viewpoint.

[0022] In this implementation, the human body part is adjusted in a 3D model based on the first image frame to obtain a processed 3D model of the human body part. Since the 3D model of the human body part is constructed based on image frames from multiple perspectives of the human body part, the size of the human body part in the 3D model is more similar to the actual size of the human body part. Therefore, the processed 3D model of the human body part can not only more accurately reflect the size of the human body part, but also be consistent with the posture of the human body part in the first image frame. This is beneficial for accurately obtaining the state information of the wearing part of the human body part when wearing the target item from the target perspective, which is also beneficial for more accurately determining whether the size of the target item is appropriate.

[0023] In one possible implementation, the method further includes: receiving target item configuration information sent by the first device and the third device, the target item configuration information being used to indicate one or more of the target item's size information, physical information, and material information, the target item's physical information including the target item's weight, texture information, hardness / softness information, color information, or other physical parameter information, etc.; and then the first device can obtain a three-dimensional model of the target item based on the target item configuration information.

[0024] In this implementation, a three-dimensional model of the target item is obtained based on one or more of its size, physical, and material information. This helps to obtain an accurate three-dimensional model of the target item, which in turn helps to obtain more accurate state information. This allows for a more accurate understanding of whether the size of the target item is suitable when worn on the wearing part.

[0025] In one possible implementation, the method further includes: a first device, based on a 3D model of at least one human body part and a 3D model of the target object, performs a physical simulation of the wearing part of the human body part wearing the target object from a target perspective, obtaining state information corresponding to the first image frame. The physical simulation in this application can also be called a 3D simulation. For example, the physical simulation algorithm used by the physical simulation engine can be the finite element method (FEM), the finite integral in time domain (FDID) method, or other types of physical simulation algorithms.

[0026] In this implementation, since the 3D model of the human body part can carry the size information of the human body part and the 3D model of the target object can carry the size information of the target object, based on the 3D model of the human body part and the 3D model of the target object, physical simulation of the wearing part of the human body part wearing the target object from the target perspective can be performed to obtain accurate state information.

[0027] In one possible implementation, the aforementioned at least one human body part includes at least one of the following: neck, wrist, leg, or foot. For example, if the first image frame is an image frame of the upper body from the target's perspective, and the target item is a necklace, then at least one human body part may include: head, neck, and torso, with the neck being the part being worn. As another example, if the first image frame is an image frame of the forearm and hand from the target's perspective, and the target item is a bracelet, then at least one human body part may include: forearm, wrist, and hand, with the wrist being the part being worn. As yet another example, if the first image frame is an image frame of the whole body from the target's perspective, and the target item is shoes, then at least one human body part may include: head, torso, arm, hand, leg, and foot, with the foot being the part being worn, etc. It should be noted that the granularity of human body part division can differ in different application scenarios. For example, in some scenarios, the head is divided into three human body parts: skull, hair, and neck, while in other scenarios, the head is considered a single human body part. The granularity of human body part division in this application can be determined based on the actual application scenario.

[0028] This implementation provides various scenarios for different body parts, allowing users to try on the target item on various body parts, thus expanding the application scenarios of this solution. In particular, for smaller body parts such as the neck, wrist, or feet, the method provided in this application can be used to determine whether the target item fits the user when trying it on, which helps improve the user experience.

[0029] Secondly, this application provides a virtual wearable device that can be used in the field of virtual try-on. The device includes: an acquisition module for acquiring a first image frame, the first image frame including an image of a human body part from a target viewpoint; a receiving module for receiving a confirmation request for wearing an item, the confirmation request being used to instruct a simulated human body part to wear a target item; and a processing module for providing a display page, the display page being used to display status information of the human body part wearing the target item from the target viewpoint, the status information being processed based on a three-dimensional model of the human body part, a three-dimensional model of the target item, and the target viewpoint, the status information indicating the force or deformation of the human body part wearing the target item from the target viewpoint.

[0030] In the second aspect of this application, the virtual wearable device can also perform the steps performed by the first device in the first aspect and various possible implementations of the first aspect. The specific implementation of the steps in the second aspect and various possible implementations of the second aspect, the meaning of the terms, and the beneficial effects brought about by each possible implementation can all be referred to the description in the first aspect and various possible implementations of the first aspect, and will not be repeated here.

[0031] Thirdly, this application provides a computing device cluster, which includes at least one computing device, each computing device including a processor and a memory: the memory is used to store instructions; the processor is used to cause the computing device cluster to perform the method described in the first aspect or any possible implementation of the first aspect according to the instructions.

[0032] Fourthly, this application provides a computer storage medium storing one or more instructions, which, when executed by one or more computers, cause the one or more computers to perform the method described in the first aspect or any possible implementation of the first aspect.

[0033] Fifthly, this application provides a computer program product that stores instructions that, when executed by a computer, cause the computer to perform the method described in the first aspect or any of the possible implementations of the first aspect.

[0034] The beneficial effects brought about by the second to fifth aspects of this application can be referred to the descriptions in the first aspect and the various possible implementations of the first aspect, and will not be repeated here. Attached Figure Description

[0035] Figure 1 is a flowchart illustrating a virtual wearable method provided in an embodiment of this application.

[0036] Figure 2 is a schematic diagram of a second image frame provided in an embodiment of this application;

[0037] Figure 3 is a schematic diagram of a stress cloud diagram used to indicate the force on the wearing part provided in an embodiment of this application;

[0038] Figure 4 is a schematic diagram of a three-dimensional model of at least one human body part provided in an embodiment of this application;

[0039] Figure 5 is a schematic diagram of a three-dimensional model of clothing and a mesh that makes up the three-dimensional model of clothing provided in an embodiment of this application;

[0040] Figure 6 is another flowchart illustrating the virtual wearable method provided in an embodiment of this application;

[0041] Figure 7 is a schematic diagram of adjusting the image of the target item according to an embodiment of this application;

[0042] Figure 8 is a schematic diagram of an image of the target item provided in an embodiment of this application;

[0043] Figure 9 is another schematic diagram of the stress cloud diagram provided in the embodiment of this application;

[0044] Figure 10 is another schematic diagram of the wearable part provided in an embodiment of this application;

[0045] Figure 11 is another schematic diagram of the image of the target item provided in the embodiment of this application;

[0046] Figure 12 is a schematic diagram of a virtual wearable device provided in an embodiment of this application;

[0047] Figure 13 is a schematic diagram of a wearable image acquisition system provided in an embodiment of this application;

[0048] Figure 14 is a schematic diagram of a computing device provided in an embodiment of this application;

[0049] Figure 15 is a schematic diagram of a computing device cluster provided in an embodiment of this application;

[0050] Figure 16 is a schematic diagram of computer devices in a computer cluster connected via a network according to an embodiment of this application. Detailed Implementation

[0051] The embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are merely some, and not all, of the embodiments of this application. Those skilled in the art will recognize that, with the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0052] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the description of embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0053] In the embodiments of this application, "instruction" can include direct and indirect instructions, as well as explicit and implicit instructions. The information indicated by a certain piece of information (hereinafter referred to as instruction information) is called the information to be instructed. In specific implementation, there are many ways to indicate the information to be instructed, such as, but not limited to, directly indicating the information to be instructed, such as the information to be instructed itself or its index. It can also indirectly indicate the information to be instructed by indicating other information, where there is an association between the other information and the information to be instructed; or it can indicate only a part of the information to be instructed, while the other parts are known or pre-agreed upon. For example, the instruction can be implemented by using a pre-agreed (e.g., protocol predefined) arrangement of various information, thereby reducing the instruction overhead to a certain extent. This application does not limit the specific method of instruction. It is understood that for the sender of the instruction information, the instruction information can be used to indicate the information to be instructed; for the receiver of the instruction information, the instruction information can be used to determine the information to be instructed.

[0054] The method provided in this application can be applied to various virtual try-on scenarios. For example, when a user tries on an item using an augmented reality (AR) device, virtual try-on technology can be used to obtain the user's appearance when trying on the item. Another example is when a user purchases an item through an online platform; the same applies when a user tries on virtual items in a game. Furthermore, the method provided in this application has other applications, which will not be exhaustively listed here.

[0055] Current virtual try-on technology can obtain virtual wear image frames when a user wears an item. However, these virtual wear image frames are merely superimposed images of the item to be worn and the user's body, and the user cannot know whether the size of the item is suitable. To solve the aforementioned problem, this application discloses: a first device acquires a first image frame, which includes an image frame of a human body part from a target viewpoint; receives a wear item confirmation request, which instructs the simulated human body part to wear the target item; and then provides a display page, which displays the status information of the human body part wearing the target item from the target viewpoint. The status information is obtained based on the three-dimensional model of the human body part, the three-dimensional model of the target item, and the target viewpoint. Since the three-dimensional model of the human body part can reflect the size of the human body part, and the three-dimensional model of the target item can reflect the size of the target item, the status information can indicate the stress or deformation of the human body part wearing the target item from the target viewpoint. Therefore, the user can know the tightness of the target item when wearing it, and thus determine whether the size of the target item is suitable.

[0056] The specific implementation flow of the method provided in the embodiments of this application will be described below. Specifically, please refer to Figure 1, which is a schematic flowchart of a virtual wearable method provided in the embodiments of this application. The virtual wearable method provided in the embodiments of this application may include:

[0057] 101. Acquire the first image frame, which includes an image of the human body part from the target viewpoint.

[0058] The first image frame includes a video frame of at least one human body part from the target viewpoint. The aforementioned at least one human body part includes at least the wearing part corresponding to the target item that the first user wants to try on. The first user refers to the user in the first image frame, that is, the human body part shown in the first image frame is the human body part of the first user.

[0059] For example, the aforementioned at least one human body part may include at least one of the following: neck, wrist, ankle, foot, head, torso, or other parts. For instance, if the target item the first user wants to try on includes a necklace, then at least one human body part in the first image frame includes the first user's neck; as another example, if the target item the first user wants to try on includes a watch, bracelet, or bangle, then at least one human body part in the first image frame includes the first user's wrist; as another example, if the target item the first user wants to try on includes an anklet, then at least one human body part in the first image frame includes the first user's ankle; as another example, if the target item the first user wants to try on includes shoes, then at least one human body part in the first image frame includes the first user's foot; as another example, if the target item the first user wants to try on includes a wig, headphones, or glasses, then at least one human body part in the first image frame includes the first user's head; as another example, if the target item the first user wants to try on includes clothing, then at least one human body part in the first image frame includes the first user's torso, etc. The specific details of the target item and the first image frame can be determined in conjunction with the actual application scenario. The examples here are only for the convenience of understanding this solution.

[0060] For example, if the first image frame is an image frame of the upper body from the target's perspective, and the target item is a necklace, then at least one human body part can include: head, neck, and torso, with the necklace worn on the neck. As another example, if the first image frame is an image frame of the head from the target's perspective, and the target item is a wig, then at least one human body part can include: skull, hair, and neck, with the necklace worn on the neck. As yet another example, if the first image frame is an image frame of the whole body from the target's perspective, and the target item is shoes, then at least one human body part can include: head, torso, arms, hands, legs, and feet, with the necklace worn on the feet. As yet another example, if the first image frame is an image frame of the forearm and hand from the target's perspective, and the target item is a bracelet, then at least one human body part can include: forearm, wrist, and hand, with the necklace worn on the wrist, and so on.

[0061] It should be noted that the granularity of human body parts division can vary in different application scenarios. For example, in some scenarios, the head is divided into three human body parts: skull, hair, and neck, while in other scenarios, the head is considered as a single human body part. The granularity of human body parts division in this application can be determined in conjunction with the actual application scenario.

[0062] This application provides various scenarios for different body parts, allowing users to try on target items on different body parts using the method provided in this application, which helps expand the application scenarios of this solution. In particular, for smaller body parts such as the neck, wrist, or feet, the method provided in this application can also be used to determine whether the target item fits the user when trying it on, which helps improve the user experience.

[0063] In one scenario, the at least one first image frame acquired by the first device in step 101 can be at least one video frame from a video, meaning each first image frame is a video frame. It should be noted that if the first device acquires a video, the video can be segmented into at least one video frame. In another scenario, each first image frame acquired by the first device in step 101 is an independently captured two-dimensional image; in other words, each acquired first image frame may not be a video frame from a video.

[0064] For example, in one scenario, such as when the first device is a cloud server, step 101 may include: the first device receiving at least one first image frame sent by the client device. In another scenario, such as when the first device is a client device that directly interacts with the user, step 101 may include: the first device obtaining the at least one first image frame from a local photo album, or step 101 may include: the first device capturing the at least one first image frame using a camera.

[0065] In step 101, optionally, after acquiring at least one first image frame, the first device may also perform image enhancement on each first image frame to improve the clarity of each first image frame.

[0066] For example, a large model for image processing can be used to enhance each first image frame. Exemplarily, a prompt message and each first image frame are input into the large model for image processing to obtain an updated first image frame output by the large model. The aforementioned prompt message can be used to prompt the large model to improve the clarity of the input image frame. The large model for image processing is a machine learning model. Optionally, the aforementioned large model for image processing can be a machine learning model based on an attention mechanism.

[0067] Alternatively, a machine learning model specifically designed to improve image clarity can be used to process each first image frame to obtain an updated first image frame. In other words, the function of the aforementioned machine learning model is to improve image clarity. For example, the aforementioned machine learning models can be convolutional neural networks, fully connected neural networks, residual neural networks, attention-based neural networks, support vector machines, or other types of machine learning models.

[0068] 102. Receive a confirmation request for wearing items. The confirmation request for wearing items is used to instruct the simulated human body parts to wear the target items.

[0069] For example, the wearable item confirmation request may include the identification information of the target item. The wearable item confirmation request is used to enable the first device to determine the target item from multiple items. Then, the first device can confirm the target item from multiple items based on the received wearable item confirmation request, and then simulate wearing the target item on the aforementioned human body parts.

[0070] In one scenario, for example, where the first device is a cloud server, step 102 may include: the first device receiving a wearable item confirmation request sent by a client device. In another scenario, for example, where the first device is a client device that directly interacts with the user, step 102 may include: when the first device receives a first operation input by the second user, it obtains the wearable item confirmation request; for example, the first operation may be a click operation on the target item, or the first operation may be a dragging operation on the target item to a preset area, or the first operation may also be a long press operation on the target item, etc. The specific implementation method can be determined according to the actual application scenario.

[0071] Optionally, after receiving the wearable item confirmation request, the first device may also obtain at least one status information corresponding to at least one first image frame (or the updated first image frame) and at least one second image frame corresponding to at least one first image frame (or the updated first image frame).

[0072] Further, in one scenario, for example, where the first device is a cloud server, or where the first device is a client device and the client device generates the second image frame and status information locally, step 102 may include: the first device generating a second image frame and status information corresponding to each first image frame (or updated first image frame) based on the acquired at least one first image frame (or updated first image frame). In another scenario, for example, where the first device is a client device and the cloud server generates the second image frame and status information, step 102 may include: the first device receiving at least one second image frame corresponding to at least one first image frame (or updated first image frame) and at least one status information corresponding to the aforementioned at least one first image frame (or updated first image frame) sent by the cloud server.

[0073] Each second image frame can be understood as a virtual image of the target item being worn on the wearable part of the human body. The second image frame includes the image of the wearable part of the human body in the first image frame superimposed with the image of the target item. The image of the target item is the image of the target item from the target perspective. For example, in one case, each second image frame is an image frame of the target item being tried on and worn on the wearable part of the human body in the first image frame, which helps to improve the realism of the second image frame and improve the user's try-on experience.

[0074] To more intuitively understand this solution, please refer to Figure 2, which is a schematic diagram of a second image frame provided in an embodiment of this application. In Figure 2, taking the wrist as the wearable part of the human body and a watch as the target item to be tried on, the first image frame can be an image frame of the user's hand plus the wrist, as shown in Figure 2. The second image frame is an image frame of the watch superimposed on the user's wrist in the first image frame. It should be understood that the example in Figure 2 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0075] In another case, each second image frame may also include: a model corresponding to a human body part in the first image frame, and the model of the human body part has a target item worn on the wearable part; wherein, the pose of the model of the human body part is the same as the pose of the human body part in the first image frame; for example, the model can be a three-dimensional (3D) model or a two-dimensional (2D) model, etc., which can be determined according to the actual application scenario.

[0076] The status information corresponding to each first image frame indicates the stress or deformation of the wearable parts of the human body in that first image frame when wearing the target item from a target viewpoint. For example, the status information corresponding to each second image frame includes at least one of the following: a stress cloud map indicating the stress on the wearable parts of the human body when wearing the target item from a target viewpoint; stress values ​​at at least one location of the wearable parts when wearing the target item from a target viewpoint; deformation maps of the wearable parts when wearing the target item from a target viewpoint; displacement information at at least one location of the wearable parts when wearing the target item from a target viewpoint, the displacement information indicating the deformation of the wearable parts when wearing the target item from a target viewpoint; strain values ​​at at least one location of the wearable parts when wearing the target item from a target viewpoint; or other information that can reflect the stress or deformation of the wearable parts, etc. These can be specifically set according to the actual application scenario, and are not exhaustively listed here.

[0077] To further understand this solution, please refer to Figure 3. Figure 3 is a schematic diagram of a stress cloud map provided by an embodiment of this application for indicating the stress condition of the wearing part. In Figure 3, the wearing part is the foot as an example. The stress cloud map can be represented as a color cloud map, that is, the stress cloud map uses different colors to show the stress magnitude of the wearing part. For example, red represents high stress and tight fit; blue represents low stress and loose fit; yellow and green represent stress between red and blue, and the degree of fit is also between blue and red. Since the color stress cloud map in Figure 3 has been grayscaled, it can be seen that the gray depth corresponding to different areas of the foot is different, which can also represent different stress and different degrees of fit at different positions in the wearing part. It should be understood that the example in Figure 3 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0078] Furthermore, when an object deforms due to external factors, internal forces are generated between the various parts of the object to resist the action of these external factors and attempt to restore the object from its deformed position to its original position. The internal force per unit area on the cross-section under consideration is called stress. The stress value at at least one location in the first user's wearing part can reflect the force situation at at least one location in that wearing part when the first user wears the target item. The larger the stress value at a certain location, the greater the force at that location; the smaller the stress value at a certain location, the smaller the force at that location.

[0079] The deformation diagram of the first user's wearing part can show the deformed wearing part when the first user wears the target item from the target perspective, and can reflect the deformation of at least one position of the wearing part when the first user wears the target item.

[0080] The displacement information may include displacement distance, and optionally, the displacement information may also include displacement velocity; the displacement distance of at least one position of the first user's wearing part can reflect what kind of deformation occurs at the aforementioned at least one position of the wearing part when the first user wears the target item from the target viewpoint; the displacement velocity of at least one position of the first user's wearing part can reflect the speed at which the deformation occurs at the aforementioned at least one position of the wearing part when the first user wears the target item from the target viewpoint.

[0081] Strain refers to the physical quantity that causes a relative change in the geometry and size of an object due to external factors. Strain can also be called relative deformation. The strain value at at least one position in the first user's wearing part can reflect the degree of deformation at at least one position in the wearing part when the first user wears the target item from the target perspective. The larger the strain value at a certain position, the greater the degree of deformation at that position; the smaller the strain value at a certain position, the smaller the degree of deformation at that position.

[0082] In this embodiment, the status information may include at least one piece of information, reducing the implementation difficulty of this solution; the stress cloud diagram can more intuitively show the stress on the wearing part when wearing the target item, the stress value can more accurately reflect the stress at at least one position of the wearing part when wearing the target item, the deformation diagram can more intuitively reflect the deformation caused to the wearing part when wearing the target item, and the displacement information can more accurately reflect the deformation caused to the wearing part when wearing the target item. The choice of which or which status information to use can be selected according to the specific application scenario, which not only improves the implementation flexibility of this solution, but also helps to expand the application scenarios of this solution.

[0083] The state information is obtained by processing a 3D model of a human body part, a 3D model of the target object, and a target viewpoint. Optionally, the first device can perform a physical simulation of the wearing part of the human body part wearing the target object from the target viewpoint based on a second 3D model of at least one human body part, a 3D model of the target object, and a target viewpoint, to obtain state information corresponding to the first image.

[0084] For example, after acquiring at least one first image frame (or an updated first image frame), the first device can generate state information corresponding to each of the at least one first image frames based on the three-dimensional model of the human body part, the three-dimensional model of the target object, and the target viewpoint.

[0085] In one scenario, the first device can independently acquire the status information corresponding to each first image frame, whether it is an independently captured image frame or at least one video frame in a video.

[0086] For ease of description, in this embodiment, any one of the at least one first image frame (or the updated first image frame) is referred to as the "target image frame". For example, the first device can acquire a first three-dimensional model corresponding to the target image frame, wherein the pose of the human body parts in the first three-dimensional model is consistent with the pose of the human body parts in the target image frame. Optionally, the size of the human body parts in the first three-dimensional model can be obtained based on the image frame of the first user's human body parts, that is, the size of the human body parts in the first three-dimensional model can be as similar as possible to the size of the first user's human body parts. The first device can also acquire a three-dimensional model of the target item; and then, based on the first three-dimensional model corresponding to the target image frame, the three-dimensional model of the target item, and the target viewpoint, it can determine the state information corresponding to the target image frame, that is, determine the state information of the wearing part when trying on the target item on the wearing part of the first three-dimensional model corresponding to the target image frame. The first device repeats the aforementioned operation at least once to obtain at least one state information corresponding one-to-one with at least one first image frame (or the updated first image frame).

[0087] The first three-dimensional model can be understood as a three-dimensional model including at least one human body part of the first user. For example, the first three-dimensional model can be represented by first information, which may include first position information of the first vertices among the plurality of first vertices used to form the first three-dimensional model. For example, the first position information of the first vertex may be the coordinates of the first vertex in three-dimensional space. The first information may also include first connection relationship information between different first vertices among the plurality of first vertices. This first connection relationship information indicates which first vertices among the plurality of first vertices are connected. The first vertices with connection relationships can be connected to obtain multiple meshes that make up the user's three-dimensional model. Optionally, the first information may also include skin color or other information, etc., which are not exhaustively listed here.

[0088] To understand this solution more intuitively, please refer to Figure 4. Figure 4 is a schematic diagram of a three-dimensional model of at least one human body part provided in an embodiment of this application. In Figure 4, the three-dimensional model of at least one human body part is taken as an example of the model of the entire human body. It should be noted that Figure 4 is only a schematic diagram obtained after visualizing the three-dimensional model of at least one human body part in this solution. The first device can be deployed with first information to represent the three-dimensional model of at least one human body part. In addition, the example in Figure 4 is only for the convenience of understanding the user's concept of a three-dimensional model and is not intended to limit this solution.

[0089] Regarding the specific implementation of the first device acquiring the first 3D model, in one implementation, the first 3D model corresponding to the target image frame (i.e., any first image frame or any updated first image frame) is obtained based on the target image frame and the second 3D model. The second 3D model is obtained based on multiple image frames of at least one human body part of the first user. The second 3D model can also be understood as a 3D model of at least one human body part of the user. For example, the first device can obtain the first 3D model corresponding to the target image frame based on the second 3D model and the target image frame. Further, the first device adjusts the posture of the second 3D model of at least one human body part according to the first image frame to obtain a processed 3D model of at least one human body part (i.e., the first 3D model of at least one human body part). The posture of the human body part in the first 3D model (i.e., the processed 3D model of at least one human body part) is the same as the posture of the human body part in the target image frame. In other words, the first 3D model can be a processed 3D model of the human body part obtained by adjusting the posture of the human body part in the second 3D model according to the posture of the human body part of the first user in the target image frame.

[0090] For example, the second three-dimensional model can be obtained by three-dimensional modeling based on multiple third image frames of the human body parts of the first user. For example, before performing step 203, the first device can also acquire multiple third image frames, which include images of at least one human body part from multiple perspectives; and perform three-dimensional reconstruction based on the multiple third image frames to obtain a second three-dimensional model of at least one human body part.

[0091] For example, the 3D modeling algorithm can be the 3D human pose and shape regression with pyramidal mesh aligns the feedback loop (PyMAF) algorithm, the 4D Humans algorithm, or other 3D modeling algorithms, etc. The specific algorithm can be determined based on the actual application scenario.

[0092] Optionally, the aforementioned multiple perspectives may include multiple image frames obtained by capturing images of the first user's body parts from multiple different angles. The term "perspective" can also be replaced with shooting angle, shooting angle, angle, etc., without specific limitation here. For example, multiple image frames of at least one body part of the first user may be obtained by capturing images of that at least one body part of the first user from at least two different shooting angles. For example, if at least one body part of the first user includes a foot, taking the direction of the toes as forward as an example, the multiple image frames corresponding to at least one body part of the first user may include: an image frame captured from the front of at least one body part of the first user, an image frame captured from the left side of at least one body part of the first user, an image frame captured from the right side of at least one body part of the first user, and an image frame captured from the rear of at least one body part of the first user. For example, if at least one human body part includes the wrist, the image frames corresponding to the at least one human body part of the first user from multiple angles may include: an image frame taken from above the at least one human body part of the first user, an image frame taken from below the at least one human body part of the first user, an image frame taken from the left side of the at least one human body part of the first user, an image frame taken from the right side of the at least one human body part of the first user, etc. It should be noted that the examples here are only for the convenience of understanding this solution, and the specifics can be determined in combination with the actual application scenario.

[0093] In this embodiment of the application, since the three-dimensional model of the human body part is constructed based on image frames from multiple perspectives of the human body part, the size of the human body part in the second three-dimensional model is more similar to the actual size of the human body part of the first user. That is, the three-dimensional model of the human body part can more accurately reflect the actual size of the human body part.

[0094] Then, based on the first image frame, the pose of the human body part in the three-dimensional model of the human body part is adjusted to obtain the processed three-dimensional model of the human body part. The processed three-dimensional model of the human body part can more accurately reflect the size of the human body part, which is conducive to obtaining more accurate state information, that is, it is conducive to more accurately knowing whether the size of the target object is appropriate.

[0095] In another implementation, the first device can also directly perform 3D modeling based on the target image frame (i.e., any first image frame or any updated first image frame) to obtain the first 3D model corresponding to the target image frame. The 3D modeling algorithm used can be referred to the above description and will not be repeated here.

[0096] For example, the first device can utilize second information to represent the three-dimensional model of the target item. The second information may include second position information for each of the plurality of second vertices constituting the three-dimensional model of the target item. For instance, the second position information for each second vertex may be the coordinates of each second vertex in three-dimensional space, such as coordinates on the x-axis, y-axis, and z-axis. The second information also includes second connection relationship information between different second vertices among the aforementioned plurality of second vertices. This second connection relationship information indicates which second vertices among the aforementioned plurality of second vertices are connected, and the connected second vertices can be connected to obtain multiple meshes constituting the three-dimensional model of the target item. Optionally, the second information may also include at least one of the following information of the three-dimensional model of the target item: material information, hardness / softness, texture information, color information, weight, or other physical parameter information, etc.

[0097] Optionally, before generating the status information, the first device may also receive target item configuration information sent by the second device. The target item configuration information indicates one or more of the target item's size, physical information, and material information. The physical information of the target item may include its weight, texture, hardness / softness, color, or other physical parameters. Then, the first device can obtain a three-dimensional model of the target item based on the target item configuration information. For example, the first device can perform three-dimensional modeling of the target item based on the target item configuration information to obtain a three-dimensional model of the target item.

[0098] To understand this solution more intuitively, please refer to Figure 5. Figure 5 is a schematic diagram of a three-dimensional model of clothing and the mesh that makes up the three-dimensional model of clothing provided in the embodiment of this application. Figure 5 includes a left sub-schematic diagram and a right sub-schematic diagram. The left sub-schematic diagram of Figure 5 shows a 5D garment, and the right sub-schematic diagram of Figure 5 shows the vertices and mesh of the collar part that makes up the three-dimensional garment. It should be noted that Figure 5 is only a schematic diagram obtained after visualizing the three-dimensional garment in this solution for the convenience of understanding. The first device can be deployed as first information for representing the three-dimensional target object. In addition, the example in Figure 5 is only for the convenience of understanding the concept of the three-dimensional target object and is not intended to limit this solution.

[0099] In this embodiment of the application, a three-dimensional model of the target item is obtained based on one or more of the target item's size information, physical information, and material information. This is beneficial for obtaining an accurate three-dimensional model of the target item, which in turn is beneficial for obtaining more accurate state information. This allows for a more accurate understanding of whether the size of the target item is suitable when the wearable part is wearing it.

[0100] For example, the first device can adjust the posture of a second 3D model of at least one human body part based on a target video frame to obtain a first 3D model corresponding to the target video frame. Then, based on the first 3D model corresponding to the target image frame, the 3D model of the target object, and the target viewpoint, a physics simulation engine is used to perform a physical simulation of the 3D model of the wearable part in the first 3D model wearing the target object, that is, to simulate wearing the target object on the wearable part, so as to determine the state information of the wearable part in the first 3D model corresponding to the target image frame when wearing the target object from the target viewpoint. The first device repeats the above operation at least once to obtain at least one state information corresponding to at least one first image frame.

[0101] The physical simulation in this application can also be referred to as three-dimensional simulation. For example, the physical simulation engine may use the finite element method (FEM), finite integral in time domain (FDID) method, or other types of physical simulation algorithms.

[0102] Optionally, the first device can also acquire the human body parameter information of the first user. For example, the human body parameter information of the first user may include weight, height, or other parameter information. Then, the first device can use a physics simulation engine to perform a physical simulation of the wearing part in the first 3D model trying on the 3D model of the target item based on the first 3D model corresponding to the target image frame, the 3D model of the target item, the target viewpoint, and the human body parameter information of the first user, to determine the state information of the wearing part when trying on the 3D model of the target item on the wearing part of the first 3D model corresponding to the target image frame. For example, if the wearing part is the foot and the target item to be tried on is shoes, the weight of the first user can be used when performing a physical simulation of the first user trying on shoes in a standing or sitting posture. It should be understood that this example is for the convenience of understanding this solution, and the specific implementation can be determined in combination with the actual application scenario.

[0103] In another scenario, where the acquired at least one first image frame (or updated first image frame) is at least one video frame in a video, the first device can further utilize a physics simulation engine to perform a physical simulation on the 3D model of the target item, including the wearing parts of the human body included in each of the aforementioned at least one first image frame (or updated first image frame), based on the first 3D model corresponding to one of the aforementioned video frames (or updated first image frames), the 3D model of the target item, and the at least one first image frame (or updated first image frame), to obtain the state information corresponding to each first image frame (or updated first image frame); one of the aforementioned at least one first image frames can be the first video frame. For example, the first device can input the first 3D model corresponding to one of the aforementioned at least one first image frame (or updated first image frame), the at least one first image frame (or updated first image frame), and the 3D model of the target item into the physics simulation engine to obtain the state information corresponding to each first image frame (or updated first image frame) generated by the physics simulation engine.

[0104] In this embodiment of the application, since the three-dimensional model of the human body part can carry the size information of the human body part and the three-dimensional model of the target item can carry the size information of the target item, based on the three-dimensional model of the human body part and the three-dimensional model of the target item, physical simulation of the wearing part of the human body part wearing the target item from the target perspective can be performed to obtain accurate state information.

[0105] For example, after acquiring at least one first image frame (or an updated first image frame), the first device may also generate a second image frame corresponding to each first image frame (or an updated first image frame).

[0106] Furthermore, in one case, the second image frame is an image of the target object superimposed on at least one human body part, including the wearing part, in the first image frame, and the image of the target object is an image of the target object from the target perspective; optionally, the image of the target object can be obtained by rendering the three-dimensional model of the target object from three dimensions to two dimensions, and the image of the target object is obtained by combining the first three-dimensional model of at least one human body part with the three-dimensional model of the target object, and the pose of the wearing part in the first three-dimensional model is consistent with the pose of the wearing part in the first image frame.

[0107] The following describes the specific implementation of generating a second image frame corresponding to a target image frame (i.e., any one of at least one first image frame or at least one updated image frame). For example, when the first device performs a physical simulation of the 3D model of the wearable part wearing the target item in the target image frame using a physical simulation engine, it can also render the scene of the 3D model of the wearable part wearing the target item in the first 3D model corresponding to the target image frame to 2D to obtain a virtual wearable image corresponding to the target image frame. The first device can also obtain a local area containing only the target item from the virtual wearable image corresponding to the target image frame, i.e., obtain an image of the target item corresponding to the target image frame. Then, the first device can superimpose the image of the target item corresponding to the target image frame onto the wearable part in the target image frame to obtain the second image frame corresponding to the target image frame. The first device repeats the aforementioned steps at least once to obtain at least one second image frame corresponding one-to-one with at least one first image frame (or an updated first image frame).

[0108] For example, semantic segmentation can be performed on the virtual wearable image corresponding to the target image frame to obtain the pixels in the virtual wearable image that have the semantic meaning of the target item, thereby obtaining the image of the target item in the virtual wearable image corresponding to the target image frame.

[0109] Optionally, the relative relationship between the image of the target item and the wearing part in the second image frame is consistent with the relative relationship between the target item and the wearing part in the virtual wearing image. For example, the relative relationship can be understood as a relative size relationship, a relative position relationship, a size relationship, etc., so that the second image frame can show whether the target item fits when the first user tries it on.

[0110] For example, in one implementation, the first device can determine the relative relationship between the image of the target item and the image of the wearing part in the second image frame based on the relative relationship between the target item and the wearing part in the virtual wearable image. The first device can identify the wearing part from the first image frame (or the updated first image frame) and then overlay the image frame of the target item onto the wearing part in the first image frame (or the updated first image frame) according to the aforementioned relative relationship. In another implementation, the first device can use a third machine learning model to perform the overlay operation. The input of the third machine learning model includes the first image frame (or the updated first image frame), the image of the target item, and the virtual wearable image. The output of the third machine learning model is the second image frame, etc. This example is only to demonstrate the feasibility of this solution; the specific implementation method can be determined in combination with the actual application scenario.

[0111] In another scenario, when the first device performs a physical simulation of the 3D model of the wearable part wearing the target item in the target image frame using a physical simulation engine, it can also render the scene of the wearable part trying on the 3D model of the target item in the first 3D model corresponding to the target image frame to a 2D model to obtain a virtual wearable image corresponding to the target image frame. The aforementioned virtual wearable image corresponding to the target image frame is considered as a second image frame corresponding to the target image frame. The second image frame includes an image of the wearable part in the human body part of the first image frame superimposed with an image of the target item. The image of the target item is an image of the target item from the target perspective, and the second image frame is an image from the target perspective. The first device repeats the aforementioned steps at least once to obtain at least one second image frame that corresponds one-to-one with at least one first image frame (or an updated first image frame).

[0112] Optionally, before performing physical simulation using the physical simulation engine, the first device can also obtain the first size information of the wearable part in the first three-dimensional model corresponding to the target image frame and the second size information of the three-dimensional model of the target item. If, based on the first and second size information, it is determined that the size difference between the wearable part and the target item exceeds a preset range, it can be determined that the target item does not fit properly, and the physical simulation engine will no longer be used to perform physical simulation on the wearable part trying on the target item. If, based on the first and second size information, it is determined that the size difference between the first user's wearable part and the target item is within a preset range, the operation of performing physical simulation using the physical simulation engine described above can be triggered. For example, the preset range may include a first threshold and a second threshold. When the first threshold is exceeded, it proves that the target item is too small and the first user's wearable part cannot wear the target item; when the second threshold is exceeded, it proves that the target item is too large for the first user's wearable part.

[0113] Optionally, if the size difference between the wearing part and the target item exceeds a preset range based on the first size information and the second size information, the user can be directly prompted that the target item does not fit in subsequent steps; for example, the user can be prompted that the target item is too small, or that the target item is too big. The specific prompting method can be determined in combination with the actual application scenario.

[0114] Optionally, if the at least one first image frame obtained in step 101 is at least one video frame in the video, after the first device obtains the second image frame corresponding to each video frame, it can compose the user's virtual wearable video by combining the second image frames corresponding to all the video frames in the at least one video frame; it can also perform inter-frame consistency processing on the user's virtual wearable video to obtain an updated user's virtual wearable video.

[0115] For example, a large image processing model can be used to perform inter-frame consistency processing on the aforementioned user's virtual wearable video. For instance, the prompt information and the user's virtual wearable video are input into the large image processing model to obtain an updated user virtual wearable video output by the model. This prompt information can be used to improve the inter-frame consistency between different video frames. The role of inter-frame consistency processing can be understood as increasing the similarity between adjacent video frames, making the transitions between different video frames in the updated user virtual wearable video smoother, and allowing the combined virtual wearable video to display the dynamic effects of the user wearing the target item.

[0116] Alternatively, a machine learning model specifically designed to improve the inter-frame consistency of video can be used to process the aforementioned user's virtual wearable video to obtain an updated version of the user's virtual wearable video. In other words, the function of the aforementioned machine learning model only includes improving the inter-frame consistency between different video frames.

[0117] 103. Provide a display page. The display page is used to display the status information of the human body parts wearing the target item from the target perspective. The status information is obtained based on the three-dimensional model of the human body parts, the three-dimensional model of the target item, and the target perspective. The status information indicates the force or deformation of the human body parts wearing the target item from the target perspective.

[0118] After acquiring the status information corresponding to each first image frame, the first device can also provide a display page to display the status information of the human body parts wearing the target item from the target viewpoint. For example, in one scenario, if the first device is a cloud server, step 103 may include: the first device sending a display interface to a client device, the display interface including the status information. In another scenario, if the first device is a client device, step 103 may include: the first device displaying the acquired status information on the display interface.

[0119] In this embodiment, a first image frame is acquired, which includes an image frame of a human body part from a target viewpoint; a confirmation request for wearing an item is received, which is used to instruct a simulated human body part to wear a target item; and a display page is provided, which is used to display the status information of the human body part wearing the target item from the target viewpoint. The status information is obtained based on the three-dimensional model of the human body part, the three-dimensional model of the target item, and the target viewpoint. The status information indicates the force or deformation of the human body part when wearing the target item from the target viewpoint. Through this status information, the tightness of the target item worn by the user can be known, thereby determining whether the size of the target item being tried on is suitable.

[0120] Based on the embodiment described in Figure 1 above, further, in one case, the acquisition operation of the first image frame, the generation operation of the second image frame and status information, and the display operation of the second image frame and status information are all performed by the client device; in another case, after the client device acquires the first image frame (or the updated first image frame), it can send the first image frame (or the updated first image frame) to the cloud server, whereby the server performs the generation operation of the second image frame and status information, and the cloud server sends the generated second image frame and status information to the client device, whereby the client device continues to perform the display operation of the second image frame and status information. Here, the virtual wearable method provided in this application is described in detail using the example that the acquisition operation of the first image frame, the generation operation of the second image frame and status information, and the display operation of the second image frame and status information are all performed by the client device. Specifically, please refer to Figure 6, which is another flowchart of the virtual wearable method provided in this application embodiment. The virtual wearable method provided in this application embodiment may include:

[0121] 601. Acquire the first image frame, which includes an image of the human body part from the target viewpoint.

[0122] The specific implementation of step 601 can be found in the description of the embodiment corresponding to Figure 2 above, and will not be repeated here.

[0123] 602. Receive a confirmation request for wearing items. The confirmation request for wearing items is used to instruct the simulated human body parts to wear the target items.

[0124] 603. Obtain a second image frame and status information corresponding to the first image frame. The second image frame includes an image of the wearing part in the human body part of the first image frame superimposed with an image of the target object. The image of the target object is an image of the target object from the target perspective. The status information is obtained based on the three-dimensional model of the human body part, the three-dimensional model of the target object, and the target perspective. The status information indicates the force or deformation of the human body part when wearing the target object from the target perspective.

[0125] The specific implementation methods of steps 601 and 602, as well as the specific generation methods of the second image frame corresponding to the first image frame and the status information, can be referred to the description in the embodiment corresponding to Figure 1 above, and will not be repeated here.

[0126] It should be understood that in the embodiment corresponding to Figure 4, the first image frame acquisition operation, the second image frame and status information generation operation, and the second image frame and status information display operation are all performed by the first device as an example. If the second image frame and status information generation operation is performed by the cloud server (that is, step 603 is performed by the cloud server), the cloud server can send the generated second image frame and status information to the client device. In step 603, the client device can receive the second image frame and status information sent by the cloud server.

[0127] 604. Provide a display page, which is used to display the status information of the human body parts wearing the target item from the target's perspective and the second image frame.

[0128] For example, after the client device obtains the second image frame and status information corresponding to each first image frame (or the updated first image frame), it can provide a display page. For example, the client device can display the obtained second image frame and status information on the display interface.

[0129] Regarding the specific implementation of step 604 by the client equipment, in one case, regardless of whether the at least one first image frame obtained is an independently captured image frame or at least one video frame in a video, the client equipment can display the second image frame corresponding to each first image frame (or the updated first image frame) and the status information corresponding to each first image frame (or the updated first image frame).

[0130] Optionally, in one implementation, the client device can display a second image frame corresponding to each first image frame (or an updated first image frame) and a stress cloud map, wherein the stress cloud map corresponding to each first image frame indicates the stress on the body parts of the first user in the first image frame (or the updated first image frame).

[0131] Furthermore, in one implementation, the stress cloud map is superimposed on the second image frame. For example, the client device can superimpose an image of the target item and the corresponding stress cloud map onto the wearable part in each first image frame (or an updated first image frame), thereby enabling the display of the second image frame and status information.

[0132] In this embodiment, since the stress cloud map can more intuitively and vividly display the stress on the body parts of the first user, the user can not only understand the effect of trying on the target item with the help of the second image frame and the stress cloud map, but also intuitively understand the stress on the body parts when wearing the target item. Thus, the user can have a more comprehensive understanding of the situation when wearing the target item, which is conducive to improving the user experience. Directly superimposing the image frame of the target item and the stress cloud map onto the body parts of the first image frame can more comprehensively restore the situation when the user wears the target item. The user can more clearly understand the stress on the body parts when wearing the target item, which is conducive to further improving the user experience.

[0133] Optionally, the stress cloud map can be segmented to obtain a stress cloud map that matches the wearing part of the first user, and then the stress cloud map that matches the wearing part of the first user can be superimposed on the human body part.

[0134] In another implementation, the client device can overlay an image of the target item that the first user wants to try on onto the wearing part in each first image frame (or an updated first image frame) (i.e., display a second image frame), and display the stress cloud map corresponding to the first image frame (or the updated first image frame) outside the display area of ​​the second image frame.

[0135] For example, the client device may display a stress cloud map on the right side of the second image frame, or on the left side of the second image frame, etc. The specific implementation method can be determined based on the actual application scenario.

[0136] Optionally, the stress cloud map has a display state and a non-display state. The switching between the display state and the non-display state of the stress cloud map is triggered by a switching operation input by the second user. The non-display state of the stress cloud map can be understood as not displaying the stress cloud map. For example, an icon for receiving the second operation can be deployed on the client device. The second operation can be a switching operation between the display state and the non-display state of the stress cloud map. The client device can adjust the stress cloud map from the display state to the non-display state based on the acquired second operation, or the client device can adjust the stress cloud map from the non-display state to the display state based on the acquired second operation. For another example, when the client device receives a click operation input by the second user in the first area of ​​the stress cloud map, it is considered to have received a switching operation that triggers the switching between the display state and the non-display state of the stress cloud map. For another example, the second user can input the aforementioned switching operation by voice input, etc. It should be understood that the examples here are only to demonstrate the feasibility of this solution, and the specific implementation method can be determined in combination with the actual application scenario.

[0137] In this embodiment, the user can control the stress cloud map to switch between a display state and a non-display state by inputting a switching operation. The user can choose whether to display the stress cloud map according to their actual needs, which improves the user's flexibility when trying on the target item and helps to further improve the user experience of this solution.

[0138] Alternatively, in another implementation, the client device can display the stress value of at least one location of the first user's wearing parts in the second image frame corresponding to each first image frame (or the updated first image frame); for example, the stress value can be displayed at at least one location of the first user's wearing parts included in the second image frame; or, for example, the stress value of at least one location of the first user's wearing parts can be displayed outside the second image frame, etc. The specific display method can be determined in combination with the actual application scenario.

[0139] Alternatively, in another implementation, the client device can display a second image frame corresponding to each first image frame (or an updated first image frame). The second image frame includes a deformed image of the first user's wearing part. That is, the second image frame not only shows the target item being tried on by the first user's wearing part, but also shows the deformation caused to the wearing part by wearing the target item, thereby realizing the display of the second image frame and status information.

[0140] Alternatively, in another implementation, the client device may display a second image frame corresponding to each first image frame (or an updated first image frame), and display an arrow corresponding to displacement information at at least one position on the first user's wearing part, the arrow indicating the displacement direction at that position of the wearing part, thereby also showing the deformation of the wearing part caused by wearing the target item.

[0141] Alternatively, in another implementation, the client device can simultaneously display at least two of the aforementioned stress cloud diagram, stress value, deformation diagram, and displacement information, etc. The specific display method can be determined based on the actual application scenario.

[0142] Optionally, the transparency of the target item's image can be adjusted. For example, a second icon can be deployed on the client device to receive a third operation, which can be an adjustment of the transparency of the target item's image. The client device can increase or decrease the transparency of the target item's image based on the acquired third operation. Another example is that the user can input the adjustment of the target item's image transparency via voice input. For instance, when the client device receives the user's voice input, "Please increase the transparency of the target item you are trying on," it can increase the transparency of the target item's image. Yet another example is that when the client device receives a click operation input by the user in a second area of ​​the target item's image, it is considered to have received an adjustment of the target item's image transparency. It should be noted that when both the target item's image and the stress cloud map are superimposed on the user's wearing area, the first area and the first area can be different areas. This example is only to demonstrate the feasibility of this solution; the specific implementation method can be determined based on the actual application scenario and is not limited here.

[0143] In this embodiment, the second image frame displayed to the user is an image of the target item superimposed on the wearing area. The transparency of the target item image is set to adjustable. After the image of the target item is made transparent, the fit between the target item and the user's wearing area can be seen more intuitively, allowing the user to understand the situation when wearing the target item more clearly. If a stress cloud map is also superimposed on the user's wearing area, after the image of the target item is made transparent, the degree of fit between the target item and the user's wearing area can be seen, along with the stress on the human body part, thus providing a more comprehensive display of the situation when the user wears the target item, further improving the user experience.

[0144] To better understand this solution, please refer to Figure 7. Figure 7 is a schematic diagram of adjusting the image of the target item according to an embodiment of this application. In Figure 7, the image of the target item and a stress cloud map are superimposed on the wearing area of ​​the first user as an example. In Figure 7, the color stress cloud map has been grayscaled. Figure 7 includes two sub-schematic diagrams, left and right. In the left sub-schematic diagram of Figure 7, the transparency of the image of the target item is 0. In the right sub-schematic diagram of Figure 7, the transparency of the image of the target item is increased. By comparing the left and right sub-schematic diagrams of Figure 7, it can be seen that after increasing the transparency of the image of the target item, the fit between the target item and the wearing area can be seen more clearly, and the stress on the foot can be seen through the shoe. The user can also understand the compression caused by the target item to the wearing area when wearing the target item. Thus, the user can have a more comprehensive understanding of the situation when wearing the target item. It should be understood that the example in Figure 7 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0145] In another scenario, if the acquired first image frame is at least one video frame in a video, after obtaining at least one second image frame corresponding one-to-one with the at least one first image frame (or an updated first image frame), the client device can assemble all the second image frames corresponding to the at least one first image frame (or an updated first image frame) into a virtual wearable video for the first user. Optionally, inter-frame consistency processing can be performed on the virtual wearable video of the first user to obtain an updated virtual wearable video for the first user. Displaying at least one second image frame by the client device may include: the client device displaying the virtual wearable video (or the updated virtual wearable video) of the first user.

[0146] For example, a large model for image processing can be used to perform inter-frame consistency processing on the aforementioned user's virtual wearable video. For instance, the prompt information and the aforementioned user's virtual wearable video are input into the large model for image processing to obtain the updated user's virtual wearable video output by the large model. The aforementioned prompt information can be used to prompt improvement of inter-frame consistency between different video frames. The role of performing inter-frame consistency processing can be understood as improving the similarity between adjacent video frames, so that the transition between different video frames in the updated user's virtual wearable video frames is smoother, and after being combined into a virtual wearable video, it can also display the dynamic effect of the user wearing the target item.

[0147] Alternatively, a machine learning model specifically designed to improve the inter-frame consistency of video can be used to process the aforementioned user's virtual wearable video to obtain an updated version of the user's virtual wearable video. In other words, the function of the aforementioned machine learning model only includes improving the inter-frame consistency between different video frames.

[0148] During the process of displaying the virtual wearable video (or the updated virtual wearable video) of the first user on the client device, the status information corresponding to each first image frame (or the updated first image frame) can be displayed. That is, as the virtual wearable video is played, the client device can display different status information. The specific display method of the status information can be referred to the above description, which will not be repeated here.

[0149] Alternatively, the client device may display only the status information corresponding to a portion of the first image frames (or updated first image frames), such as displaying only the status information corresponding to the first video frame in at least one of the first image frames; for example, if the posture of the first user in the video remains unchanged, the stress cloud map corresponding to the first video frame can be continuously displayed next to the second image frame. Alternatively, the status information corresponding to the first video frame in at least one of the first image frames may be displayed, and when the posture of the first user in the video changes, the status information corresponding to the changed posture may be displayed, etc. The specific display method can be determined based on the actual application scenario, and is not limited in this embodiment.

[0150] To provide a more intuitive understanding of this solution, the following examples illustrate the method provided in this application using two specific examples: trying on shoes on the feet and trying on watches on the wrists.

[0151] 1. Try the shoes on your feet.

[0152] The client equipment can capture third image frames of the first user's feet from multiple angles. Based on the third image frames from multiple angles, a 3D model is created to obtain a second 3D model of the first user's feet. The 3D modeling algorithm used is described above and will not be repeated here.

[0153] The client device acquires a first image frame of the first user's foot. Based on the second 3D model and the first image frame of the first user's foot, a first 3D model of the first user's foot is obtained. The posture of the foot in the first 3D model is consistent with the posture of the foot in the first image frame. The client device acquires a 3D model of the shoe. Based on the first 3D model of the foot, the 3D model of the shoe, and the target viewpoint, a physics simulation engine is used to perform a physical simulation on the 3D model of the first user's foot wearing the shoe, obtaining the state information of the first user's foot when wearing the shoe. The physical simulation algorithm used by the physics simulation engine can be referred to the above description and will not be repeated here. For example, in this embodiment, the state information includes a stress cloud map.

[0154] When the client device uses a physics simulation engine to perform a physics simulation on a first 3D model of the first user's foot wearing the 3D model of the shoe, it can also render the scene of the first user's foot wearing the 3D model of the shoe on the first 3D model of the first user's foot to 2D, to obtain a virtual wearing image of the first user's foot trying on the shoe. The image of the shoe is then segmented from the virtual wearing image. For a more intuitive understanding of this solution, please refer to Figure 8. Figure 8 is a schematic diagram of an image of a target item provided in an embodiment of this application. In Figure 8, the target item being tried on is a shoe. It should be understood that this example is only for the convenience of understanding this solution.

[0155] To overlay the stress cloud map onto the foot in the first image frame, the client device can also perform image segmentation on the stress cloud map to obtain a stress cloud map corresponding to the foot. For a more intuitive understanding of this solution, please refer to Figure 9. Figure 9 is another schematic diagram of the stress cloud map provided in this embodiment. In Figure 9, the foot is taken as the wearing part, and the stress cloud map corresponding to the foot is also the shape of the foot. For an introduction to stress cloud maps, please refer to the above description; it will not be repeated here. It should be understood that this example is only for the convenience of understanding this solution.

[0156] The client device identifies the image area of ​​the foot covered by the shoe in the acquired first image frame (or the updated first image frame). For a more intuitive understanding of this solution, please refer to Figure 10. Figure 10 is another schematic diagram of the wearing part provided in the embodiment of this application. As shown in Figure 10, the image area framed by lines is the image area of ​​the foot in the first image frame (or the updated first image frame). It should be understood that this example is only for the convenience of understanding this solution.

[0157] The client device overlays the image of the shoe and the stress cloud map corresponding to the wearing part onto the image area of ​​the foot in the first image frame (or the updated first image frame), and adjusts the transparency of the shoe image to make the shoe image semi-transparent, thereby obtaining a second image frame with superimposed state information, and displays the second image frame with superimposed state information.

[0158] 2. Try on the watch on your wrist.

[0159] The client device can capture a third image frame of the first user's wrist from multiple angles. Based on the third image frames from multiple angles, a 3D model is created to obtain a second 3D model of the first user's wrist. The 3D modeling algorithm used is described above and will not be repeated here.

[0160] The client device acquires a first image frame of the first user's wrist. Based on the second 3D model and the first image frame of the first user's wrist, it obtains a first 3D model of the first user's wrist. The posture of the wrist in the first 3D model is consistent with the posture of the wrist in the first image frame. The client device acquires a 3D model of a watch and uses a physics simulation engine to perform a physical simulation on the 3D model of the first user's wrist wearing the watch, obtaining the state information of the first user's wrist when wearing the watch. The physical simulation algorithm used by the physics simulation engine can be referred to the above description and will not be repeated here. For example, in this embodiment, the state information includes a stress cloud map.

[0161] When the client device uses a physics simulation engine to perform a physical simulation of the three-dimensional model of the first user's wrist wearing the watch on the first three-dimensional model, it can also render the scene of the first user's wrist wearing the watch on the first three-dimensional model to two dimensions to obtain a virtual wearing image of the first user's wrist trying on the watch. The image of the watch is then segmented from the virtual wearing image. For a more intuitive understanding of this solution, please refer to Figure 11. Figure 11 is another schematic diagram of the image of the target item provided in the embodiment of this application. In Figure 11, the target item being tried on is a watch. It should be understood that this example is only for the convenience of understanding this solution.

[0162] To overlay the stress cloud image onto the wrist in the first image frame, the client device can further segment the stress cloud image to obtain a stress cloud image corresponding to the wrist. The client device identifies the image area of ​​the wrist covered by the watch in the acquired first image frame (or an updated first image frame). The client device overlays the watch image and the stress cloud image corresponding to the wearing part onto the wrist image area in the first image frame (or an updated first image frame), and adjusts the transparency of the watch image to make it semi-transparent, thus obtaining a second image frame overlaid with state information, which is then displayed.

[0163] Based on the embodiments corresponding to Figures 1 to 11, in order to better implement the above-mentioned solutions of the embodiments of this application, related equipment for implementing the above solutions is also provided below. Specifically, refer to Figure 12, which is a structural schematic diagram of a virtual wearable device provided in the embodiments of this application. The virtual wearable device 1200 includes: an acquisition module 1201, used to acquire a first image frame, the first image frame including an image of a human body part from a target perspective; a receiving module 1202, used to receive a wearing item confirmation request, the wearing item confirmation request being used to instruct the simulated human body part to wear a target item; and a processing module 1203, used to provide a display page, the display page being used to display the status information of the human body part wearing the target item from the target perspective, the status information being obtained based on the three-dimensional model of the human body part, the three-dimensional model of the target item, and the target perspective, the status information indicating the force or deformation of the human body part wearing the target item from the target perspective.

[0164] Optionally, the status information includes at least one of the following: a stress cloud diagram indicating the stress condition of the wearing part in the human body; a stress value at at least one location in the wearing part; a deformation diagram of the wearing part; and displacement information at at least one location in the wearing part, wherein the displacement information indicates the deformation of the wearing part.

[0165] Optionally, the display page is also used to display a second image frame, which includes an image of the wearable part in the human body part of the first image frame superimposed with an image of the target object, and the image of the target object is an image of the target object from the target viewpoint.

[0166] Optionally, the state information includes a stress contour map, which is superimposed on the second image frame.

[0167] Optionally, the stress cloud diagram has a display state and a non-display state, and the switching between the display state and the non-display state of the stress cloud diagram is triggered by a user input switching operation.

[0168] Optionally, the acquisition module 1201 is further configured to acquire multiple third image frames, which include images of human body parts from multiple perspectives; the processing module 1203 is further configured to perform three-dimensional reconstruction based on the multiple third image frames to obtain a three-dimensional model of the human body parts.

[0169] Optionally, the processing module 1203 is further configured to adjust the posture of the three-dimensional model of the human body part according to the first image frame to obtain the processed three-dimensional model of the human body part. The posture of the human body part in the processed three-dimensional model of the human body part is consistent with the posture of the human body part in the first image frame. The state information is obtained based on the processed three-dimensional model of the human body part, the three-dimensional model of the target object, and the target viewpoint.

[0170] Optionally, the receiving module 1202 is further configured to receive target item configuration information, which indicates one or more of the target item's size information, physical information, and material information; the processing module 1203 is further configured to obtain a three-dimensional model of the target item based on the target item configuration information.

[0171] Optionally, the processing module 1203 is also used to perform physical simulation of the wearing parts of the human body wearing the target object from the target perspective based on the three-dimensional model of the human body parts and the three-dimensional model of the target object, so as to obtain the state information corresponding to the first image frame.

[0172] Optionally, the body part includes at least one of the following: neck, wrist, leg, or foot.

[0173] The acquisition module 1201, receiving module 1202, and processing module 1203 can all be implemented in software or in hardware. For example, the implementation of the acquisition module 1201 will be described below. Similarly, the implementation of the receiving module 1202 and processing module 1203 can refer to the implementation of the acquisition module 1201.

[0174] As an example of a software functional unit, module 1201 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, module 1201 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed within the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0175] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0176] As an example of a hardware functional unit, the acquisition module 1201 may include at least one computing device, such as a server. Alternatively, the acquisition module 1201 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0177] The multiple computing devices included in the acquisition module 1201 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the acquisition module 1201 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the acquisition module 1201 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0178] It should be noted that, in other embodiments, the acquisition module 1201 can be used to execute any step in the method for acquiring wearable images, the receiving module 1202 can be used to execute any step in the method for acquiring wearable images, and the processing module 1203 can be used to execute any step in the method for acquiring wearable images. The steps implemented by the acquisition module 1201, the receiving module 1202, and the processing module 1203 can be specified as needed. By implementing different steps in the method for acquiring wearable images through the acquisition module 1201, the receiving module 1202, and the processing module 1203, all functions of the virtual wearable device 1200 can be realized.

[0179] This application also provides a system for acquiring wearable images. Figure 13 is a schematic diagram of a system for acquiring wearable images provided in an embodiment of this application. As shown in Figure 13, the system for acquiring wearable images includes: a virtual wearable device 1200 and a device for acquiring wearable information 1300.

[0180] The wearable information acquisition device 1300 is used to send a first image frame and a wearable item confirmation request to the virtual wearable device 1200;

[0181] The virtual wearable device 1200 is used to acquire a first image frame, which includes an image of a human body part from a target viewpoint; receive a confirmation request for wearing an item, which is used to instruct the simulated human body part to wear a target item; and provide a display page, which is used to display the status information of the human body part wearing the target item from the target viewpoint. The status information is obtained based on the three-dimensional model of the human body part, the three-dimensional model of the target item, and the target viewpoint. The status information indicates the force or deformation of the human body part wearing the target item from the target viewpoint.

[0182] The wearable information acquisition device 1300 is used to receive the display page sent by the virtual wearable device 1200 and also to display the status information.

[0183] Both the virtual wearable device 1200 and the wearable information acquisition device 1300 can be implemented in software or in hardware. For example, the implementation of the virtual wearable device 1200 will be described below. Similarly, the implementation of the wearable information acquisition device 1300 can refer to the implementation of the virtual wearable device 1200.

[0184] As an example of a software functional unit, the virtual wearable device 1200 may include code running on a computing instance. This computing instance can be at least one of a physical host (computing device), a virtual machine, a container, or other computing devices. Furthermore, the aforementioned computing device may be one or more. For example, the virtual wearable device 1200 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the application can be distributed within the same region or in different regions. The multiple hosts / virtual machines / containers used to run the code can be distributed within the same Availability Zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0185] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same VPC or across multiple VPCs. Typically, a VPC is set up within a single region. Communication between two VPCs within the same region, and between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0186] As an example of a hardware functional unit, the virtual wearable device 1200 may include at least one computing device, such as a server. Alternatively, the virtual wearable device 1200 may also be a device implemented using an ASIC or a PLD. The aforementioned PLD may be implemented using a CPLD, FPGA, GAL, or any combination thereof.

[0187] The virtual wearable device 1200 includes multiple computing devices that can be distributed within the same region or in different regions. Similarly, the multiple computing devices can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices can be distributed within the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0188] This application also provides a computing device 100. As shown in FIG14, FIG14 is a structural schematic diagram of a computing device provided in an embodiment of this application. The computing device 100 includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other through the bus 102. The computing device 100 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 100.

[0189] Bus 102 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 14, but this does not imply that there is only one bus or one type of bus. Bus 102 can include pathways for transmitting information between various components of computing device 100 (e.g., memory 106, processor 104, communication interface 108).

[0190] The processor 104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0191] Memory 106 may include volatile memory, such as random access memory (RAM). Processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0192] The memory 106 stores executable program code, which the processor 104 executes to implement the functions of the aforementioned acquisition module, receiving module, and processing module, thereby realizing the method for acquiring wearable images. In other words, the memory 106 stores instructions for executing the method for acquiring wearable images.

[0193] Alternatively, the memory 106 stores executable code, which the processor 104 executes to implement the functions of the aforementioned virtual wearable device 1200 and the wearable information acquisition device 1300, thereby realizing the method for acquiring wearable images. That is, the memory 106 stores instructions for executing the method for acquiring wearable images.

[0194] The communication interface 108 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 100 and other devices or communication networks.

[0195] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0196] As shown in Figure 15, which is a schematic diagram of a computing device cluster provided in an embodiment of this application, the computing device cluster includes at least one computing device 100. The memory 106 in one or more computing devices 100 in the computing device cluster may store the same instructions for executing a method for acquiring wearable images.

[0197] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the method of acquiring wearable images. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for executing the method of acquiring wearable images.

[0198] It should be noted that the memory 106 in different computing devices 100 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the virtual wearable device 1200. That is, the instructions stored in the memory 106 of different computing devices 100 can implement the functions of one or more modules among the acquisition module, receiving module, and processing module.

[0199] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 16 illustrates one possible implementation, and is a schematic diagram of computer devices in a computer cluster provided in this application embodiment connected via a network. As shown in Figure 16, two computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 106 in computing device 100A stores instructions for performing the functions of the acquisition module and the receiving module. Simultaneously, the memory 106 in computing device 100B stores instructions for performing the functions of the processing module.

[0200] The connection method between the computing device clusters shown in Figure 16 can be considered as follows: taking into account that the wearable image acquisition method provided in this application requires the independent management of the processing model that consumes a lot of computer resources, the function implemented by the processing module is considered to be executed by the computing device 100B.

[0201] It should be understood that the functions of computing device 100A shown in Figure 16 can also be performed by multiple computing devices 100. Similarly, the functions of computing device 100B can also be performed by multiple computing devices 100.

[0202] This application embodiment also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection method of the computing device clusters described in Figures 15 and 16. The difference is that the memory 106 in one or more computing devices 100 in this computing device cluster can store the same instructions for executing the method of acquiring wearable images.

[0203] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the method of acquiring wearable images. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for executing the method of acquiring wearable images.

[0204] It should be noted that the memory 106 in different computing devices 100 within the computing device cluster can store different instructions for executing some functions of the wearable image acquisition system. That is, the instructions stored in the memory 106 of different computing devices 100 can implement the functions of one or more of the virtual wearable device 1200 and the wearable information acquisition device 1300.

[0205] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform a method for acquiring wearable images.

[0206] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to perform a method for acquiring wearable images.

[0207] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A virtual wearable method, characterized in that, The method includes: Acquire a first image frame, the first image frame including an image of a human body part from the target viewpoint; Receive a confirmation request for wearing an item, the confirmation request for wearing an item being used to instruct the simulated human body part to wear the target item; A display page is provided to display the status information of the human body part wearing the target item from the target perspective. The status information is obtained based on the three-dimensional model of the human body part, the three-dimensional model of the target item, and the target perspective. The status information indicates the force or deformation of the human body part wearing the target item from the target perspective.

2. The method according to claim 1, characterized in that, The status information includes at least one of the following: Stress cloud diagram used to indicate the stress conditions of the wearable parts in the human body; The stress value at at least one location in the wearable part; Deformation diagram of the wearable part; Displacement information at at least one location of the wearable part, the displacement information indicating the deformation of the wearable part.

3. The method according to claim 1 or 2, characterized in that, The display page is also used to display a second image frame, which includes an image of the wearable part of the human body in the first image frame superimposed with an image of the target item, wherein the image of the target item is an image of the target item from the target perspective.

4. The method according to claim 3, characterized in that, The status information includes a stress cloud map, which is superimposed on the second image frame.

5. The method according to claim 4, characterized in that, The stress cloud diagram has a display state and a non-display state, and the switching between the display state and the non-display state of the stress cloud diagram is triggered by a user input switching operation.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Acquire multiple third image frames, the multiple third image frames including images of the human body parts from multiple viewpoints; Three-dimensional reconstruction is performed based on the multiple third image frames to obtain a three-dimensional model of the human body part.

7. The method according to claim 6, characterized in that, The method further includes: Based on the first image frame, the pose of the three-dimensional model of the human body part is adjusted to obtain the processed three-dimensional model of the human body part. The pose of the human body part in the processed three-dimensional model of the human body part is consistent with the pose of the human body part in the first image frame. The state information is obtained based on the processed three-dimensional model of the human body part, the three-dimensional model of the target object, and the target viewpoint.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Receive target item configuration information, wherein the target item configuration information is used to indicate one or more of the target item's size information, physical information, and material information; Based on the target item configuration information, a three-dimensional model of the target item is obtained.

9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: based on the three-dimensional model of the human body part and the three-dimensional model of the target object, performing a physical simulation of the wearing part of the human body part wearing the target object from the target perspective, to obtain the state information corresponding to the first image frame.

10. The method according to any one of claims 1 to 9, characterized in that, The body part includes at least one of the following: neck, wrist, leg, or foot.

11. A virtual wearable device, characterized in that, The device includes: The acquisition module is used to acquire a first image frame, wherein the first image frame includes an image of a human body part from the target viewpoint; The receiving module is used to receive a confirmation request for wearing an item, wherein the confirmation request for wearing an item is used to instruct the simulated human body part to wear a target item; The processing module is used to provide a display page, which is used to display the status information of the human body part wearing the target item from the target perspective. The status information is obtained based on the three-dimensional model of the human body part, the three-dimensional model of the target item, and the target perspective. The status information indicates the force or deformation of the human body part wearing the target item from the target perspective.

12. The apparatus according to claim 11, characterized in that, The status information includes at least one of the following: Stress cloud diagram used to indicate the stress conditions of the wearable parts in the human body; The stress value at at least one location in the wearable part; Deformation diagram of the wearable part; Displacement information at at least one location of the wearable part, the displacement information indicating the deformation of the wearable part.

13. The apparatus according to claim 11 or 12, characterized in that, The display page is also used to display a second image frame, which includes an image of the wearable part of the human body in the first image frame superimposed with an image of the target item, wherein the image of the target item is an image of the target item from the target perspective.

14. The apparatus according to claim 13, characterized in that, The status information includes a stress cloud map, which is superimposed on the second image frame.

15. The apparatus according to claim 14, characterized in that, The stress cloud diagram has a display state and a non-display state, and the switching between the display state and the non-display state of the stress cloud diagram is triggered by a user input switching operation.

16. The apparatus according to any one of claims 11 to 15, characterized in that, The acquisition module is further configured to acquire multiple third image frames, the multiple third image frames including images of the human body parts from multiple perspectives; The processing module is also used to perform three-dimensional reconstruction based on the plurality of third image frames to obtain a three-dimensional model of the human body part.

17. The apparatus according to claim 16, characterized in that, The processing module is further configured to adjust the posture of the three-dimensional model of the human body part according to the first image frame to obtain a processed three-dimensional model of the human body part. The posture of the human body part in the processed three-dimensional model of the human body part is consistent with the posture of the human body part in the first image frame. The state information is obtained based on the processed three-dimensional model of the human body part, the three-dimensional model of the target object, and the target viewpoint.

18. The apparatus according to any one of claims 11 to 17, characterized in that, The receiving module is further configured to receive target item configuration information, wherein the target item configuration information is used to indicate one or more of the target item's size information, physical information, and material information; The processing module is also used to obtain a three-dimensional model of the target item based on the target item configuration information.

19. The apparatus according to any one of claims 11 to 18, characterized in that, The processing module is further configured to perform physical simulation of the wearing part of the human body wearing the target item from the target perspective based on the three-dimensional model of the human body part and the three-dimensional model of the target item, so as to obtain the state information corresponding to the first image frame.

20. The apparatus according to any one of claims 11 to 19, characterized in that, The body part includes at least one of the following: neck, wrist, leg, or foot.

21. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 10.

22. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1 to 10.

23. A computer-readable storage medium, characterized in that, Includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Wearing effect determination method and device and electronic equipment

    CN113240819A

  • Image processing method and device, electronic equipment and storage medium

    CN115965826A

  • Virtual fitting system

    JP2020097803A