Method and device for generating virtual image

By combining real leg images with three-dimensional shoe models, using machine learning and transparency processing, the occluded areas are accurately divided, solving the problem of poor virtual image effects in virtual shoe-trying technology and achieving more realistic virtual image generation.

CN112330784BActive Publication Date: 2025-09-16BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011134938.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-21
Publication Date
2025-09-16
Estimated Expiration
2040-10-21

AI Technical Summary

Technical Problem

In existing virtual shoe-trying technologies, the virtual images of human feet wearing shoes generated through approximate modeling lack real three-dimensional structural information, resulting in poor results.

Method used

A virtual image is synthesized using real leg images and three-dimensional shoe models. The leg area is extracted through a machine learning model, and combined with the transparent three-dimensional shoe model outline, the occluded area is accurately divided to generate a virtual image with a leg occlusion effect.

Benefits of technology

The authenticity and effect of virtual images are improved, and the spatial layering and realism of virtual images are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112330784B_ABST
    Figure CN112330784B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and device for generating a virtual image, and relates to the field of computer technology. The method comprises: obtaining a leg region from an image to be processed that includes the leg and foot; rendering a three-dimensional shoe model into a two-dimensional shoe image corresponding to the image to be processed based on the posture parameters of the foot in the image to be processed and the internal parameters of a camera; determining the portion of the shoe region that is occluded by the leg based on the overlap between the leg region and the shoe region in the two-dimensional shoe image; and rendering a composite image of the image to be processed and the two-dimensional shoe image based on the occluded portion to generate a virtual image with a leg occlusion effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a method for generating a virtual image, a device for generating a virtual image, and a non-volatile computer-readable storage medium. Background Art

[0002] With the advancement of VR (Virtual Reality) and AR (Augmented Reality) technologies, the ability to convert shoes into purchases through virtual try-ons is becoming increasingly popular. Virtual shoe try-ons, enabled by combining AR with smartphone cameras, allow users to see how a shoe will look on their feet.

[0003] In order to achieve the visual effect of human feet wearing shoes, virtual shoe-trying technology requires virtual occlusion of the three-dimensional shoe model to replace the real shoes on the human feet.

[0004] In the related art, a virtual leg model is modeled outside the shoe opening area of ​​a three-dimensional shoe model, and the virtual leg model and the three-dimensional shoe model are rendered on a screen to generate a virtual image of a human foot wearing shoes. Summary of the Invention

[0005] The inventors of the present disclosure have discovered that the above-mentioned related technologies have the following problems: virtual images of human feet wearing shoes are generated only through approximate modeling, which lacks real three-dimensional structural information, resulting in poor results in generating virtual images.

[0006] In view of this, the present disclosure proposes a technical solution for generating a virtual image, which can synthesize a virtual image by using a real leg image and a three-dimensional shoe model, thereby improving the effect of the virtual image.

[0007] According to some embodiments of the present disclosure, a method for generating a virtual image is provided, including: obtaining a leg area in an image to be processed that includes legs and feet; rendering a three-dimensional shoe model into a two-dimensional shoe image corresponding to the image to be processed based on posture parameters of the foot in the image to be processed and internal parameters of a camera; determining a portion of the shoe area that is obscured by the legs based on an overlapping portion between the leg area and the shoe area in the two-dimensional shoe image; and rendering a composite image of the image to be processed and the two-dimensional shoe image based on the portion obscured by the legs to generate a virtual image with a leg occlusion effect.

[0008] In some embodiments, determining the portion of the shoe area obscured by the legs based on the overlapping portion of the leg area and the shoe area in the two-dimensional shoe image includes: determining the outer contour of the shoe based on the position of the shoe body area in the two-dimensional shoe image, and determining the inner contour of the shoe based on the position of the shoe opening area in the two-dimensional image; determining the portion of the shoe area obscured by the legs based on the intersection of the contour of the leg area with the outer contour and the inner contour.

[0009] In some embodiments, determining the portion of the shoe area obscured by the legs based on the intersection of the outline of the leg area with the outer outline and the inner outline includes: determining the point on the inner outline that is closest to the intersection; and determining the portion of the shoe area obscured by the legs based on the intersection and the point that is closest.

[0010] In some embodiments, rendering a three-dimensional shoe model into a two-dimensional shoe image corresponding to the image to be processed includes: making the shoe opening area in the three-dimensional shoe model transparent; rendering the three-dimensional shoe model after transparency into a two-dimensional shoe image; and determining the shoe body area and the shoe opening area based on the binary image of the two-dimensional shoe image.

[0011] In some embodiments, making the shoe opening area in the three-dimensional shoe model transparent includes: detecting the shoe opening area of ​​the three-dimensional shoe model, covering the shoe opening area with a closed grid, and making the portion covered by the closed grid transparent.

[0012] In some embodiments, obtaining a leg region in an image to be processed that includes legs and feet includes: inputting the image to be processed into a machine learning model to determine the leg region in the image to be processed.

[0013] In some embodiments, the machine learning model includes a convolutional neural network module and a spatial pyramid pooling module connected in sequence.

[0014] In some embodiments, the convolutional neural network module is set according to a Fast-SCNN (Fast Segmentation Convolutional Neural Network) model.

[0015] In some embodiments, the images to be processed are frames of images in a video; the generation method further includes: generating a video with a leg masking effect based on virtual images corresponding to the generated frames of images.

[0016] According to other embodiments of the present disclosure, a device for generating a virtual image is provided, including: a determination unit for obtaining a leg area in an image to be processed that includes the leg and the foot, and determining a portion of the shoe area that is occluded by the leg based on an overlapping portion between the leg area and the shoe area in a two-dimensional shoe image; a processing unit for rendering a three-dimensional shoe model into a two-dimensional shoe image corresponding to the image to be processed based on posture parameters of the foot in the image to be processed and internal parameters of a camera, and rendering a composite image of the image to be processed and the two-dimensional shoe image based on the portion occluded by the leg, to generate a virtual image with a leg occlusion effect.

[0017] In some embodiments, the determination unit determines the outer contour of the shoe based on the position of the shoe body area in the two-dimensional shoe image, and determines the inner contour of the shoe based on the position of the shoe mouth area in the two-dimensional image; and determines the part of the shoe area obscured by the leg based on the intersection of the contour of the leg area with the outer contour and the inner contour.

[0018] In some embodiments, the determination unit determines, on the inner contour, a point closest to the intersection point; and determines a portion of the shoe region blocked by the leg based on the intersection point and the point closest to the intersection point.

[0019] In some embodiments, the processing unit performs transparency processing on the shoe opening area of ​​the three-dimensional shoe model and renders the three-dimensional shoe model after transparency processing into the two-dimensional shoe image. The determination unit determines the shoe body area and the shoe opening area based on the binary image of the two-dimensional shoe image.

[0020] In some embodiments, the processing unit detects the shoe opening area of ​​the three-dimensional shoe model, covers the shoe opening area with a closed grid, and performs transparency processing on the portion covered by the closed grid.

[0021] In some embodiments, the determining unit determines the position of the closed grid covering the shoe opening area according to the size of the portion of the preset shoe area that is not blocked by the leg.

[0022] In some embodiments, the determination unit inputs the image to be processed into a machine learning model to determine the leg area in the image to be processed.

[0023] In some embodiments, the machine learning model includes a convolutional neural network module and a spatial pyramid pooling module connected in sequence.

[0024] In some embodiments, the convolutional neural network module is configured according to a Fast-SCNN model.

[0025] In some embodiments, the images to be processed are frames of images in a video; the processing unit generates a video with a leg masking effect based on the generated virtual images corresponding to the frames of images.

[0026] According to some further embodiments of the present disclosure, a device for generating a virtual image is provided, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the method for generating a virtual image in any of the above embodiments based on instructions stored in the memory device.

[0027] According to some further embodiments of the present disclosure, a non-volatile computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method for generating a virtual image in any of the above embodiments is implemented.

[0028] In the above embodiment, the virtual occlusion portion is accurately determined based on the visual cues of the real leg and the position of the shoe in the 2D shoe image. In this way, the real leg image and the 3D shoe model can be combined to form a virtual image, improving the quality of the virtual image. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0030] The present disclosure can be more clearly understood from the following detailed description with reference to the accompanying drawings:

[0031] Figure 1 A flowchart illustrating some embodiments of the method for generating a virtual image of the present disclosure;

[0032] Figure 2 Show Figure 1 Flowcharts of some embodiments of step 130;

[0033] Figures 3a to 3c Schematic diagrams showing some embodiments of the method for generating a virtual image of the present disclosure;

[0034] Figure 4 A block diagram illustrating some embodiments of a device for generating a virtual image according to the present disclosure;

[0035] Figure 5 A block diagram showing some other embodiments of the apparatus for generating a virtual image according to the present disclosure;

[0036] Figure 6 A block diagram showing some further embodiments of the apparatus for generating a virtual image according to the present disclosure. DETAILED DESCRIPTION

[0037] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure.

[0038] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0039] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.

[0040] Technologies, methods and equipment known to ordinary technicians in the relevant field may not be discussed in detail, but where appropriate, technologies, methods and equipment should be considered part of the authorization specification.

[0041] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0042] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0043] As mentioned above, the virtual leg model is fixed relative to the shoe model and cannot reflect the actual three-dimensional structure of the leg (such as the user's actual trouser position, shape, posture, leg position, etc.). This will lead to a decrease in the quality of the virtual image.

[0044] To address the above technical issues, visual cues of the legs can be extracted from the real scene (e.g., using a neural network model) and combined with the outline of the 3D shoe model after the shoe opening area is made transparent, the area that needs to be blocked in the virtual image can be accurately demarcated. For example, the technical solution of this disclosure can be implemented through the following embodiments.

[0045] Figure 1 A flowchart illustrating some embodiments of the method for generating a virtual image of the present disclosure.

[0046] like Figure 1 As shown, the production method includes: step 110, obtaining a leg area; step 120, rendering a two-dimensional shoe image; step 130, determining a portion blocked by the leg; and step 140, generating a virtual image.

[0047] In step 110 , a leg region in an image to be processed containing legs and feet is obtained.

[0048] In some embodiments, the image to be processed is input into a machine learning model to determine the leg region in the image to be processed. For example, the machine learning model can be a convolutional neural network module (e.g., configured according to a Fast-SCNN model).

[0049] In this way, a skeleton network for leg region extraction is constructed based on the lightweight Fast-SCNN model. A simple yet efficient composite coefficient allows for a more structured network architecture, compressing the model's parameters and increasing training speed. Furthermore, Fast-SCNN reduces the model's floating-point operations, thereby improving computational performance.

[0050] In some embodiments, the machine learning model includes a convolutional neural network module and an SPP (Spatial Pyramid Pooling) module connected in sequence. For example, the SPP module includes a convolution processing module, an upsampling module, and a concatenation module connected in sequence.

[0051] In some embodiments, the machine learning model includes a convolutional neural network module, a first SPP module, and a second SPP module. The convolutional neural network module is connected to the convolution processing module and the connection module of the first SPP module, and is connected to the connection module of the second SPP module; the first SPP module is connected to the second SPP module.

[0052] In this way, the SPP module can well maintain the complete inter-context information and avoid misclassification in image processing. Moreover, the SPP module has better robustness for small-sized and insignificant object recognition, and can pay attention to different sub-regions containing insignificant objects, thereby improving the accuracy of leg region recognition.

[0053] In some embodiments, the above-mentioned machine learning model can be trained using SoftMax Loss to set the loss function.

[0054] In step 120, the 3D shoe model is rendered into a 2D shoe image corresponding to the image to be processed based on the foot posture parameters in the image to be processed and the camera's internal parameters. For example, the camera's internal parameters are parameters related to the camera's own characteristics, such as the camera's focal length and pixel size.

[0055] In some embodiments, the PnP (Perspective-n-Point) algorithm can be used to determine the posture parameters.

[0056] In some embodiments, a rendering tool such as OpenGL may be used to render the three-dimensional shoe model into a two-dimensional shoe image.

[0057] In some embodiments, the shoe opening area of ​​the 3D shoe model is detected and covered with a closed mesh as a barrier. The covered portion of the closed mesh is rendered transparent. For example, a transparent mesh serving as a barrier can be provided on the 3D shoe model and placed inside the shoe opening area to render the shoe opening area transparent.

[0058] In some embodiments, the position of the closed grid covering the shoe opening area is determined based on the size of the portion of the shoe area not blocked by the leg. For example, the blocking piece can be moved a predetermined distance toward the sole so that the edge thickness of the shoe opening area exceeds a threshold.

[0059] In this way, according to the different depths of the blocking piece positions, the uncropped portion of the shoe opening area has a certain thickness, which can increase the spatial layering of the virtual image and enhance the real effect.

[0060] In step 130 , the portion of the shoe region blocked by the leg is determined based on the overlapping portion of the leg region and the shoe region in the two-dimensional shoe image.

[0061] In some embodiments, the deep learning model can be used to accurately segment the leg area, thereby determining the boundary area between the leg model and the shoe model, and then accurately obtaining the occluded part of the shoe opening area as the cropping area. After rendering and display on the screen, a virtual occlusion visual effect will be produced. For example, Figure 2 The embodiment in implements step 130.

[0062] Figure 2 Show Figure 1 Flowchart of some embodiments of step 130 in FIG.

[0063] like Figure 2 As shown, step 130 includes: step 1310, determining the inner contour and outer contour of the shoe; and step 1320, determining the portion blocked by the leg.

[0064] In step 1310, the outer contour of the shoe is determined according to the position of the shoe body region in the two-dimensional shoe image, and the inner contour of the shoe is determined according to the position of the shoe opening region in the two-dimensional image.

[0065] In some embodiments, the Figure 3a In the embodiment, the shoe body area and the shoe opening area are determined to determine the inner contour and the outer contour.

[0066] Figure 3a Schematic diagrams showing some embodiments of the method for generating a virtual image according to the present disclosure.

[0067] like Figure 3a As shown, the shoe opening area in the three-dimensional shoe model is made transparent and then rendered into a two-dimensional shoe image. The shoe body area 31 and the shoe opening area 32 are determined based on the binary image of the two-dimensional shoe image.

[0068] After the shoe body area 31 and the shoe opening area 32 are determined, the outer contour and the inner contour can be determined, and then Figure 2 The remaining steps in determine the occluded part.

[0069] In step 1320 , the portion of the shoe region that is blocked by the leg is determined based on the intersection of the outline of the leg region and the outer outline and the inner outline.

[0070] In some embodiments, the Figure 3b 、 3cIn the embodiment, the shoe body area and the shoe opening area are determined to determine the inner contour and the outer contour.

[0071] Figure 3b Schematic diagrams showing some embodiments of the method for generating a virtual image according to the present disclosure.

[0072] like Figure 3b As shown, the segmentation mask binary image of the leg area 30 can be inferred by using the neural network module. Figure 3a The binary image of the shoe and Figure 3b The binary image of the leg in FIG3 can determine the intersection of the outline of the leg area 30 and the outline of the shoe body area 31.

[0073] Figure 3c Schematic diagrams showing some embodiments of the method for generating a virtual image according to the present disclosure.

[0074] like Figure 3c As shown, the outline of the shoe body area 31 is the outer outline 311, and the outline of the shoe opening area 32 is the inner outline 321. The intersection points of the outline of the leg area 30 and the outer outline 311 are 3a and 3b, and the point on the inner outline 321 closest to the intersection 3a is 3d, and the point closest to the intersection 3b is 3c.

[0075] The connection points 3a, 3b, 3c, 3d form a closed area (located on the left side of the shoe body), which is defined as the portion of the shoe area that is hidden by the leg.

[0076] In some embodiments, the portion of the shoe region obscured by the leg may also be determined based on the intersection of the outline of the leg region 30 and the inner outline 321 and the closed region formed by 3a and 3b.

[0077] The part of the shoe area that is blocked by the leg is determined by Figure 1 Step 140 in the process generates a virtual image.

[0078] In step 140 , a composite image of the image to be processed and the two-dimensional shoe image is rendered according to the portion blocked by the leg, to generate a virtual image with a leg blocking effect.

[0079] In some embodiments, the images to be processed are frames of images in a video. A video with a leg masking effect can be generated based on the generated virtual images corresponding to the frames of images.

[0080] Figure 4 A block diagram illustrating some embodiments of an apparatus for generating a virtual image according to the present disclosure.

[0081] like Figure 4 As shown, the virtual image generating device 4 includes a determining unit 41 and a processing unit 42 .

[0082] The determining unit 41 obtains a leg region in the image to be processed that includes the leg and the foot; and determines a portion of the shoe region that is blocked by the leg based on an overlapping portion between the leg region and the shoe region in the two-dimensional shoe image.

[0083] In some embodiments, the determination unit 41 determines the outer contour of the shoe based on the position of the shoe body area in the two-dimensional shoe image, and determines the inner contour of the shoe based on the position of the shoe mouth area in the two-dimensional image; and determines the part of the shoe area blocked by the leg based on the intersection of the contour of the leg area with the outer contour and the inner contour.

[0084] In some embodiments, the determination unit 41 determines the point on the inner contour that is closest to the intersection point; and determines the portion of the shoe area that is blocked by the leg based on the intersection point and the point that is closest.

[0085] In some embodiments, the determination unit 41 inputs the image to be processed into a machine learning model to determine the leg area in the image to be processed.

[0086] In some embodiments, the machine learning model includes a convolutional neural network module and a spatial pyramid pooling module connected in sequence.

[0087] In some embodiments, the convolutional neural network module is configured according to a Fast-SCNN model.

[0088] The processing unit 42 renders the three-dimensional shoe model into a two-dimensional shoe image corresponding to the image to be processed based on the posture parameters of the foot in the image to be processed and the internal parameters of the camera; and renders a composite image of the image to be processed and the two-dimensional shoe image based on the part occluded by the legs to generate a virtual image with a leg occlusion effect.

[0089] In some embodiments, the processing unit 42 performs transparency processing on the shoe opening area of ​​the 3D shoe model and renders the 3D shoe model after transparency processing into the 2D shoe image. The determination unit 41 determines the shoe body area and the shoe opening area based on the binary image of the 2D shoe image.

[0090] In some embodiments, the processing unit 42 detects the shoe opening area of ​​the three-dimensional shoe model, covers the shoe opening area with a closed grid, and performs transparency processing on the portion covered by the closed grid.

[0091] In some embodiments, the determining unit 41 determines the position of the closed grid covering the shoe opening area according to the size of the portion of the preset shoe area that is not blocked by the leg.

[0092] In some embodiments, the images to be processed are frames of images in a video; the processing unit 42 generates a video with a leg masking effect based on the generated virtual images corresponding to the frames of images.

[0093] Figure 5A block diagram showing some other embodiments of the apparatus for generating a virtual image according to the present disclosure.

[0094] like Figure 5 As shown, the virtual image generating device 5 of this embodiment includes: a memory 51 and a processor 52 coupled to the memory 51, and the processor 52 is configured to execute the virtual image generating method in any embodiment of the present disclosure based on the instructions stored in the memory 51.

[0095] The memory 51 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory may store, for example, an operating system, application programs, a boot loader, a database, and other programs.

[0096] Figure 6 A block diagram showing some further embodiments of the apparatus for generating a virtual image according to the present disclosure.

[0097] like Figure 6 As shown, the virtual image generation device 6 of this embodiment includes: a memory 610 and a processor 620 coupled to the memory 610, and the processor 620 is configured to execute the virtual image generation method in any of the aforementioned embodiments based on the instructions stored in the memory 610.

[0098] The memory 610 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory may store, for example, an operating system, application programs, a boot loader, and other programs.

[0099] The virtual image generation device 6 may further include an input / output interface 630, a network interface 640, a storage interface 650, and the like. These interfaces 630, 640, 650, as well as the memory 610 and the processor 620, may be connected, for example, via a bus 660. The input / output interface 630 provides a connection interface for input / output devices such as a display, mouse, keyboard, touch screen, microphone, and speakers. The network interface 640 provides a connection interface for various networked devices. The storage interface 650 provides a connection interface for external storage devices such as SD cards and USB flash drives.

[0100] Those skilled in the art will appreciate that embodiments of the present disclosure may be provided as methods, systems, or computer program products. Thus, the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable non-transitory storage media, including but not limited to magnetic disk storage, CD-ROMs, optical storage, and the like, containing computer-usable program code.

[0101] The virtual image generation method, virtual image generation device, and non-volatile computer-readable storage medium according to the present disclosure have been described in detail. To avoid obscuring the concepts of the present disclosure, some details known in the art have been omitted. Based on the above description, those skilled in the art will fully understand how to implement the technical solutions disclosed herein.

[0102] The methods and systems of the present disclosure may be implemented in many ways. For example, the methods and systems of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above unless otherwise specified. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the methods according to the present disclosure. Therefore, the present disclosure also covers recording media that store programs for executing the methods according to the present disclosure.

[0103] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will appreciate that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A method for generating a virtual image, comprising: Obtaining a leg region in the image to be processed that includes the leg and the foot; Rendering the three-dimensional shoe model into a two-dimensional shoe image corresponding to the image to be processed according to the posture parameters of the foot in the image to be processed and the internal parameters of the camera; determining a portion of the shoe region blocked by the leg according to an overlapping portion of the leg region and the shoe region in the two-dimensional shoe image; Rendering a composite image of the image to be processed and the two-dimensional shoe image according to the portion blocked by the leg to generate a virtual image with a leg blocking effect; The determining, based on an overlapping portion of the leg region and the shoe region in the two-dimensional shoe image, a portion of the shoe region blocked by the leg comprises: determining an outer contour of the shoe according to a position of a shoe body region in the two-dimensional shoe image, and determining an inner contour of the shoe according to a position of a shoe opening region in the two-dimensional shoe image; determining a portion of the shoe region obscured by the leg according to an intersection of the outline of the leg region with the outer outline and the inner outline; The step of determining the portion of the shoe region that is blocked by the leg according to the intersection of the outline of the leg region with the outer outline and the inner outline comprises: Determine, on the inner contour, a point closest to the intersection point; The portion of the shoe area blocked by the leg is determined according to the intersection point and the point with the shortest distance.

2. The generation method according to claim 1, wherein: Rendering the three-dimensional shoe model into a two-dimensional shoe image corresponding to the image to be processed includes: Performing transparency processing on the shoe opening area in the three-dimensional shoe model; Rendering the three-dimensional shoe model after the transparency processing into the two-dimensional shoe image; A shoe body region and a shoe opening region are determined according to the binary image of the two-dimensional shoe image.

3. The generation method according to claim 2, wherein: The transparent processing of the shoe opening area in the three-dimensional shoe model includes: Detect the shoe opening area of ​​the 3D shoe model and cover the shoe opening area with a closed mesh; The closed grid covered portion is made transparent.

4. The generation method according to claim 3, wherein: The detecting of the shoe opening area of ​​the three-dimensional shoe model and covering the shoe opening area with a closed grid includes: The position where the closed grid covers the shoe opening area is determined according to the size of a preset portion of the shoe area that is not blocked by the leg.

5. The generation method according to any one of claims 1 to 4, wherein: The step of obtaining a leg region in an image to be processed containing legs and feet includes: The image to be processed is input into a machine learning model to determine a leg area in the image to be processed.

6. The generation method according to claim 5, wherein: The machine learning model includes a convolutional neural network module and a spatial pyramid pooling module connected in sequence.

7. The generation method according to claim 6, wherein: The convolutional neural network module is set according to the fast segmentation convolutional neural network Fast-SCNN model.

8. The generation method according to any one of claims 1 to 4, wherein: The images to be processed are frames of images in the video; Also includes: A video with a leg masking effect is generated based on the generated virtual images corresponding to the respective frame images.

9. A device for generating a virtual image, comprising: a determining unit, configured to obtain a leg region in the image to be processed that includes the leg and the foot, and determine a portion of the shoe region that is blocked by the leg based on an overlap between the leg region and the shoe region in the two-dimensional shoe image; a processing unit configured to render the three-dimensional shoe model into a two-dimensional shoe image corresponding to the image to be processed based on the posture parameters of the foot in the image to be processed and the internal parameters of the camera, and to render a composite image of the image to be processed and the two-dimensional shoe image based on the portion occluded by the leg, thereby generating a virtual image with a leg occlusion effect; The determining unit determines the outer contour of the shoe according to the position of the shoe body region in the two-dimensional shoe image, determines the inner contour of the shoe according to the position of the shoe opening region in the two-dimensional shoe image, and determines the portion of the shoe region blocked by the leg according to the intersection of the contour of the leg region with the outer contour and the inner contour. The determining unit determines, on the inner contour, a point closest to the intersection point; and determines a portion of the shoe region blocked by the leg according to the intersection point and the point closest to the intersection point.

10. A device for generating a virtual image, comprising: Memory; and A processor coupled to the memory, wherein the processor is configured to execute the method for generating a virtual image according to any one of claims 1 to 8 based on instructions stored in the memory.

11. A non-volatile computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method for generating a virtual image according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • AR imaging virtual shoe trying method and device capable of processing local shelter

    CN111369686A

  • Method for virtually trying on footwear

    EP2647305A1