System and method for virtual fitting during live broadcast

By trying on the product 3D model onto the user's 3D model and rendering it according to the presenter's pose in the streaming data, the problem of users' unimaginable product appearance on themselves is solved, and a more accurate product evaluation and user experience is achieved.

CN120017918AActive Publication Date: 2025-05-16GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510101413.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-03-10
Filing Date
2021-02-26
Publication Date
2025-05-16
Estimated Expiration
2041-02-26

AI Technical Summary

Technical Problem

When purchasing clothing or wearable products online, users find it difficult to imagine the appearance or performance of the product on themselves, and the existing technology cannot effectively solve this problem.

Method used

By evaluating the streaming data, obtain the user's 3D model, and try the product model onto the user's model, place the model according to the presenter's pose in the streaming data, render it and present it to the user with the streaming data.

Benefits of technology

It provides users with more accurate evaluation of the appearance of clothing or wearable products on themselves, enhancing users' experience of products and the basis for purchasing decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017918A_ABST
    Figure CN120017918A_ABST
Patent Text Reader

Abstract

Methods and systems for enhancing streaming media data to include virtual fitting data are described herein. The method includes evaluating the streaming data to obtain a first three-dimensional (3D) model associated with the product. The method also includes obtaining a second 3D model associated with the user. The first 3D model is then tried on to the second 3D model and the tried-on model is gestures in a manner estimated from the presenter in the streaming data. The postulated model is then rendered and presented to the viewer along with the streaming data. The embodiment of the invention is suitable for various virtual reality applications and computer-based fitting systems.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Description of the case

[0002] This application is a divisional application of the Chinese patent with an application date of February 26, 2021, application number 202180014731.X, and invention name “System and method for virtual fitting during live broadcast”.

[0003] CROSS-REFERENCE TO RELATED APPLICATIONS

[0004] This application is based on the U.S. provisional patent application with application number 62 / 987,474, application date March 10, 2020, and application name “System and method for virtual fitting during live broadcast”, and claims priority of the U.S. provisional patent application. The entire contents of the U.S. provisional patent application are hereby incorporated into this application by introduction. Technical Field

[0005] The present invention generally relates to methods and systems related to virtual fitting applications. More specifically, embodiments of the present invention provide methods and systems for enhancing streaming media data to include virtual fitting data. Background Art

[0006] When considering purchasing a product online, users often struggle to visualize how it will look or perform. This is particularly true for clothing and other wearable products. To better understand a product's attributes, users often view images and / or videos of a person evaluating and reviewing the product. These images and videos feature the product on the person, who discusses its various advantages and disadvantages. However, because people have different body types, even seeing the product on a person may not help users fully visualize how it will look or perform on themselves.

[0007] Embodiments of the present invention address these and other problems, individually and collectively. Summary of the Invention

[0008] The method includes evaluating streaming data to obtain a first three-dimensional (3D) model associated with a product. The method also includes obtaining a second 3D model associated with a user. The first 3D model is then tried on the second 3D model and the tried-on model is posed in a manner estimated from a presenter in the streaming data. The posed model is then rendered and presented to a viewer along with the streaming data. Embodiments of the present invention are suitable for various applications in virtual reality and computer-based fitting systems.

[0009] One embodiment of the present invention relates to a method comprising receiving an indication of media content that a user is viewing, identifying a product associated with the media content, obtaining a first 3D model representing the product, obtaining a second 3D model representing the user, determining a presentation pose based on the media content, applying the presentation pose to the second 3D model, generating a third 3D model by having the second 3D model try on the first 3D model, and presenting the third 3D model to the user in the presentation pose.

[0010] Another embodiment of the present invention relates to a system comprising a processor and a memory comprising instructions that, when executed by the processor, cause the system to at least receive an indication of media content being viewed by a user, identify a product associated with the media content, obtain a first 3D model representing the product, obtain a second 3D model representing the user, determine a presentation pose based on the media content, apply the presentation pose to the second 3D model, generate a third 3D model by having the second 3D model try on the first 3D model, and present the third 3D model to the user in the presentation pose.

[0011] Yet another embodiment of the present disclosure relates to a non-transitory computer-readable medium storing specific computer-executable instructions that, when executed by a processor, cause a computer system to at least receive an indication of media content being viewed by a user, identify a product associated with the media content, obtain a first 3D model representing the product, obtain a second 3D model representing the user, determine a presentation pose based on the media content, apply the presentation pose to the second 3D model, generate a third 3D model by having the second 3D model try on the first 3D model, and present the third 3D model to the user in the presentation pose.

[0012] The present system has several advantages over conventional systems. For example, embodiments of the present invention relate to methods and systems for providing a user with a more accurate assessment of how a garment or other wearable product would look on him / her. In the described system, streaming media data is augmented using virtual fitting data of the user. To this end, a product model associated with the streaming media data is identified, and a user model associated with a viewer of the streaming media data is obtained. The product model is tried on the user model, with the user model posing in a manner similar to the outline of a presenter in the streaming media data. The product model and the user model are rendered and presented with the streaming media data (e.g., augmented in the streaming media data). BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1An illustrative example of a system that can enhance streaming video using virtual fitting information in accordance with at least some embodiments is shown.

[0014] Figure 2 A system architecture for augmenting streaming data with virtual fitting information is shown, in accordance with at least some embodiments.

[0015] Figure 3 A simplified flow chart of a method for presenting a data stream enhanced with virtual fitting data according to an embodiment of the present invention is shown.

[0016] Figure 4 Illustrative examples of techniques for obtaining a 3D model using sensor data are shown, in accordance with at least some embodiments.

[0017] Figure 5 An example of a graphical user interface (GUI) is shown for demonstrating achievable features according to an embodiment of the present invention.

[0018] Figure 6 A flow chart illustrating a process for presenting virtual fitting data to a user, according to at least some embodiments.

[0019] Figure 7 An example of components of a computer system according to some embodiments is shown.

[0020] Figure 8 A block diagram of an apparatus for presenting virtual fitting data to a user is shown, according to at least some embodiments. DETAILED DESCRIPTION

[0021] The present invention generally relates to methods and systems related to virtual reality applications. More specifically, embodiments of the present invention provide methods and systems for determining the fit between a user and a product. Embodiments of the present invention are suitable for various applications in virtual reality and computer-based fitting systems.

[0022] Figure 1 An illustrative example of a system that can enhance streaming video using virtual fitting information in accordance with at least some embodiments is shown. Figure 1 In the example embodiment, the user device 102 is used to provide a request for virtual fitting information to the mobile application server 104. In some cases, the user device can be used to obtain user data 106, which can be provided to the mobile application server 104 for generating the virtual fitting information.

[0023] In one example, user device 102 represents a suitable computing device that includes one or more graphics processing units (GPUs), one or more general-purpose processors (GPPs), and one or more memories storing computer-readable instructions, wherein the computer-readable instructions can be executed by at least one of the processors to implement various functions of embodiments of the present invention. For example, user device 102 can be any one of a smartphone, a tablet, a laptop, a personal computer, a game console, or a smart TV. User device 102 can also include a range-finding camera (i.e., a depth sensor) and / or an RGB optical sensor (e.g., a camera).

[0024] The user device can be used to capture and / or generate user data 106. The user data 106 may include information related to a specific user (e.g., the user of the user device 102) for whom virtual fitting data should be created. The user data 106 may include data about the user that can be used to generate virtual fitting data. For example, the user data 106 may include the user's dimensions. The user data 106 can be captured in any suitable format. For example, the user data 106 may include a point cloud, a 3D mesh or model, or a string containing measurements at predetermined locations. In some cases, capturing the user data 106 includes receiving information about the user that is manually entered into the user device 102. For example, the user may enter measurements of various parts of the user's body via a keyboard. In some cases, acquiring the user data 106 may include using a camera and / or a depth sensor to acquire image / depth information related to the user. The user device 102 may also be configured to generate a 3D model based on the acquired image / depth information. Figure 1 Explain the process in more detail.

[0025] The mobile application server 104 comprises any computing device capable of generating a data stream for a user enhanced with virtual fitting data according to the techniques described herein. To generate the enhanced data stream, the mobile application server 104 may receive user data 106 from the user device 102. It should be noted that while the mobile application server 104 may receive the user data 106 concurrently with a request to generate virtual fitting data, the mobile application server 104 may also receive the user data 106 prior to and independently of any request to generate virtual fitting data. For example, the mobile application server 104 may receive the user data 106 during a registration phase when a user establishes an account with the mobile application server 104.

[0026] The request for virtual fitting data may refer to stream data 108. Stream data 108 may be streaming video (e.g., a live stream) or other suitable dynamic media content. Stream data 108 may show at least one presenter 110 and at least one product 112. Mobile application server 104 may obtain an identifier (product identifier 114) of at least one product 112 and data related to the presenter's posture (gesture data 116) from stream data 108. In some embodiments, one or more of product identifier 114 or posture data 116 may be associated with stream data 108 via metadata attached to stream data 108. In some embodiments, one or more machine vision techniques may be used to determine one or more of product identifier 114 and / or posture data 116 from images within stream data 108.

[0027] The mobile application server 104 may include or access object model data 118, from which product data 120 may be obtained to fulfill the request. The object model data 118 may include any computer-readable storage medium having one or more 3D models stored thereon. For example, the object model data 118 may be a database maintained by the mobile application server 104 or another server. The 3D models stored in the object model data 118 may represent products that can be worn by a user, such as clothing (e.g., clothes) or accessories. In some embodiments, the object model data 118 may store 3D models of multiple versions (e.g., different sizes and / or styles) of a product. Upon receiving a product identifier 114 for a particular product, the mobile application server 104 retrieves the product data 120 from the object model data 118, which includes the 3D models associated with the particular product.

[0028] The mobile application server 104 can be configured to combine the user data 106 and the product data 120 to generate a try-on avatar for the user. The mobile application server 104 can also pose the try-on avatar based on the pose data 116. Once the try-on avatar is generated, the mobile application server 104 can use the try-on avatar to enhance the stream data 108 to generate enhanced stream data 122. Once the enhanced stream data 122 is generated, the enhanced stream data 122 can be sent back to the user device 102, where it can be rendered on a display for the user to view.

[0029] For clarity, Figure 1 A certain number of components are shown in FIG. However, it should be understood that the number of each component in the embodiments of the present invention can be more than one. In addition, some embodiments of the present invention may include less than or more than Figure 1 All components shown. In addition, Figure 1The components in the can communicate using any suitable communication protocol over any suitable communication medium, including the Internet.

[0030] Figure 2 The architecture of a system for augmenting streaming data with virtual fitting information according to at least some embodiments is shown. Figure 2 In the embodiment, the user device 202 can communicate with a plurality of other components including at least a mobile application server 204. The mobile application server 204 can execute at least a portion of the processing functions required by the mobile application installed on the user device. The user device 202 and the mobile application server 204 can be reference Figure 1 Examples of user device 102 and mobile application server 104 are described separately.

[0031] User device 202 can be any suitable electronic device that has at least some of the functionality described herein. Specifically, user device 202 can be any electronic device that can capture user data and / or present an enhanced data stream on a display. In some embodiments, the user device can establish a communication session with another electronic device (e.g., mobile application server 204) and send / receive data to / from the electronic device. The user device has the ability to download and / or execute mobile applications. User devices include mobile communication devices as well as personal computers and thin client devices. In some embodiments, the user device includes any portable electronic device that has basic communication-related functionality. For example, the user device can be a smartphone, a personal data assistant (PDA), or any other suitable handheld device. The user device can be implemented as a self-contained unit having various components (e.g., input sensors, one or more processors, memory, etc.) integrated into the user device. References to the "output" of a component or the "output" of a sensor in the present invention do not necessarily mean that the output is sent outside the user device. The outputs of the various components can remain within the self-contained unit that defines the user device.

[0032] In one illustrative configuration, the user device 202 may include at least one memory 206 and one or more processing units (or processors) 208. The processor 208 may be suitably implemented as hardware, computer-executable instructions, firmware, or a combination thereof. The computer-executable instructions or firmware implementation of the processor 208 may include computer-executable instructions or machine-executable instructions written in any suitable programming language for performing the various functions described. The user device 202 may also include one or more input sensors 210 for receiving user input and / or environmental input. There may be various input sensors 210 capable of detecting user input or environmental input, such as accelerometers, camera devices, depth sensors, microphones, global positioning system (e.g., GPS) receivers, etc. The one or more input sensors 210 may include a ranging camera device (e.g., a depth sensor) capable of generating a depth image and a camera device for acquiring image information.

[0033] For the purposes of the present invention, a rangefinder camera (e.g., a depth sensor) can be any device used to identify the distance or range of one or more objects from the rangefinder camera. In some embodiments, the rangefinder camera can generate a depth image (or depth map), where the pixel value in the image (or depth map) corresponds to the detected distance of that pixel. Pixel values ​​can be directly obtained in physical units (e.g., meters). In at least some embodiments of the present invention, a user device can employ a rangefinder camera that uses structured light. In a rangefinder camera that uses structured light, a projector projects light in a structured pattern onto one or more objects. The light can be outside the visible light range (e.g., infrared or ultraviolet). The rangefinder camera has one or more camera devices that capture an image of the object with the reflected pattern. Distance information can then be generated based on the distortion in the detected pattern. It should be noted that while the present invention focuses on the use of rangefinder cameras using structured light, any suitable type of rangefinder camera, including those that use stereo triangulation, light sheet triangulation, time-of-flight, interferometry, coded aperture, or any other suitable distance detection technique, can be used with the described system.

[0034] Memory 206 stores program instructions that can be loaded and executed on processor 208, as well as data generated during the execution of these programs. Depending on the configuration and type of user device 202, memory 206 can be volatile (e.g., random access memory (RAM)) and / or non-volatile (e.g., read-only memory (ROM), flash memory, etc.). User device 202 also includes additional storage 212, such as removable or non-removable storage, including but not limited to magnetic storage, optical disks, and / or tape storage. Disk drives and their associated computer-readable media can provide non-volatile storage for computer-readable instructions, data structures, program modules, and other data for the computing device. In some embodiments, memory 206 can include multiple different types of memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), or ROM. The contents of memory 206 are described in detail below. Memory 206 includes an operating system 214 and one or more applications or services used to implement the features disclosed herein, including at least mobile application 216. Memory 206 also includes application data 218, which provides information generated and / or used by mobile application 216. In some embodiments, application data 218 may be stored in a database.

[0035] For purposes of the present invention, a mobile application may be any set of computer-executable instructions that are installed and executed on a user device 202. The mobile application may be installed on the user device by the user device manufacturer or by another entity. In some embodiments, the mobile application 216 may enable the user device to establish a session with a mobile application server 204, which provides backend support for the mobile application 216. The mobile application server 204 may maintain account information associated with a particular user device and / or user. In some embodiments, the user may be required to log in to the mobile application to access functionality provided by the mobile application 216.

[0036] According to at least some embodiments, the mobile application 216 is configured to provide user information to the mobile application server 204 and to present to the user information received from the mobile application server 204. More specifically, the mobile application 216 is configured to obtain measurement data of the user and submit the measurement data to the mobile application server 204 in association with a request for streamed data enhanced with virtual fitting data. In some embodiments, the mobile application 216 may also receive an indication of a data stream enhanced with virtual fitting data.

[0037] According to at least some embodiments, the mobile application 216 can receive output from the input sensor 210 and generate a 3D model based on the output. For example, the mobile application 216 can receive depth information (e.g., a depth image) from a depth sensor (e.g., a ranging camera), wherein the depth sensor can be, for example, the depth sensor mentioned in the previous description of the input sensor 210, and the mobile application 216 can also receive image information from the camera input sensor. Based on this information, the mobile application 216 can determine the edges of the object to be identified (e.g., a user). For example, a sudden change in depth within the depth information can indicate the edge or outline of the object. In another example, the mobile application 216 can use one or more machine vision techniques and / or machine learning to identify the edges of the object. In this example, the mobile application 216 can receive image information from the camera input sensor 210 and can identify potential objects within the image information based on differences in color or texture data detected within the image or based on learned patterns. In some embodiments, the mobile application 216 may cause the user device 202 to send output obtained from the input sensor 210 to the mobile application server 204 , which may then perform one or more object recognition techniques on the output to generate a 3D model of the object.

[0038] User device 202 also includes a communication interface 220 that enables user device 202 to communicate with any other suitable electronic device. In some embodiments, communication interface 220 may enable user device 202 to communicate with other electronic devices on a network (e.g., a private network). For example, user device 202 may include a Bluetooth® interface that allows it to communicate with another electronic device. TM (BLUETOOTH TM ) wireless communication module. The user device 202 also includes input / output (I / O) devices and / or ports 222, for example, for connecting to a keyboard, mouse, pen, voice input device, touch input device, display, speaker, printer, etc.

[0039] In some embodiments, user device 202 can communicate with mobile application server 204 via a communication network. The communication network can include any one or a combination of multiple different types of networks, such as cable networks, the Internet, wireless networks, cellular networks, and other private and / or public networks. In addition, the communication network includes a variety of different networks. For example, user device 202 can communicate with a wireless router via a wireless local area network (WLAN), and the wireless router can then route the communication to mobile application server 204 via a public network (e.g., the Internet).

[0040] The mobile application server 204 can be any computing device or multiple computing devices used to perform one or more calculations for the mobile application 216 on the user device 202. In some embodiments, the mobile application 216 can communicate with the mobile application server 204 periodically. For example, the mobile application 216 can receive updates, push notifications, or other instructions from the mobile application server 204. In some embodiments, the mobile application 216 and the mobile application server 204 can use proprietary encryption and / or decryption schemes to protect communications between the two. In some embodiments, the mobile application server 204 can be executed by one or more virtual machines implemented in a hosted computing environment. The hosted computing environment includes one or more computing resources that are quickly provisioned and released, and the computing resources may include computing, networking, and / or storage devices. A hosted computing environment may also be referred to as a cloud computing environment.

[0041] In one illustrative configuration, the mobile application server 204 may include at least one memory 224 and one or more processing units (or processors) 226. The processor 226 may be suitably implemented as hardware, computer-executable instructions, firmware, or a combination thereof. The computer-executable instructions or firmware implementation of the processor 226 may include computer-executable instructions or machine-executable instructions written in any suitable programming language for performing the various functions described.

[0042] Memory 224 can store program instructions that are loadable and executable on processor 226, as well as data generated during the execution of these programs. Depending on the configuration and type of mobile application server 204, memory 224 can be volatile (e.g., RAM) and / or non-volatile (e.g., ROM, flash memory, etc.). Mobile application server 204 also includes additional memory 228, such as removable or non-removable memory, including but not limited to magnetic storage, optical disks, and / or tape storage. Disk drives and their associated computer-readable media can provide non-volatile storage for computer-readable instructions, data structures, program modules, and other data for computing devices. In some embodiments, memory 224 can include multiple different types of memory, such as SRAM, DRAM, or ROM. The contents of memory 224 are described in detail below. Memory 224 includes an operating system 230 and one or more applications or services for implementing the features disclosed herein, including at least a module for fitting a product 3D model onto a user 3D model (fitting module 232) and / or a module for determining and applying gestures to the product 3D model and the user 3D model (gesture module 234). Memory 224 also includes account data 236, which provides information associated with user accounts maintained by the system; user model data 238, which maintains 3D models associated with each user of the account; and / or object model data 240, which maintains 3D models associated with a plurality of objects (products). In some embodiments, one or more of account data 236, user model data 238, or object model data 240 may be stored in a database. In some embodiments, object model data 240 may be an electronic catalog that includes data related to objects available for sale from a resource provider (e.g., a retailer or other suitable merchant).

[0043] Memory 224 and additional memory 228 are examples of computer-readable storage media, which may be removable or non-removable. For example, computer-readable storage media may include volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. As used herein, the term "module" refers to a programming module executed by a computing system (e.g., a processor) installed on and / or executed from the mobile application server 204. The mobile application server 204 also includes a communication connection 242, which allows the mobile application server 204 to communicate with a stored database, another computing device or server, a user terminal and / or other components of the system. The mobile application server 204 also includes input / output (I / O) devices and / or ports 244, for example, for implementing connections with a keyboard, mouse, pen, voice input device, touch input device, display, speaker, printer, etc.

[0044] The contents of the memory 224 are described in detail below. The memory 224 includes a try-on module 232, a gesture module 234, a database containing account data 236, a database containing user model data 238, and / or a database containing object model data 240.

[0045] In some embodiments, the try-on module 232 may be configured to, in conjunction with the processor 226, deform the product 3D model in order to fit it onto the user 3D model. The try-on module 232 may access one or more rules that describe how to deform (e.g., stretch and / or bend) a specific product type (e.g., a shirt, pants, etc.) in order to fit it onto the user model. To fit the product 3D model onto the user 3D model, the try-on module 232 may align certain portions of the product 3D model to specific portions of the user 3D model. For example, the 3D model of the shirt may be positioned so that the sleeves of the 3D model of the shirt wrap around the arms of the user 3D model. Furthermore, the 3D model of the shirt may be positioned so that the collar of the 3D model of the shirt wraps around the neck of the user 3D model. The remaining portions of the 3D model of the shirt may then be deformed by stretching and bending portions of the 3D model of the shirt so that the interior of the 3D model of the shirt is outside or along the outside of the 3D model of the user.

[0046] In some embodiments, the posture module 234 may be configured to, together with the processor 226, identify the posture of the presenter (i.e., the human body) in the streaming data and apply the posture to the combination of the product 3D model and the user 3D model generated by the try-on module 232. This includes using one or more posture estimation techniques to determine the current posture of the presenter in the data stream. For example, the posture module 234 may use machine learning to determine the posture of the presenter in the data stream. Those skilled in the art will appreciate that a variety of suitable posture estimation techniques may be employed. In some embodiments, the posture module 234 may apply the determined posture to the user model (on which the product model has been tried on). This includes repositioning one or more accessories or body parts of the user model until the determined posture is achieved. In some embodiments, the posture module 234 may monitor the posture of the presenter in the streaming data and may adjust the posture of the user model to match the posture of the presenter when a change in the posture of the presenter is detected.

[0047] After adjusting the posture of the combined user model and product model, the posture module 234 can render the combined user model and product model. In some embodiments, the combined user model and product model can be rendered in a small window, which is then placed in an inconspicuous position within the streaming data. For example, in the case where the streaming data is a video, the combined user model and product model can be rendered in a window in the lower corner of the video. This rendering can enable the user to imagine the product being worn on him / herself. It should be noted that although the posture module 234 and the try-on module 232 are described with reference to the mobile application server 204, the functions described as being performed by one or more modules can be performed by a mobile application on the user device 202.

[0048] In some embodiments, each object entry within object model database 240 can be associated with a 3D model of the object. In these embodiments, the 3D model can be combined with a second 3D model of the user and provided to mobile application 216, so that user device 202 displays the combined 3D model on the user device's display as an enhancement to the streamed data. As the presenter's pose within the streamed data is updated, mobile application 216 can dynamically update the pose of the combined 3D model on the user device's display.

[0049] Figure 3A simplified flowchart of a method for presenting a data stream enhanced with virtual fitting data provided by an embodiment of the present invention is shown. The process is introduced in conjunction with a computer system as an example of a computer system introduced in the present invention. Some or all of the operations of the process can be implemented by specific hardware on the computer system and / or as computer-readable instructions stored on a non-transitory computer-readable medium of the computer system. As stored, the computer-readable instructions represent a programmable module including code that can be executed by a computer system processor. The execution of such instructions configures the computer system to perform the corresponding operation. Each programmable module and the processor together represent a device for performing the corresponding operation. Although these operations are shown in a specific order, it should be understood that the specific order is not required, and one or more operations can be omitted, skipped and / or reordered.

[0050] At the beginning of process 300, one or more products 302 are scanned to generate product model data 304. Generated product models 304 serve as 3D virtual representations of products 302. Product models 304 may be generated by scanning products 302 from multiple perspectives using a camera and / or a depth sensor. In step 1 of process 300, the multiple generated product models 304 are provided to mobile application server 204 for storage in object model data 240. Product models 304 may be generated by a variety of different entities. For example, a product model for a particular product may be generated by the manufacturer of that product.

[0051] In addition, in step 2 of the process 300, the user is scanned using a camera and / or a depth sensor installed on the user device 202 to generate user model data 308. Figure 4 Some example techniques for generating an object (e.g., user) model are explained in more detail. In step 3 of process 300, user model data 308 is sent to the mobile application server 204 for storage in the user model data 238. In some cases, the user model data 308 can be stored in association with an account maintained for the scanned user.

[0052] The mobile application server 204 receives a request from a user to use the streaming data 306. In step 4 of process 300, upon receiving the request to use the streaming data 306, the mobile application server 204 retrieves the streaming data 306 from its location. In some embodiments, the streaming data 306 may be maintained by the mobile application server 204. In some embodiments, the streaming data 306 may be maintained by an entity separate from the mobile application server 204. For example, a user may request through the mobile application server 204 to view a video hosted by YOUTUBE.com. TMThe mobile application server 204 provides support for a mobile application installed on a user's mobile device. In this example, the user can provide a uniform resource locator (URL) or other identifier of a video. The mobile application server 204 can then retrieve the video file by accessing the URL. Once the stream data 306 is retrieved, the mobile application server 204 identifies one or more related products and the pose of the presenter. As described elsewhere, the above-mentioned Figure 2 The described try-on module 232 and / or gesture module 234 are used to complete the process.

[0053] The steps performed by the gesture module 234 are described below. In step 5 of process 300, process 300 includes determining the gesture of a presenter within the stream data 306. This involves first identifying the presenter within the stream data (e.g., using one or more machine vision techniques), and then estimating the gesture of the presenter using any suitable gesture estimation technique. Those skilled in the art will appreciate that a variety of suitable techniques may be employed. Generally speaking, the gesture of an object indicates the position and orientation of the object. For a speaker, the estimated gesture includes a record of the position and orientation of various body parts or joints of the presenter.

[0054] After estimating the presenter's pose, the pose module 234 applies the pose to the user model. To do this, in step 6 of process 300, the pose module 234 retrieves the user model. A user model is a three-dimensional model representing a person that is stored in user model data 238 in association with the person or an account linked to the person. Upon receiving a request for usage stream data 306 regarding a particular user, the pose module 234 can retrieve the user model associated with the person from the user model data 238. Once retrieved, the pose module 234 applies the presenter's estimated pose to the retrieved user model. To do this, the pose module 234 repositions the various body parts of the user model to match the record of the position and orientation of the various body parts of the presenter. The posed user model is then provided to the try-on module 232 in step 7 of process 300.

[0055] The following describes the steps performed by try-on module 232. In step 8 of process 300, process 300 includes retrieving one or more product models. First, try-on module 232 identifies one or more products associated with stream data 306. In some embodiments, stream data 306 may include an indication of one or more products. For example, stream data 306 may include additional metadata indicating a stock keeping unit (SKU) or other product identifier associated with stream data 306. In some embodiments, one or more products may be identified from stream data 306 using machine vision techniques (e.g., object recognition). For example, try-on module 232 may identify a specific product (e.g., a shirt or pants) worn by a presenter within stream data 306 by comparing the product's visual attributes with attributes associated with multiple products stored in an electronic catalog maintained by mobile application server 204. In this example, try-on module 232 may identify the product in the electronic catalog that best matches the product worn by the presenter. Once try-on module 232 has identified one or more products associated with stream data 306, it retrieves product models for these products from the object model data.

[0056] In step 9 of process 300, one or more product models obtained in step 8 are fitted onto the posed user model provided in step 7. The product models are fitted onto the posed user model by adjusting a set of parameters that control the deformation of one or more regions of interest of the product model until the product model matches the user model. This set of parameters can be defined as a set of measurements, such as the displacement of each vertex of the product model. This process can be described as an optimization process, in which several different optimization algorithms can be used to find the optimal set of parameters that minimizes one or more cost functions. The cost function can be defined as the number of penetrations between the meshes of the two 3D models, the average distance between the vertices of the body mesh and the vertices of the clothing mesh, etc. Further examples of techniques for rendering product models fitted onto a user model are described in more detail in U.S. patent application Ser. No. 62 / 987,196, entitled “SYSTEM AND METHOD FOR VIRTUAL FITTING,” which is hereby incorporated by reference in its entirety for all purposes.

[0057] Once the product models have been tried on the user model, they are rendered at step 10. Rendering is the process of giving a 3D model a physical appearance using shading and color. Those skilled in the art will recognize that there are a variety of suitable techniques for rendering the product models tried on the posed user model. In some embodiments, the rendered models are used to enhance the stream data 306. For example, the rendered models can be placed as augmented visual data in a small window within the stream data 306, allowing a viewer of the stream data 306 (e.g., a user) to view the rendered models while viewing the stream data 306.

[0058] Then, at step 11 of process 300, the rendered model (e.g., enhanced streaming data) is provided to user device 202. Upon receiving the rendered model, user device 202 may present the rendered model to the user. For example, user device 202 may play the enhanced streaming data via a media player application.

[0059] In some embodiments, additional processing may be performed when using the enhanced streaming data 306. For example, the user device 202 that is presenting the enhanced streaming data 306 may capture image information of the user viewing the enhanced streaming data 306 using a front-facing camera installed on the user device 202. In this example, the user's facial data may be extracted from the image information and overlaid onto the rendered user model, thereby providing the user's facial and facial expression data to the user model.

[0060] It should be understood that Figure 3 The specific steps shown provide a specific method for presenting a data stream enhanced with virtual fitting data according to an embodiment of the present invention. As described above, according to alternative embodiments, the steps may also be performed in other orders. For example, alternative embodiments of the present invention may perform the above steps in a different order. In addition, Figure 3 The various steps shown may include multiple sub-steps that may be performed in a variety of sequences suitable for the various steps. Additionally, steps may be added or deleted based on specific applications. Many variations, modifications, and alternatives are available to those skilled in the art.

[0061] Figure 4 An illustrative example of a technique for obtaining a 3D model using sensor data, according to at least some embodiments, is shown. According to at least some embodiments, sensor data 402 may be obtained from one or more input sensors mounted on a user device. The captured sensor data 402 includes image information 404 captured by a camera device and depth map information 406 captured by a depth sensor.

[0062] As described above, sensor data 402 includes image information 404. One or more image processing techniques can be applied to image information 404 to identify one or more objects within image information 404. For example, edge detection can be used to identify regions 408 within image information 404 that include objects. To this end, discontinuities in brightness, color, and / or texture can be identified within the image to detect the edges of various objects within the image. Region 408 shows an illustrative example image of a chair that highlights such discontinuities.

[0063] As described above, sensor data 402 includes depth information 406. In depth information 406, each pixel can be assigned a value representing the distance between the user device and a specific point corresponding to the pixel's location. Depth information 406 can be analyzed to detect sudden changes in depth within depth information 406. For example, a sudden change in distance can indicate an edge or boundary of an object within depth information 406.

[0064] In some embodiments, sensor data 402 includes both image information 404 and depth information 406. In at least some of these embodiments, an object can be first identified in either the image information 404 or the depth information 406, and various attributes of the object can be determined from the other information. For example, edge detection techniques can be used to identify a region 408 within the image information 404 that includes an object. Region 408 can then be mapped to a corresponding region 410 in the depth information to determine depth information (e.g., a point cloud) for the identified object. In another example, region 410 can first be identified within the depth information 406 that includes an object. In this example, region 410 can then be mapped to a corresponding region 408 in the image information to determine appearance attributes (e.g., color or texture values) of the identified object.

[0065] In some embodiments, various attributes of objects identified in the sensor data 402 (e.g., color, texture, point cloud data, object edges) can be used as input to a machine learning module to identify or generate a 3D model 412 that matches the identified object. In some embodiments, a point cloud of the object can be generated from the depth information and / or image information and compared to point cloud data stored in a database to identify the best matching 3D model. Alternatively, a 3D model of an object (e.g., a user or a product) can be generated using the sensor data 402. To do this, a mesh can be created using point cloud data obtained from region 410 of the depth information 406. The system can then map appearance data from the region corresponding to region 410 in the image information 404 to the mesh to generate a basic 3D model. Although specific techniques are described, it should be noted that there are many techniques for identifying specific objects from sensor output.

[0066] As described elsewhere, by a user device (e.g. Figure 1 Sensor data captured by a user device 102 (e.g., a user's device) can be used to generate a user 3D model using the techniques described above. This user 3D model can then be provided to the mobile application server as user data. In some embodiments, the sensor data can be used to generate a 3D model of a product, which can be stored in user model data 238. For example, a user wishing to sell a product can capture sensor data related to the product from a user device. The user's user device can then generate a 3D model in the manner described above and provide the 3D model to the mobile application server.

[0067] Figure 5 An example of a graphical user interface (GUI) is shown for demonstrating some example features that may be implemented according to an embodiment of the present invention. Figure 5 In FIG, an example user device 502 is shown having a display screen on which visual data can be presented. The user device 502 is the same as that described above with reference to FIG. Figure 2 An example of user equipment 202 is described.

[0068] like Figure 5 As shown, a GUI of a software application (e.g., a media viewer application) installed on a user device 502 can be used to present streaming data 504. Streaming data 504 includes at least one presenter 506 and a product 508, where the presenter is the person shown in the streaming data 506. Product 508 can be worn or otherwise presented by presenter 506 within the streaming data 504.

[0069] As described elsewhere, a posed and tried-on model 510 can be presented with the streaming data 504. For example, the model 510 can be presented within a separate window 512 positioned to minimize any obstruction to viewing the streaming data 504, sometimes referred to as a picture-in-picture. The model 510 includes a user model representing the current viewer of the streaming data 504, the user model having posed in a manner similar to the presenter 508 and having tried on a product model representing the product 508.

[0070] Figure 6 A flow chart illustrating a process for presenting virtual fitting data to a user, as provided by at least some embodiments. Figure 6 The process 600 shown in FIG. 6 may be performed by a user device (eg, Figure 2 The user device 202) communicates with the mobile application server (eg, Figure 2 The mobile application server 204) executes.

[0071] At 602, process 600 includes receiving an indication of media content being used by a user. For example, an indication may be received that the user is watching a streaming video, where streaming video is a type of media content. The indicated media content may include a display of a presenter, where the presenter is a person other than the user. The indicated media content may also include a description of a product presented by the presenter. For example, the product may be clothing worn by the presenter in the media content.

[0072] At 604, process 600 includes identifying a product associated with the media content. In some embodiments, the product associated with the media content is identified by an identifier associated with the product contained in metadata of the media content. In one example, the identifier associated with the product is a SKU number. In some embodiments, the product associated with the media content is identified by object recognition.

[0073] At 606, process 600 includes obtaining a first 3D model representing the product. To this end, the first 3D model is obtained from a computer having object model data stored therein (e.g., Figure 2 The 3D model associated with the product identified in 604 may be retrieved from a database of object model data 240. In some embodiments, an appropriate size and / or style of the product may be selected based on stored information about the user who is viewing the media content.

[0074] At 608, process 600 includes obtaining a second 3D model representing the user. In some embodiments, the user model may be stored in association with one or more accounts. In these embodiments, the second 3D model may be identified and retrieved by being stored in association with the account used to view the media content. In some embodiments, the second 3D model representing the user may be received from a user device currently being used to view the media content.

[0075] At 610, process 600 includes determining a presentation pose based on the media content. The presentation pose is determined to be the presenter's current pose within the media content. This can be accomplished using any suitable pose estimation technique. The determined presentation pose includes an indication of various parts (e.g., body parts) of the user model and their respective positions and orientations.

[0076] At 612, process 600 includes applying the demonstration gesture to the second 3D model. To this end, the positions and orientations of various parts (e.g., body parts) of the second 3D model may be adjusted so that they match corresponding positions and orientations in the demonstration gesture data.

[0077] At 614 , process 600 includes generating a third 3D model by fitting the second 3D model through the first 3D model. Here, the first 3D model is deformed to minimize a distance between the first 3D model and the second 3D model.

[0078] At 616, process 600 includes presenting the third 3D model to the user. This involves rendering the third 3D model and providing the third 3D model to a user device that is presenting the media content. The third 3D model is presented along with the media content. For example, the media content may be enhanced to include the third 3D model (e.g., in a separate window within the media content).

[0079] It should be understood that according to the embodiments of the present invention, Figure 6 The specific steps shown provide a specific method for presenting virtual fitting data to the user. As mentioned above, other step sequences may also be performed according to alternative embodiments. For example, alternative embodiments of the present invention may perform the above steps in a different order. In addition, Figure 6 The various steps shown may include multiple sub-steps that may be performed in a variety of sequences suitable for the various steps. Additionally, steps may be added or deleted based on specific applications. Many variations, modifications, and alternatives are available to those skilled in the art.

[0080] Figure 7 1. An example of computer system components provided by some embodiments is shown. Computer system 700 is an example of a computer system described in this disclosure. Although these components are shown as belonging to the same computer system 700, the computer system 700 can also be distributed.

[0081] The computing system 700 includes at least a processor 702, a memory 704, a storage device 706, an input / output (I / O) peripheral device 708, a communication peripheral device 710, and an interface bus 712. The interface bus 712 can be used to communicate, send, and transmit data, control, and commands between the various components of the computing system 700. The memory 704 and the storage device 706 can include computer-readable storage media, such as RAM, ROM, electrically erasable programmable read-only memory (EEPROM), hard drives, CD-ROMs, optical storage devices, magnetic storage devices, electronic non-volatile computer storage (e.g., memory), and other tangible storage media. Any such computer-readable storage media can be used to store instructions or program code that implement various aspects of the present disclosure. The memory 704 and the storage device 706 can also include computer-readable signal media. Computer-readable signal media include propagated data signals containing computer-readable program code. Such propagated signals can take any of a variety of forms, including but not limited to electromagnetic, optical, or any combination thereof. Computer-readable signal media includes any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transport a program for use in connection with computing system 700 .

[0082] In addition, the memory 704 package may include an operating system, programs, and applications. The processor 702 may be used to execute stored instructions and include, for example, a logic processing unit, a microprocessor, a digital signal processor, and other processors. The memory 704 and / or the processor 702 may be virtualized and may be hosted in, for example, another computing system in a cloud network or a data center. The input / output peripherals 708 may include a user interface, such as a keyboard, a screen (e.g., a touch screen), a microphone, a speaker, other input / output devices, and computing components, such as a graphics processing unit, a serial port, a parallel port, a universal serial bus, and other input / output peripherals. The input / output peripherals 708 are connected to the processor 702 via any port coupled to the interface bus 712. The communication peripherals 710 may be used to facilitate communication between the computing system 700 and other computing devices through a communication network, and include, for example, a network interface controller, a modem, wireless and wired interface cards, antennas, and other communication peripherals.

[0083] Figure 8 A block diagram of an apparatus for presenting virtual fitting data to a user, as provided by at least some embodiments, is shown. Figure 8 The apparatus 800 shown in FIG. 8 may be implemented as a user device (eg, Figure 2The user device 202) communicates with the mobile application server (eg, Figure 2 Mobile application server 204).

[0084] The apparatus 800 includes a receiving module 802 configured to receive an indication of media content being viewed by a user. For example, an indication may be received that the user is viewing a streaming video, where streaming video is a type of media content. The indicated media content may include a description of a presenter, where the presenter is a person other than the user. The indicated media content may also include a description of a product presented by the presenter. For example, the product may be clothing worn by the presenter in the media content.

[0085] The device 800 also includes an identification module 804 that is configured to identify a product associated with the media content. In some embodiments, the product associated with the media content is identified by an identifier that is associated with the product contained in metadata of the media content. In one example, the identifier associated with the product is a stock keeping unit (SKU) number. In some embodiments, the product associated with the media content is identified by object recognition.

[0086] The apparatus 800 further comprises an acquisition module 806 configured to acquire a first three-dimensional (3D) model representing the product. Figure 2 The 3D model associated with the product identified in 604 may be retrieved from a database of object model data 240. In some embodiments, an appropriate size and / or style of the product may be selected based on stored information about the user who is viewing the media content.

[0087] The acquisition module 806 is further configured to acquire a second 3D model representing the user. In some embodiments, the user model may be stored in association with one or more accounts. In these embodiments, the second 3D model may be identified and retrieved by being stored in association with the account used to view the media content. In some embodiments, the second 3D model representing the user may be received from a user device currently being used to view the media content.

[0088] The apparatus 800 further includes a determination module 808 configured to determine a presentation gesture based on the media content. The presentation gesture is determined to be the presenter's current gesture within the media content. Any suitable gesture estimation technique may be used to accomplish this. The determined presentation gesture may include an indication of various parts (e.g., body parts) of the user model and their respective positions and orientations.

[0089] The device 800 also includes an application module 810, which is configured to apply the demonstration posture to the second 3D model; to this end, the position and orientation of various parts (e.g., body parts) of the second 3D model can be adjusted to match the corresponding position and orientation in the demonstration posture data.

[0090] The apparatus 800 further includes a generating module 812 configured to generate a third 3D model by fitting the second 3D model onto the first 3D model, wherein the first 3D model is deformed to minimize the distance between the first 3D model and the second 3D model.

[0091] The apparatus 800 further includes a presentation module 814 configured to present the third 3D model to the user. This involves rendering the third 3D model and providing the third 3D model to a user device that is presenting the media content. The third 3D model is presented along with the media content. For example, the media content may be enhanced to include the third 3D model (e.g., in a separate window within the media content).

[0092] Although the subject matter has been described in detail with respect to specific embodiments thereof, it will be understood that those skilled in the art, after obtaining an understanding of the foregoing, can easily generate changes, variations and equivalents to these embodiments. Therefore, it will be understood that the present disclosure is presented for purposes of illustration and not limitation, and does not exclude the inclusion of such modifications, variations and / or additions to the subject matter that are obvious to a person of ordinary skill in the art. In fact, the methods and systems described in the present disclosure can be implemented in a variety of other forms; in addition, various omissions, substitutions and changes in the form of the methods and systems described in the present disclosure can be made without departing from the spirit of the present disclosure. The accompanying claims and their equivalents are intended to cover such forms or modifications that fall within the scope and spirit of the present disclosure.

[0093] Unless expressly stated otherwise, it should be understood that terms such as "process," "compute," "calculate," "determine," and "identify" are used throughout the discussion of this specification to refer to the actions or processes of a computing device (e.g., one or more computers or similar electronic computing devices) that manipulate or transform data represented as physical electronic or magnetic quantities in a memory, register, or other information storage device, transmission device, or display device of a computing platform.

[0094] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provides a result conditioned on one or more inputs. Suitable computing devices include multi-purpose microprocessor-based computer systems that access stored software that programs or configures the computing system from a general-purpose computing device to a special-purpose computing device that implements one or more embodiments of the present subject matter. Any suitable programming, scripting, or other type of language or combination of languages ​​may be used to implement the teachings contained herein in software used to program or configure a computing device.

[0095] Embodiments of the methods disclosed herein can be performed in the operation of such a computing device. The order of the blocks presented in the above examples can be changed, for example, the blocks can be reordered, combined and / or decomposed into sub-blocks. Some blocks or processes can be executed in parallel.

[0096] Conditional language used herein, such as "can," "might," "may," "could," "for example," and the like, unless expressly stated otherwise or otherwise understood in the context of use, is generally intended to convey that some examples include and other examples do not include certain features, elements, and / or steps. Thus, such conditional language generally does not imply that one or more examples require the features, elements, and / or steps in any way, or that one or more examples must include logic for determining, with or without author input or prompting, whether such features, elements, and / or steps are included or will be performed in any particular example.

[0097] The terms "comprise", "include", "have" and the like are synonymous and are used inclusively in an open manner and do not exclude other elements, features, actions, operations and the like. In addition, the term "or" is used in its inclusive (rather than exclusive) manner so that when, for example, it is used to connect a list of elements, the term "or" represents one, some or all of the elements in the list. "Applicable to" or "for" as used herein refers to open and inclusive language and does not exclude devices that are applicable to or used to perform additional tasks or steps. In addition, the use of "based on" means open and inclusive because the process, step, calculation or other action that is "based on" one or more listed conditions or values ​​may actually be based on additional conditions or values ​​other than those listed. Similarly, the use of "based at least in part on" means open and inclusive because the process, step, calculation or other action that is "based at least in part on" one or more listed conditions or values ​​may actually be based on additional conditions or values ​​other than those listed. The titles, lists and numbers included herein are for ease of explanation only and are not meant to be limiting.

[0098] The various features and processes described above can be used independently of each other, or can be used in combination in various ways. All possible combinations and sub-combinations are intended to fall within the scope of this disclosure. In addition, certain method or process blocks can be omitted in some embodiments. The methods and processes described herein are not limited to any particular order, and the blocks or states associated therewith can be performed in other appropriate orders. For example, the blocks or states described can be performed in an order different from that specifically disclosed, or multiple blocks or states can be combined in a single block or state. The example blocks or states can be executed serially, in parallel, or in some other manner. Blocks or states can be added to or deleted from the disclosed examples. Similarly, the example systems and components described herein can be configured differently from what is described. For example, elements can be added, removed, or rearranged compared to the disclosed examples.

Claims

1. A method for virtual fitting, comprising: Obtaining products associated with the media content being viewed by the user; Acquire a first three-dimensional (3D) model representing the product; Acquire a second 3D model representing the user; The third 3D model is generated by fitting the second 3D model onto the first 3D model.

2. The method according to claim 1, wherein: The method further comprises: determining a presentation gesture according to the media content; applying the demonstration pose to the second 3D model; A third 3D model is presented to the user in the presentation posture.

3. The method according to claim 1, wherein: The obtaining of the product associated with the media content being watched by the user comprises: receiving an indication of media content being viewed by a user; A product associated with the media content is identified.

4. The method according to claim 1, wherein: The product associated with the media content is identified by an identifier associated with the product contained in metadata of the media content.

5. The method according to claim 2, wherein: The product is a garment worn by a presenter within the media content, and the presentation pose includes a pose of the presenter.

6. The method according to claim 5, wherein: The presenters include a second user different from the user.

7. The method according to claim 1, wherein: The step of generating a third 3D model by making the second 3D model try on the first 3D model comprises: The first 3D model is fitted onto the second 3D model by adjusting parameters that control deformation of one or more regions of interest of the first 3D model until the first 3D model matches the second 3D model to generate the third 3D model.

8. The method according to claim 2, wherein: The applying the demonstration gesture to the second 3D model comprises: The second 3D model is retrieved from the user model data, the body parts of the second 3D model are repositioned to match the positions and orientations of the body parts of the presenter, and the presentation gesture is applied to the second 3D model.

9. The method according to claim 1, wherein: The method further comprises: Collecting image information of the user; extracting facial data of the user from the image information of the user; The facial data of the user is overlaid on the second 3D model of the user.

10. A system for virtual fitting, comprising: processor; as well as A memory including instructions that, when executed by the processor, cause the system to at least: Obtaining products associated with the media content being viewed by the user; Acquire a first three-dimensional (3D) model representing the product; Acquire a second 3D model representing the user; The third 3D model is generated by fitting the second 3D model onto the first 3D model.

11. A non-transitory computer-readable medium storing specific computer-executable instructions that, when executed by a processor, cause a computer system to at least: Obtaining products associated with the media content being viewed by the user; Acquire a first three-dimensional (3D) model representing the product; Acquire a second 3D model representing the user; The third 3D model is generated by fitting the second 3D model onto the first 3D model.

12. A device for virtual fitting, comprising: An acquisition module, configured to acquire a product associated with the media content being viewed by the user, a first three-dimensional 3D model representing the product, and a second 3D model representing the user; The generating module is configured to generate a third 3D model by making the second 3D model try on the first 3D model.

13. A computer program, wherein: When the computer program is executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method and system for realizing three-dimensional fitting based on intelligent three-dimensional television

    CN103533449A

  • Virtual fitting system in shopping mall

    CN105843386A

  • Real-time shopping method and system, intelligent network television and storage medium

    CN109963201A

  • Gesture interaction method and device based on AR scene, storage medium and communication terminal

    CN110221690A

  • Calibration of model data based on user feedback

    WO2019013736A1