Systems and methods for extracting whole body measurement results
Through deep learning technology and computer vision, the limitations of user posture, clothing and background in the prior art are solved, and the method of accurately extracting body measurement results from 2D images is realized, improving the user experience and reliability of measurement results.
Patent Information
- Application Number
- CN202210093559.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-11-19
- Filing Date
- 2019-04-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2039-04-15
AI Technical Summary
When extracting body measurement results from 2D images, the prior art requires users to adopt specific postures, wear tights, and have to be blank in the background, resulting in poor user experience and inaccurate measurement results.
Deep learning technology combined with computer vision, by identifying and labeling body features, the full-body measurement results are extracted from 2D user images, allowing users to take any posture, wear any type of clothing, and take pictures in front of any background.
It realizes accurate extraction of body measurement results without being restricted by user posture, clothing and background, improving the reliability of user experience and measurement results.
Smart Images

Figure CN114419677B_ABST
Abstract
Description
[0001] This application is a divisional application of a Chinese national stage patent application with the application number 201980000731.7 after the PCT application with the international application number PCT / US2019 / 027564, the international filing date of April 15, 2019, and the invention title of "Systems and methods for fullbody measurements extraction" entered the Chinese national stage on May 27, 2019.
[0002] Citation of priority applications
[0003] This application claims priority under the Patent Cooperation Treaty (PCT) to U.S. Patent Application No. 16 / 195,802, filed on November 19, 2018, with the title "Systems and methods for fullbody measurements extraction" and U.S. Patent Application No. 62 / 660,377, filed on April 20, 2018, with the title "Systems and methods for full body measurements extraction using a 2D phone camera". Technical field
[0004] Embodiments of the present invention pertain to the field of automated body measurements and, more particularly, to using images captured by a mobile device to extract body measurements of a user. Background of the invention
[0005] The statements in the background of the invention are provided to assist in understanding the present invention and its applications and uses and may not constitute prior art.
[0006] Generally, there are three methods that have been attempted to generate or extract body measurements from an image of a user. The first method is to use a 3D camera that provides depth data, such as the MICROSOFT KINECT camera. Through depth sensing, a 3D body model can be established to capture body dimensions. However, not everyone has access to a 3D camera, and since there is no significant widespread adoption at present, it is not currently envisioned that such 3D cameras will become ubiquitous.
[0007] The second method is to use a 2D camera to capture 2D video and use 2D-to-3D reconstruction techniques to recreate a 3D body model to capture body dimensions. Companies such as MTAILOR and 3DLOOK use such techniques. In the 2D video method, a 3D body model is recreated, and the method attempts to perform a "point cloud matching technique" to match an existing 3D body template with the point cloud pre-populated onto the newly created 3D body. However, when attempting to fit an existing template to the 3D body of a unique user, the results may not be accurate. After the matching of the template 3D body to the user's 3D body is completed, the dimensions and measurements are obtained, but the dimensions and measurements are generally inaccurate.
[0008] The third method is to use a 2D camera to capture 2D photos instead of 2D video and, similar to the previous method, utilize 2D-to-3D reconstruction techniques to capture body dimensions. AGISOFT, for example, is a company that has developed 3D reconstruction from 2D photos to 3D models and uses such techniques. Using 2D photos instead of 2D video may involve photos captured at a higher resolution, thus producing slightly more accurate results, but the other aforementioned problems still exist.
[0009] In existing methods using 2D video or photos, 3D body models are produced, and these methods typically require the user to pose in a specific position, stand at a specific distance from the camera, in front of a blank background, wear tight clothing and / or be partially nude wearing only underwear. Such requirements for a controlled environment and significant user discomfort are undesirable.
[0010] Therefore, the advancement of the prior art is to provide a system and method for accurately extracting body measurements from 2D photos in which 1) the user takes any pose, 2) the user stands in front of any type of background, 3) the photo is taken at any distance, and 4) the user is wearing any type of clothing, such that everyone can easily take their own photo and benefit from full body measurement extraction.
[0011] The present invention has been developed from this background. SUMMARY OF THE INVENTION
[0012] The present invention relates to a method and system for extracting full body measurements using 2D user images, such as obtained from a mobile device camera.
[0013] More specifically, in various embodiments, the present invention is a computer-implemented method for generating human body size measurement results, the computer-implemented method being executable by a hardware processor, the method comprising the steps of: receiving one or more user parameters; receiving at least one image containing the person and the background; identifying one or more body features associated with the person; performing body feature annotation on the identified body features to generate annotation lines corresponding to body feature measurement results on each body feature, the body feature annotation using an annotation deep learning network trained with annotation training data, the annotation training data including one or more images of one or more sample body features and the annotation lines of each body feature; using a sizing machine learning module based on the annotated body features and the one or more user parameters to generate body feature measurement results from the one or more annotated body features; and generating body size measurement results by aggregating the body feature measurement results of each body feature.
[0014] In one embodiment, the step of identifying the one or more body features associated with the person comprises: performing body segmentation on the at least one image to identify the one or more body features associated with the person from the background, the body segmentation using a segmentation deep learning network trained with segmentation training data, and the segmentation training data including one or more images of one or more sample persons and the body feature segmentations of each body feature of the one or more sample persons.
[0015] In one embodiment, the body feature segmentations are extracted under the clothing, and the segmentation training data includes body segmentations of the person's body estimated under the clothing by an annotator.
[0016] In one embodiment, the annotation lines on each body feature include one or more line segments corresponding to a given body feature measurement result, and generating the body feature measurement results from the one or more annotated body features utilizes the annotation lines on each body feature.
[0017] In one embodiment, the at least one image includes at least a front view image and a side view image of the person, and the method further comprises the following steps performed after the body feature annotation step: calculating at least one perimeter of at least one annotated body feature using the front view image and the side view image with line annotations and the person's height; and using the sizing machine learning module based on the at least one perimeter, the height, and the one or more user parameters to generate the body feature measurement results from the at least one perimeter.
[0018] In one embodiment, the sizing machine learning module includes a random forest algorithm and is trained with real data including one or more sample body size measurements of one or more sample persons.
[0019] In one embodiment, the one or more user parameters are selected from the group consisting of height, weight, gender, age, and demographic information.
[0020] In one embodiment, receiving the one or more user parameters includes receiving user input of the one or more user parameters via a user device.
[0021] In one embodiment, receiving the one or more user parameters includes receiving measurements performed by a user device.
[0022] In one embodiment, the at least one image is selected from the group consisting of a front view image of the person and a side view image of the person.
[0023] In one embodiment, the at least one image further includes additional images of the person taken at a 45-degree angle relative to the front view image of the person.
[0024] In one embodiment, performing the body segmentation on the at least one image further includes receiving user input to improve the accuracy of the body segmentation, and the user input includes user selection of one or more parts of the body features corresponding to a given region of the person's body.
[0025] In one embodiment, the at least one image includes at least one image of a fully dressed user or a partially dressed user, and generating the body feature measurements further includes generating the body feature measurements based on the at least one image of the fully dressed user or the partially dressed user.
[0026] In one embodiment, the body size measurements include first body size measurements, and the method further includes using a second sizing machine learning module to generate second body size measurements, and the accuracy of the second body size measurements is higher than the accuracy of the first body size measurements.
[0027] In another embodiment, the method further comprises: determining whether a given body feature measurement of the body feature corresponds to a confidence level below a predetermined value; and in response to determining that the given body feature measurement corresponds to a confidence level below the predetermined value, performing 3D model matching on the body feature using a 3D model matching module to determine a matching 3D model of the person, wherein one or more high-confidence body feature measurements are used to guide the 3D model matching module, performing body feature measurement based on the matching 3D model, and replacing the given body feature measurement with an expected body feature measurement from the matching 3D model.
[0028] In another embodiment, the method further comprises: determining whether a given body feature measurement of the body feature corresponds to a confidence level below a predetermined value; and in response to determining that the given body feature measurement corresponds to a confidence level below the predetermined value, performing bone detection on the body feature using a bone detection module to determine the joint positions of the person, wherein one or more high-confidence body feature measurements are used to guide the bone detection module, performing body feature measurement based on the determined joint positions, and replacing the given body feature measurement with an expected body feature measurement from the bone detection module.
[0029] In another embodiment, the method further comprises preprocessing the at least one image of the person and the background before performing the body segmentation. In one embodiment, the preprocessing at least includes perspective correction of the at least one image. In one embodiment, the perspective correction is selected from the group consisting of perspective correction using the head of the person, perspective correction using a gyroscope of the user device, and perspective correction using another sensor of the user device.
[0030] In another embodiment, the step of identifying one or more body features further comprises: generating a segmentation map of the body feature on the person; and cropping the one or more identified body features from the person and the background before performing the body feature annotation step, and performing the body feature annotation step using a plurality of annotated deep learning networks that have been trained separately for each body feature.
[0031] In various embodiments, a computer program product is disclosed. The computer program can be used to generate human body dimension measurement results and can include a computer-readable storage medium having program instructions or program code incorporated therein, the program instructions being executable by a processor to cause the processor to perform the following steps: receiving one or more user parameters; receiving at least one image containing the human and the background; identifying one or more body features associated with the human; performing body feature annotation on the extracted body features to generate annotation lines corresponding to body feature measurement results on each body feature, the body feature annotation using an annotation deep learning network trained with annotated training data, wherein the annotated training data includes one or more images of one or more sample body features and annotation lines for each body feature; using a sizing machine learning module based on the annotated body features and the one or more user parameters to generate body feature measurement results from the one or more annotated body features; and generating body dimension measurement results by aggregating the body feature measurement results of each body feature.
[0032] According to another embodiment, the present invention is a computer-implemented method for generating human body dimension measurement results, the computer-implemented method being executable by a hardware processor, the method including the following steps: receiving one or more user parameters from a user device; receiving at least one image from the user device, the at least one image containing the human and the background; performing body segmentation on the at least one image to extract one or more body features associated with the human from the background, the body segmentation using a segmentation deep learning network trained with segmentation training data; performing body feature annotation on the extracted body features to annotate annotation lines corresponding to body feature measurement results on each body feature, the body feature annotation using an annotation deep learning network trained with annotated training data, the annotated training data including one or more images of one or more sample body features and annotation lines for each body feature; using a sizing machine learning module based on the annotated body features and the one or more user parameters to generate body feature measurement results from the one or more annotated body features; and generating body dimension measurement results by aggregating the body feature measurement results of each extracted body feature.
[0033] In one embodiment, the segmentation training data includes one or more images of one or more sample humans and manually determined body feature segmentations for each body feature of the one or more sample humans.
[0034] In one embodiment, the manually determined body feature segmentation is extracted under the clothing, and the segmentation training data includes a manually determined body segmentation of the person's body estimated under the clothing by an annotator.
[0035] In one embodiment, the annotation lines on each body feature include one or more line segments corresponding to a given body feature measurement, and generating the body feature measurement from the one or more annotated body features utilizes the annotation lines on each body feature.
[0036] In one embodiment, the at least one image includes at least a front view image and a side view image of the person, and wherein the method further includes the following steps performed after the body feature annotation step: calculating at least one perimeter of at least one annotated body feature using the front view image and the side view image with line annotations and the height of the person; and generating the body feature measurement from the at least one perimeter based on the at least one perimeter, the height, and the one or more user parameters using the sizing machine learning module.
[0037] In one embodiment, the sizing machine learning module includes a random forest algorithm, and the sizing machine learning module is trained with real data, the real data including one or more sample body size measurements of one or more sample persons.
[0038] In one embodiment, the one or more user parameters are selected from the group consisting of height, weight, gender, age, and demographic information.
[0039] In one embodiment, receiving the one or more user parameters from the user device includes receiving user input of the one or more user parameters through the user device.
[0040] In one embodiment, receiving the one or more user parameters from the user device includes receiving measurements performed by the user device.
[0041] In one embodiment, the at least one image is selected from the group consisting of a front view image of the person and a side view image of the person.
[0042] In one embodiment, the at least one image further includes an additional image of the person taken at a 45-degree angle relative to the front view image of the person.
[0043] In one embodiment, performing the body segmentation on the at least one image further includes: receiving user input to improve the accuracy of the body segmentation, and the user input includes a user selection of one or more parts of the extracted body features, the one or more parts corresponding to a given region of the person's body.
[0044] In one embodiment, the at least one image includes at least one image of a fully dressed user or a partially dressed user, and generating the body feature measurement results further includes generating the body feature measurement results based on the at least one image of the fully dressed user or the partially dressed user.
[0045] In one embodiment, the body size measurement results include a first body size measurement result, and the method further includes using a second machine learning module to generate a second body size measurement result, and the accuracy of the second body size measurement result is higher than the accuracy of the first body size measurement result.
[0046] In another embodiment, the method further includes: determining whether a given body feature measurement result of the extracted body features corresponds to a confidence level lower than a predetermined value; and in response to determining that the given body feature measurement result corresponds to a confidence level lower than the predetermined value, performing 3D model matching on the extracted body features using a 3D model matching module to determine a matching 3D model of the person, wherein one or more high-confidence body feature measurement results are used to guide the 3D model matching module, performing body feature measurement based on the matching 3D model, and replacing the given body feature measurement result with an expected body feature measurement result from the matching 3D model.
[0047] In another embodiment, the method further includes: determining whether a given body feature measurement result of the extracted body features corresponds to a confidence level lower than a predetermined value; and in response to determining that the given body feature measurement result corresponds to a confidence level lower than the predetermined value, performing bone detection on the extracted body features using a bone detection module to determine the joint positions of the person, wherein one or more high-confidence body feature measurement results are used to guide the bone detection module, performing body feature measurement based on the determined joint positions, and replacing the given body feature measurement result with an expected body feature measurement result from the bone detection module.
[0048] In another embodiment, the method further includes preprocessing the at least one image of the person and the background before performing the body segmentation, the preprocessing including at least perspective correction of the at least one image, and the perspective correction is selected from the group consisting of perspective correction utilizing the person's head, perspective correction utilizing a gyroscope of the user device, and perspective correction utilizing another sensor of the user device.
[0049] In another embodiment, a computer program product can be used to generate body size measurements of a person and can include a computer-readable storage medium having program instructions or program codes embodied therein, the program instructions being executable by a processor to cause the processor to perform the following steps: receiving one or more user parameters from a user device; receiving at least one image from the user device, the at least one image containing the person and a background; performing body segmentation on the at least one image to extract one or more body features associated with the person from the background, the body segmentation utilizing a segmentation deep learning network trained with segmentation training data; performing body feature annotation on the extracted body features to draw an annotation line corresponding to the body feature measurement on each body feature, the body feature annotation utilizing an annotation deep learning network trained with annotated training data, wherein the annotated training data includes one or more images of one or more sample body features and an annotation line for each body feature; generating body feature measurements from the one or more annotated body features using a sizing machine learning module based on the annotated body features and the one or more user parameters; and generating body size measurements by aggregating the body feature measurements for each body feature.
[0050] In various embodiments, a system is described that includes: a memory that stores computer-executable components; a hardware processor that is operably coupled to the memory and executes the computer-executable components stored in the memory, wherein the computer-executable components may include components communicatively coupled to the processor that performs the aforementioned steps.
[0051] In another embodiment, the present invention is a non-transitory computer-readable storage medium storing executable instructions that, when executed by a processor, cause the processor to perform a process for generating body measurement results, the instructions causing the processor to perform the aforementioned steps.
[0052] In another embodiment, the present invention is a system for extracting whole body measurement results using a 2D phone camera, the system comprising: a user device having a 2D camera, a processor, a display, and a first memory; a server comprising a second memory and a data warehouse; a telecommunications link between the user device and the server; and a plurality of computer codes stored on the first and second memories of the user device and the server, the plurality of computer codes, when executed, causing the server and the user device to perform a process comprising the foregoing steps.
[0053] In another embodiment, the present invention is a computerized server comprising at least one processor, a memory, and a plurality of computer codes stored on the memory, the plurality of computer codes, when executed, causing the processor to perform a process comprising the foregoing steps.
[0054] Other aspects and embodiments of the present invention include methods, processes, and algorithms comprising the steps described herein, and also include the operating procedures and modes of operation of the systems and servers described herein.
[0055] Other aspects and embodiments of the present invention will become apparent from the detailed description of the present invention when read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The embodiments of the present invention described herein are exemplary and not restrictive. Embodiments will now be described by way of example with reference to the drawings, in which:
[0057] Figure 1A An exemplary flowchart of body measurement determination using a deep learning network (DLN) and machine learning according to one embodiment of the present invention is shown.
[0058] Figure 1B Another exemplary flowchart of body measurement determination using a deep learning network (DLN) and machine learning according to another embodiment of the present invention is shown.
[0059] Figure 1C A detailed flowchart of human body measurement determination using a deep learning network (DLN) and machine learning according to another embodiment of the present invention is shown.
[0060] Figure 1D A detailed flowchart of human body part segmentation and annotation using a deep learning network (DLN) according to one embodiment of the present invention is shown.
[0061] Figure 1ESchematic diagram of a machine learning algorithm for determining body dimensions based on one or more eigenvalue obtained from a deep learning network (DLN) according to another embodiment of the present invention.
[0062] Figure 2 Exemplary flowchart for training a deep learning network (DLN) and a machine learning module according to an exemplary embodiment of the present disclosure, the deep learning network and the machine learning module being used together with Figure 1A a flowchart for anthropometric result determination.
[0063] Figure 3 Schematic diagram showing a captured user image (front view) for training a segmentation and annotation DLN, the user image showing a human body wearing clothes.
[0064] Figure 4 Schematic diagram showing an annotator manually segmenting one or more features of a human body under clothing from the background for training a segmentation DLN.
[0065] Figure 5 Schematic diagram showing the body features of a human body segmented from the background for training a segmentation DLN.
[0066] Figure 6 Schematic diagram showing an annotator manually annotating annotation lines for training an annotation DLN.
[0067] Figure 7 Illustrative client - server diagram for implementing body measurement result extraction according to an embodiment of the present invention.
[0068] Figure 8 Exemplary flowchart for body measurement result determination (showing a separate segmentation DLN, annotation DLN, and dimension - setting machine learning module) according to an embodiment of the present invention.
[0069] Figure 9 Another exemplary flowchart for body measurement result determination (showing a combined segmentation - annotation DLN and dimension - setting machine learning module) according to another embodiment of the present invention.
[0070] Figure 10 Another exemplary flowchart for body measurement result determination (showing a combined dimension - setting DLN) according to another embodiment of the present invention.
[0071] Figure 11 Another exemplary flowchart for body measurement result determination (showing a 3D human model and a bone - joint position model) according to another illustrative embodiment of the present disclosure.
[0072] Figure 12 Shows an illustrative hardware architecture diagram of a server for implementing an embodiment of the present invention.
[0073] Figure 13 Shows an illustrative system architecture diagram for implementing an embodiment of the present invention in a client-server environment.
[0074] Figure 14 Shows a schematic diagram of a use case of the present invention, where a single camera on a mobile device is used to capture anthropometric measurements, showing a front view of a person wearing ordinary clothes standing in front of a normal background.
[0075] Figure 15 Shows a schematic diagram of a mobile device graphical user interface (GUI) showing user instructions for capturing a front-facing photo according to an embodiment of the present invention.
[0076] Figure 16 Shows a schematic diagram of a mobile device GUI that requests a user to enter their height (and optionally other user parameters such as weight, age, gender, etc.) and select their preferred style type (tight, regular, or loose style) according to an embodiment of the present invention.
[0077] Figure 17 Shows a schematic diagram of a mobile device GUI for capturing a front-facing photo according to an embodiment of the present invention.
[0078] Figure 18 Shows another schematic diagram of a mobile device GUI for capturing a front-facing photo according to an embodiment of the present invention.
[0079] Figure 19 Shows a schematic diagram of a mobile device GUI for capturing a side-facing photo according to an embodiment of the present invention.
[0080] Figure 20 Shows a schematic diagram of a mobile device GUI displayed when the system processes the captured photo to extract anthropometric measurements according to an embodiment of the present invention.
[0081] Figure 21 Shows a schematic diagram of a mobile device GUI showing a notification screen when anthropometric measurements have been successfully extracted according to an embodiment of the present invention. Detailed Description
[0082] Overview
[0083] Referring to the provided drawings, embodiments of the present invention will now be described in detail.
[0084] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present invention. It will be apparent, however, to one skilled in the art that the present invention may be practiced without these specific details. In other instances, schematic diagrams, usage scenarios, and / or flowcharts are used to illustrate structures, apparatuses, activities, and methods in order to avoid obscuring the present invention. Although the following description contains many specificities for purposes of illustration, one skilled in the art will appreciate that many variations and / or alterations to the presented details are within the scope of the present invention. Similarly, although many features of the present invention are described in relation to one another or in combination with one another, one skilled in the art will appreciate that many of these features may be provided independently of other features. Accordingly, this description of the present invention is presented without loss of generality and without imposing limitations on the present invention.
[0085] Others have tried many different types of methods to generate or extract body measurements from a user's image. All of these methods generally require the user to have a specific pose, stand at a specific distance from the camera, in front of a blank background, wear a tight shirt and / or be partially nude wearing only underwear. Such requirements for a controlled environment and significant user discomfort are undesirable.
[0086] The present invention solves the foregoing problems by providing a system and method for accurately extracting body measurements from 2D photographs, where 1) the user assumes any pose, 2) the user stands in front of any type of background, 3) the photograph is taken at any distance, and 4) the user wears any type of clothing, such that everyone can easily take their own photograph and benefit from full body measurement extraction. Some embodiments of the present invention neither involve any 3D reconstruction or 3D body models, nor do they require a specialized hardware camera. Instead, advanced computer vision combined with deep learning techniques is used to generate accurate body measurements from photographs provided by a simple mobile device camera, regardless of what the user is wearing. In the present disclosure, the term "2D photograph camera" is used to denote any conventional camera embedded in or connected to a computing device such as a smartphone, tablet computer, laptop computer, or desktop computer.
[0087] Deep Learning Networks and Machine Learning for Body Measurements
[0088] Figure 1AA diagram showing an exemplary process for body measurement result determination operations according to an exemplary embodiment of the present disclosure. In some embodiments of the present invention, computer vision technology and deep learning are applied to a frontal photo and a side photo of a user, plus the user's height and possibly other user parameters (such as weight, gender, age, etc.), and one or more deep learning networks are used to generate full-body measurement results. The deep learning networks are trained with annotated body measurement results collected and annotated for thousands of sample people. As the system collects more data, the accuracy of body measurement automatically improves. In some other embodiments, perspective correction, person-background subtraction, bone detection, and 3D model matching methods using computer vision technology are used to improve any low-confidence body measurement results from the deep learning method. This hybrid method significantly improves the accuracy of body measurement results and increases user satisfaction with the body measurement results. In cases where body measurement results are used for customized clothing production, the resulting accuracy improves customer service and reduces the return rate of the customized clothing produced.
[0089] The entire process begins at step 101. At step 102, normalized data (one or more user parameters), such as the user's height, is obtained, generated, and / or measured for performing normalization or scaling. In another embodiment, weight can also be used in combination with height. The two user parameters can be determined automatically (e.g., using computer vision algorithms or mined from one or more databases) or determined according to the user (e.g., user input). In one embodiment, based on these user parameters, the body mass index (BMI) can be calculated. The BMI can be used to calibrate the extraction of body measurement results using weight and height. Additional user parameters can include at least one of the following: height, weight, gender, age, race, country of origin, athletic ability, and / or other demographic information associated with the user, etc. The user's height is used to normalize or scale the frontal and / or side photos and provide a known size reference for the person in the photo. Other user parameters (such as weight, BMI index, age, gender, etc.) are used as additional inputs into the system to optimize body size measurement. In one embodiment, other user parameters can also be automatically obtained from the user device, from one or more third-party data sources, or from the server.
[0090] In step 104, one or more user photographs may be received; for example, at least one front view and / or side view photograph of a given user may be received. In another embodiment, the photographs may be obtained from a user device (e.g., a mobile phone, a laptop computer, a tablet computer, etc.). In another embodiment, the photographs may be obtained from a database (e.g., a social media database). In another embodiment, the user photographs include a photograph showing a front view and a photograph showing a side view of the user's entire body. In some embodiments, only one photograph, such as a front view, is utilized, and one photograph is sufficient to perform accurate body measurement result extraction. In other embodiments, three or more photographs are utilized, including, in some embodiments, a front view photograph, a side view photograph, and a photograph taken at a 45-degree angle. As will be recognized by those of ordinary skill in the art, other combinations of user photographs are within the scope of the present invention. In some embodiments, a user video may be received, such as a front view, 90, 180, or even 360-degree view of the user. From the user video, one or more static frames or photographs, such as a front view, a side view, and / or a 45-degree view, are extracted and the static frames or photographs are used in subsequent processes. Steps 102 and 104 may be performed in any order in various embodiments of the present invention, or the two steps may be implemented simultaneously.
[0091] In one embodiment, as further described below in connection with the following steps, the system may use the photographs and the normalized data to automatically calculate (e.g., using an algorithm of one or more AIs) body measurement results. In another embodiment, to obtain more accurate results, the user may indicate whether the user is wearing tight clothing, normal clothing, or loose clothing.
[0092] In one embodiment, an image can be taken at a specified distance (e.g., about 10 feet away from the camera of the user's device). In one embodiment, an image can be taken when the user assumes a specific pose (e.g., arms in a predetermined position, legs spread shoulder-width apart, back straight, "A pose", etc.). In another embodiment, multiple images of a given position (e.g., a front view photo and a side view photo) can be taken, and an average image of each position can be determined. This process can be performed to improve accuracy. In another embodiment, the user can be located in front of a specific type of background (e.g., a neutral color or having a predetermined background image). In some embodiments, the user can be located in front of any type of background. In one embodiment, the front view photo and the side view photo can be taken under similar lighting conditions (e.g., a given brightness, shadows, and the like). In another embodiment, the front view photo and the side view photo can include an image of the user wearing normally fitting clothes (e.g., not too loose or too tight). Optionally or additionally, depending on the needs of the AI-based algorithm and the associated process, the front view photo and the side view photo can include an image of the user partially dressed (e.g., bare-chested) or wearing different types of styles (e.g., tight, loose, etc.).
[0093] In some embodiments, when needed, one or more preprocessing operations can be performed on the front view photo and the side view photo of the user (not shown in Figure 1A ). For example, the system can use OpenCV, an open-source machine vision library, and can perform perspective correction using the features of the head in the front view photo and the side view photo and the user's height as a reference. In this way, embodiments of the present disclosure can avoid inaccurate measurement results such as determining the proportions of the lengths of the body (such as the trunk length and the leg length). Optionally, a perspective side view photo showing where the camera is located relative to the person being photographed can result in a more accurate perspective correction by allowing the system to calculate the distance between the camera and the user. In some embodiments, the system can instead use gyroscope data provided by the user device (or a peripheral device connected to the user device, such as an additional computer device) to detect the photo perspective angle and perform perspective correction based on this photo perspective angle.
[0094] In some embodiments, one or more additional preprocessing steps (not shown in Figure 1A ) can be performed on one or more photos of the user. Various computer vision techniques can be utilized to further preprocess the one or more images. In addition to perspective correction, examples of preprocessing steps can include contrast, lighting, and other image processing techniques to improve the quality of the one or more images before further processing.
[0095] In step 106, a first deep learning network (DLN) can be used to identify or extract body features from the image, such as a person's body parts (e.g., neck, arms, legs, etc.), and the first deep learning network is referred to as a segmentation DLN. In one embodiment, "deep learning" can refer to a class of machine learning algorithms that use a series of multiple layers of non-linear processing units for feature extraction and transformation, modeled according to neural networks. In one embodiment, consecutive layers can use the output of the previous layer as input. In one embodiment, the "depth" in "deep learning" can refer to the number of layers used to transform the data. In the following, reference is made to Figures 3 to 4 to illustrate and show examples of body feature extraction.
[0096] Before performing this segmentation step on data from a real user, as described with respect to Figure 2 the system can first be trained (e.g.) with sample images of people posing in different environments while wearing different clothes (e.g., with hands at a 45-degree angle, sometimes referred to as the "A pose"). In some embodiments, any suitable deep learning architecture can be used, such as a deep neural network, a deep belief network, and / or a recurrent neural network. In another embodiment, the deep learning algorithm can learn in a supervised (e.g., classification) and / or unsupervised (e.g., pattern analysis) manner. Additionally, the deep learning algorithm can learn multiple representation levels corresponding to different levels of abstraction of the information encoded in the image (e.g., body, body parts, etc.). In another embodiment, the image (e.g., a front view image and a side view image) can be represented as a matrix of pixels. For example, in one embodiment, a first representation level can extract pixels and encode edges; a second level can construct edges and encode the arrangement of the edges; a third level can encode the nose and eyes; and a fourth level can identify that the image contains a face, etc.
[0097] In one embodiment, as described below with respect to Figure 2As described, the segmentation DLN algorithm can be trained using segmented training data. In some embodiments, the segmented training data can include thousands of sample individuals with manually segmented body features. In some embodiments, the training data includes, for example, medical data from CAT scans, MRI scans, etc. In some embodiments, the training data includes data from previous tailors or includes 3D body measurements and "ground truth" data from 3D body scans of 3D body scanners. In some embodiments, in cases where frontal and side view images are not explicitly available, 3D body scans can be used to extract approximate frontal and / or side view images. In some embodiments, the ground truth data includes data measured by human tailors; while in other embodiments, the ground truth data includes automatically extracted 1D body size measurements from 3D body scans. In some embodiments, 3D body scan data from the "SizeUSA" dataset can be utilized, which is a commercial sample of 3D body scans obtained on approximately 10,000 subjects (male and female). In other embodiments, 3D body scan data from the "CAESAR" dataset can be utilized, which is another commercial sample of 3D body scans obtained on approximately 4,000 subjects and also includes ground truth data measured manually by human tailors. In other embodiments, the organization using the present invention can capture their own frontal and side view images and appropriate ground truth data using human tailors for training the segmentation DLN. In other embodiments, the segmentation training data can be automatically generated by one or more algorithms, including one or more deep learning networks, rather than being manually segmented by a human operator.
[0098] In one embodiment (not shown in Figure 1A ), the identified body parts are segmented, separated, or cropped from the rest of the person and the background using the segmentation map generated in step 106. The cropping can be physical or virtual. The portion of the image corresponding to each identified body part can be cropped, segmented, or separated from the rest of the image, and that portion of the image is passed to the annotation step 107. By cropping or separating the identified body parts from the rest of the image, the annotation DLN used in the annotation step 107 can be trained specifically or individually with each separate body part, thereby improving accuracy and reliability.
[0099] In step 107, one or more additional deep learning networks (DLNs), such as an annotation DLN, can be used to determine the annotation lines for each body part identified or extracted in step 106. In one embodiment, there is only one body feature annotation DLN for the entire body. In another embodiment, there is a separate body feature annotation DLN for each body part. The advantage of using a separate body feature annotation DLN for each body part is improved accuracy and reliability of body part measurements. Each body part DLN can be trained separately with separate and unique data for each body part. The specificity of the data for each body part improves the accuracy and reliability of the DLN and also increases the convergence speed of neural network layer training. An example of body feature annotation is illustrated and shown below with reference to Figures 5 to 6 to illustrate and show examples of body feature annotation.
[0100] In one embodiment, the system can generate and extract body feature measurements by using an AI-based algorithm, such as an annotation DLN, e.g., by first generating annotation lines based on signals obtained from the body features. Each annotation line can be different for each body feature and can be drawn differently. For example, for the width or circumference of the biceps, the system can draw a line perpendicular to the bone line at the biceps location; for the chest, the system can instead connect two chest points. Based on the annotation of each body feature, as further described below, body feature measurements can then be obtained by normalizing the height of the user received in step 102.
[0101] Before performing this annotation step on data from real users, as further described below with reference to Figure 2 the system can first be trained, for example, with sample images of people posing in different clothes and environments (e.g., with hands at a 45-degree angle, sometimes referred to as the "A pose"). In other embodiments, annotation training data can be automatically generated by one or more algorithms, including one or more deep learning networks, rather than being manually annotated by a human operator. Segmentation and annotation DLNs are described in more detail with reference to Figure 1B , Figure 1C and Figure 1D to illustrate and show examples of body feature annotation.
[0102] In step 108, one or more machine learning algorithms (e.g., sizing machine learning (ML) algorithms) can be used to estimate the body feature measurements of each body part having the annotation lines generated in step 107. In one embodiment, the sizing ML algorithm includes a random forest machine learning module. In one embodiment, there is a separate sizing ML module for each body part. In some embodiments, the entire body uses one sizing ML module. In one embodiment, the system can use the height received in step 102 as an input to normalize the size estimate to determine the size of the body features. To this end, in one embodiment, the annotation DLN draws a "full body" annotation line indicating the position of the subject's height, where one point represents the subject's sole and the other point represents the subject's head. This "full body" annotation line is used to normalize the other annotation lines according to the known height of the subject provided in step 102. In other words, the height of the subject in the detected image is detected, and this height is used together with the known actual height to normalize all annotation line measurements. This process can be considered "height reference normalization" that normalizes using the known height of the subject as a standard measurement result. In another embodiment, instead of or in addition to the user's height, an object of known size (such as a letter paper or A4-sized paper or a credit card) can be used as a normalization reference.
[0103] In another scenario, additional user demographic data (such as but not limited to the height, BMI index, gender, age, and / or other demographic information associated with the user) received in step 102 is used as an input to the sizing ML algorithm (such as random forest), regarding Figure 1E which is described in more detail below.
[0104] The system can also use other algorithms, means, and media to perform each body feature measurement. The annotation DLN and sizing ML can be implemented as a sizing DLN that annotates and measures each body feature; or can be implemented as two separate modules, namely, an annotation DLN that annotates each body feature and a separate sizing ML module that measures the annotated body features. Similarly, various alternative architectures for implementing the segmentation DLN in step 106, the annotation DLN in step 107, and the sizing ML module in step 108 are described below with respect to Figures 8 to 10 For example, Figure 8 corresponds to Figure 1A the architecture shown in, where the segmentation DLN, annotation DLN, and sizing ML module are separate modules. Conversely, Figure 9 corresponds to an alternative architecture (not shown in Figure 1Ain which the segmentation DLN and the annotation DLN are combined into a single annotation DLN (effectively performing segmentation and annotation), followed by a sizing ML module. Finally, Figure 10 corresponds to another alternative architecture (not shown in Figure 1A in which the segmentation DLN, the annotation DLN, and the sizing ML module are all combined into a single sizing DLN that effectively performs all the functions of segmentation, annotation, and sizing.
[0105] Optionally, at step 110, a confidence level for each body feature measurement can be determined, obtained, or received from the sizing ML module in step 108. In addition to outputting the predicted body measurements for each body feature, the sizing ML module also outputs a confidence level for each predicted body feature measurement, as described below, and the confidence level is subsequently used to determine whether any other method will be used to improve the output. In another embodiment, the confidence level can be based on a confidence interval. Specifically, a confidence interval can refer to a type of interval estimate calculated from the statistical data of observed data (e.g., the encoded image data of the frontal and side photographs), and the interval estimate may contain the true value of an unknown population parameter (e.g., the measurement of a body part). The interval can have an associated confidence level that can quantify the confidence level that the parameter is within the interval. More precisely, the confidence level represents the frequency (i.e., proportion) of the possible confidence intervals that contain the true value of the unknown population parameter. In other words, if confidence intervals are constructed from an infinite number of independent sample statistics using a given confidence level, then the proportion of those intervals that contain the true value of the parameter will be equal to the confidence level. In another embodiment, the confidence level can be specified before examining the data (e.g., the images and the measurements extracted therefrom). In one embodiment, a 95% confidence level is used. However, other confidence levels can be used, e.g., 90%, 99%, 99.5%, etc.
[0106] In various embodiments, the confidence interval and the corresponding confidence level can be determined based on a determination of effectiveness and / or optimality. In another embodiment, effectiveness can refer to the confidence level maintained by the confidence interval, accurately or to a good approximation. In one embodiment, optimality can refer to the construction rule that the confidence interval should use as much information as possible from the dataset (the images and the extracted features and measurements).
[0107] In step 112, it can be determined whether the confidence level is greater than a predetermined value. If it is determined that the confidence level is greater than the predetermined value, then the process can proceed to step 114, where a body feature measurement result with a high confidence level can be output. If it is determined that the confidence level is less than the predetermined value, then the method can proceed to step 116 or step 118. Steps 116 and 118 illustrate one or more optional fallback algorithms for estimating body feature measurement results for those body features for which the deep learning method has a low confidence level. As described below, the high-confidence body feature measurement results from the deep learning method (shown in dashed lines) are then combined with the predicted body feature measurement results from the optional fallback algorithms for the low-confidence body feature measurement results to form a complete set of high-confidence body feature measurement results. As noted, in another embodiment, the confidence level can be specified before examining the data (e.g., images and the measurements extracted therefrom).
[0108] Specifically, in steps 116 and 118, other optional models (e.g., AI-based or computer vision-based models) can be applied. In step 116, and according to one optional embodiment, a 3D human model matching algorithm can be applied. For example, the system can first use OpenCV and / or deep learning techniques to extract the human body from the background. The extracted human body is then matched with one or more known 3D human models to obtain body feature measurement results. Using this technique and a database of existing 3D body scans (e.g., a database of thousands of 3D body scans), the system can match the detected closest body with the points of the 3D body scan. Using the closest matching 3D model, the system can then extract body feature measurement results from the 3D model. This technique is described in more detail below with respect to Figure 11 to describe this technique in more detail.
[0109] Optionally and / or additionally, in step 118, other models such as a bone / joint position model can be applied. In one embodiment, OpenPose (discussed further below), an open-source algorithm for pose detection, can be used to perform bone / joint detection. Using this technique to obtain the bone and joint positions, the system can then use an additional deep learning network (DLN) to draw lines between appropriate points, if needed, where the lines indicate the position in the middle of the bones, which are drawn on top of the user's photograph to indicate the various key bone structures, thereby showing the positions of various body parts (such as the shoulders, neck, and arms). Based on this information, body feature measurement results can be obtained from the appropriate lines. For example, the length of the arm can be determined using the line connecting the shoulder to the wrist. This technique is described in more detail below with respect to Figure 11 to describe this technique in more detail.
[0110] In one embodiment, the 3D model algorithm is combined with the skeleton / joint position model as follows (although this is not explicitly shown in Figure 1A ). Using a database of existing 3D body scans, e.g., a database of thousands of 3D body scans, the system can match the closest bone detection to the bone points of the 3D body scan, showing points and lines indicating the positions of the bones, the positions indicating various key bone structures, thereby showing the positions of various body parts (such as the shoulders, neck, and arms). After matching the closest matching 3D model, the system can extract body feature measurements from the 3D model.
[0111] In either or both cases, at step 120, or at step 122, or at both steps, body feature measurements of high confidence can be inferred (e.g., estimated). Specifically, a process different from the first, lower-confidence deep learning process (e.g., shown and described above in connection with step 108) can be used to perform the estimation of the body feature measurements of high confidence.
[0112] An advantageous feature of this method is that the body feature measurements of high confidence from step 114 (shown as dashed lines) can be used as input to help calibrate other models, such as the 3D human model algorithm in step 116 and the skeleton / joint position model in step 118. That is, the body feature measurements of high confidence obtained from the deep learning method in step 108 can be used to assist other models, e.g., the 3D human model 116 and / or the skeleton / joint position model 118. Subsequently, the other models (116 and / or 118) can be used to obtain predicted body feature measurements of high confidence for those body feature measurements determined in step 112 to have a confidence level below a predetermined value. Thereafter, the predicted body feature measurements of high confidence can replace or supplement the low-confidence body feature measurements from the deep learning method.
[0113] Additionally, at step 124, the body feature measurements of high confidence determined at step 120 and / or step 122 can be used to determine body feature measurements of high confidence. In this way, various models, that is, the 3D human model and the skeleton / joint position model, can all be used to further improve the accuracy of the body feature measurements obtained at step 114. Thus, the body feature measurements of high confidence are aggregated—the body feature measurements of high confidence from step 114 (e.g., the deep learning method) are combined with the predicted body feature measurements of high confidence from step 120 and 122 (e.g., other models).
[0114] In step 126, the high-confidence body feature measurement results are aggregated into complete body measurement results for the entire human body and are then output for use. Specifically, the body measurement results can be output to, for example, a user device and / or a corresponding server associated with a company that produces clothing based on the measurement results. In one embodiment, the output can be in the form of a text message, an email, a written description on a mobile application or website, a combination thereof, and the like. The complete body measurement results can then be used for any purpose, including but not limited to custom clothing production. Those of ordinary skill in the art will recognize that the output of the complete body measurement results can be used for any purpose for which it is useful to achieve accurate and simple body measurements, such as but not limited to fitness, health, shopping, and the like.
[0115] Figure 1B Another exemplary flowchart for determining body measurement results using a deep learning network (DLN) and machine learning according to another embodiment of the present invention is shown. In step 151, input data 152 is received, the input data including a front photo, a side photo, and user parameters (height, weight, age, gender, etc.). In step 153, one or more image processing steps are applied. First, optional image preprocessing (perspective correction, person cropping, size resizing, etc.) steps can be performed. Next, as described in more detail with respect to Figure 1D a deep learning network (DLN) 154 is applied to the image to segment and label body features. Next, as described in more detail with respect to Figure 1E a dimensional setting machine learning module (ML) 156 is applied to the labeled body features to determine body size measurement results based on one or more of the labeling lines and the user parameters. Finally, in step 155, the body size measurement results (e.g., 16 standard body part sizes) are output, schematically shown as output data 158. The output 158 can include dimensional results (a set of standard body size measurement results, such as neck, shoulder, sleeve, height, outside leg seam of trousers, inside leg seam of trousers, etc.) and can also include the front photo and the side photo labeled with the labeling lines.
[0116] Figure 1C A detailed illustrative flowchart for determining body measurement results using a deep learning network (DLN) and machine learning according to another embodiment of the present invention is shown. The inputs into the body measurement process include a front photo 161, a side photo 162, height 163, and other user parameters (weight, age, gender, etc.) 164. The front photo 161 is preprocessed in step 165, while the side photo 162 is preprocessed in step 166. Examples of preprocessing steps, such as perspective correction, person cropping, image size resizing, etc., have been discussed previously. In step 167, the preprocessed front photo is used as DLN1 (with respect toFigure 1D The input of the segmentation-annotation DLN (described in more detail) is used to generate the annotation lines of the front photo 161. In step 168, the preprocessed side photo is used as the input of the DLN 2 (segmentation-annotation DLN) to similarly generate the annotation lines of the side photo 161. The annotation lines 169 of each body part from the front view are output from the DLN 1, and the annotation lines 170 of each body part from the side view are output from the DLN 2. In step 171, the perimeters of each body part are calculated using two sets of annotation lines from the front photo 161 and the side photo 162, as well as the height normalization reference 175 received from the height input 163. In step 172, in a machine learning algorithm (such as random forest (described in more detail)) Figure 1E the perimeters of each body part and the height and other user parameters 176 received from the inputs 163 and 164 are used to calculate one or more body size measurement results. In step 173, the body size measurement results (the length of each standard measurement) are output. Finally, the body measurement process ends in step 174.
[0117] Exemplary deep learning network and machine learning architecture
[0118] Figure 1D A detailed flowchart showing body part segmentation and annotation according to an embodiment of the present invention is shown. In one embodiment, a deep learning network (DLN) using training data as described above is used to complete body part segmentation and annotation. In one embodiment, a convolutional neural network (CNN) is combined with a pyramid scene parsing network (PSPNet) to perform body part segmentation and annotation to obtain improved global and local context information. In PSPNet, the process can utilize the global and local context information from regions of different sizes aggregated by the pyramid pooling module 184. As Figure 1D shown, the input image 181 is first passed through a convolutional neural network (CNN) 182 to obtain a feature map 183, which classifies or segments each pixel into a given body part and / or annotation line. Next, the pyramid pooling module 184 is used to extract the global and local context information from the feature map 183, which aggregates the information from the image at different size scales. Finally, the data is passed through the final convolutional layer 185 to classify each pixel into the body part segmentation line and / or annotation line 186.
[0119] More specifically, according to the input image 181, first, the CNN 182 is used to obtain the feature map 183, and then the pyramid pooling module 184 is used to extract the features of different sub-regions; afterwards, the final feature representation is formed through upsampling and concatenation layers, and the feature representation carries local and global context information. Finally, the feature representation is fed into the last convolutional layer 185 to obtain the final per-pixel prediction. In Figure 1D In the example shown in, the pyramid pooling module 184 combines features according to four different scales. The largest scale is global; subsequent levels divide the feature map into different sub-regions. The outputs of different levels in the pyramid pooling module 184 include feature maps according to different scales. In one embodiment, in order to maintain the weight of the global features, as Figure 1D shown in, a convolutional layer can be used after each pyramid level to reduce the dimension of the context representation. Next, the low-dimensional feature map is upsampled to obtain a feature of the same size as the original feature map. Finally, the different feature levels are concatenated with the original feature map 183 to obtain the output of the pyramid pooling module 184. In one embodiment, by using a four-layer pyramid, as shown, the pooling window covers the whole, half, and smaller parts of the original image 181.
[0120] In one embodiment, the PSPNet algorithm is implemented as described by Hengshuang Zhao et al. in "Pyramid Scene Parsing Network" (CVPR 2017, December 4, 2016, available at arXiv:1612.01105). PSPNet is just an illustrative deep learning network algorithm within the scope of the present invention, and the present invention is not limited to using PSPNet. Other deep learning algorithms are also within the scope of the present invention. For example, in one embodiment of the present invention, a convolutional neural network (CNN) is used to extract body segments (segmentation), and a separate CNN is used to label each body segment (annotation).
[0121] Figure 1ESchematic diagram of a machine learning algorithm for determining body measurement results based on one or more eigenvalue 191 obtained from a deep learning network (DLN) according to another embodiment of the present invention. In one embodiment, a random forest algorithm, an illustrative machine learning algorithm, is used to determine body part dimensions. The random forest algorithm uses multiple decision tree predictors such that each decision tree depends on the values of a random subset of the training data, which minimizes the probability of overfitting the training data set. In one embodiment, the random forest algorithm is implemented as described by Leo Breiman in "Random Forests" (Machine Learning, 45, 5-32, 2001, Kluwer Academic Publishers, Netherlands, available at doi.org / 10.1023 / A:1010933404324). The random forest is just one illustrative machine learning algorithm within the scope of the present invention, and the present invention is not limited to using the random forest. Other machine learning algorithms, including but not limited to nearest neighbor, decision tree, support vector machine (SVM), adaptive boosting, Bayesian network, various neural networks including deep learning networks, evolutionary algorithms, etc., are within the scope of the present invention. The input to the machine learning algorithm is the eigenvalue (x) 191, as described with respect to Figure 1C which includes the circumference, height, and other user parameters of the body part obtained from the deep learning network. The output of the machine learning algorithm is the predicted value of the dimension measurement result (y) 192.
[0122] As noted, embodiments of the devices and systems (and their various components) described herein may employ artificial intelligence (AI) to assist in automating one or more features described herein (e.g., providing body extraction, body segmentation, measurement result extraction, and the like). The components may employ various AI-based solutions to implement the various embodiments / instances disclosed herein. To permit or assist with many of the determinations described herein (e.g., determining, assessing, inferring, calculating, predicting, envisioning, estimating, obtaining, forecasting, detecting, computing), the components described herein may examine all or a subset of the data to which they are authorized access and may permit inferring or determining the state, environment, etc. of the system based on a set of observed data such as via events and / or data capture. The determination may be used to identify a particular context or action, or may produce (e.g.) a probability distribution of a state. The determination may be probability-based—that is, the calculation of the relevant state probability distribution is based on the consideration of data and events. The determination may also refer to techniques for constructing higher-level events from a set of events and / or data.
[0123] Such determination can result in constructing new events or actions from a set of observed events and / or stored event data, regardless of whether the events are temporally close, and regardless of whether the events and data are from one or several event sources and data sources. The components described herein can use various classification (explicit training (e.g., via training data) and implicit training (e.g., via observed behavior, preferences, historical information, received eigen information, etc.)) schemes and / or systems (e.g., support vector machines, neural networks, expert systems, Bayesian belief networks, fuzzy logic, data fusion engines, etc.) to perform automation and / or the determined actions in connection with the claimed subject matter. Thus, classification schemes and / or systems can be used to automatically learn and perform many functions, actions, and / or determinations.
[0124] A classifier can map an input attribute vector z = (z1, z2, z3, z4, …, zn) to a confidence that the input belongs to a class, e.g., as f(z) = confidence(class). Such classification can employ probability-based and / or statistics-based analysis (e.g., taking into account analysis utility and cost) to determine the actions to be automatically performed. Another example of a classifier that can be employed is a support vector machine (SVM). The SVM operates by finding a hyperplane in the space of possible inputs, where the hyperplane attempts to separate triggering criteria from non-triggering events. Intuitively, this corrects the classification for test data that is close to but not the same as the training data. Other directed and undirected model classification methods include (e.g.) naive Bayes, Bayesian networks, decision trees, neural networks, fuzzy logic modules, and / or probability classification models that can provide different independent modes. Classification as used herein also includes statistical regression for forming a priority model.
[0125] Training deep learning networks and machine learning modules
[0126] Figure 2 A diagram showing an exemplary flowchart for training a segmentation DLN, annotation DLN, and sizing ML used in generating body measurement results according to an exemplary embodiment of the present invention. The training process begins at step 201. At step 202, one or more images are received. For example, a front view image and a side view image of a given user can be received. In another embodiment, the images can be obtained from a user device (e.g., a mobile phone, a laptop computer, a tablet computer, etc.). In another embodiment, the images can be obtained from a database (e.g., a social media database). In another embodiment, the images from the user include an image showing a front view of the entire body of the user and an image showing a side view.
[0127] In step 204, the annotator can use human intuition to segment the body features (such as body parts) under the clothes. Specifically, body segmentation can be performed by a human to extract the human body from the background of the photograph, removing the clothes. For example, a human annotator can visually edit (e.g., trace and color-code) the photograph and indicate which body parts correspond to which parts of the photograph to extract the human body from the background, removing the clothes. In one embodiment, the photograph can include a person posing in different clothes in different environments, with both hands at a 45-degree angle (“A pose”). As noted, the accurate body outline can be manually drawn by a human annotator from the background. The ability of a human annotator to determine the body shape of the photographed person under any kind of clothes (especially an experienced and skilled annotator who can provide accurate and reliable body shape annotations) ensures the high performance of the system. The body outline can be drawn on any suitable software platform, and peripheral devices (e.g., a smart pen) can be used to facilitate the annotation. In another embodiment, a printout of the image can be used and manually segmented with a pen / pencil, and the segmented printout can be scanned and recognized by the system using one or more AI-based algorithms (e.g., computer vision-based algorithms). Additionally, at least a portion of such segmented images can be used as training data, which can be fed into a deep learning network in step 208 such that the GPU can learn from the outlines of people in an A pose wearing any clothes in any background. In another embodiment, the segmented body features can be automatically generated (e.g.) according to another training dataset, generated according to a known 3D model, or otherwise generated using another algorithm (including another deep learning network). The manner of training the segmentation DLN is not a limitation of the present invention. In one embodiment, the segmentation DLN used in step 106 of Figure 1A is trained using the segmented images from step 204.
[0128] In step 205, the annotator can then use human intuition to draw estimated annotation (measurement) lines for each body feature under the clothing. As noted, accurate annotation lines can be manually drawn by a human annotator from the background. The ability of a human worker, especially an experienced and skilled annotator who can provide accurate and reliable body shape annotations, to determine the correct annotation lines of the person being photographed under any kind of clothing ensures the high performance of the system. The annotation lines can be drawn on any suitable software platform, and peripheral devices (e.g., smart pens) can be used to facilitate the annotation. In another embodiment, a printout of the image can be used and manually annotated with a pen / pencil, and one or more AI-based algorithms (e.g., computer vision-based algorithms) can be used by the system to scan and recognize the annotated printout. Additionally, at least a portion of such annotated images can be used as training data, which can be fed into a deep learning network in step 210 below, such that the GPU can learn from the annotation lines of a person in an A pose wearing any clothing in any background. In another embodiment, the annotation of the body features can be determined automatically (e.g., according to another training dataset), generated according to a known 3D model, or otherwise generated using another algorithm (including another deep learning network). The way of training the annotation DLN is not a limitation of the present invention.
[0129] The starting point for any machine learning method, such as those used by the deep learning components described above, is a recorded dataset containing multiple instances of system inputs and correct results (e.g., training data). This dataset can be used to train a machine learning system and to evaluate and optimize the performance of the trained system using methods known in the art (including but not limited to standard machine learning methods such as parametric classification methods, nonparametric methods, decision tree learning, neural networks, methods that combine inductive learning with analytical learning, and modeling methods such as regression models). The quality of the output of a machine learning system depends on (a) pattern parameterization, (b) learning machine design, and (c) the quality of the training database. Various methods can be used to refine and optimize these components. For example, the database can be refined by adding a database for new recorded subjects. The quality of the database can be improved, for example, by populating the database with cases customized by one or more professionals customized for clothing. Thus, the database will better represent the knowledge of professionals. In one embodiment, the database includes, for example, data on ill-fitting designs, which can help evaluate the trained system.
[0130] In step 206, actual human measurements of each body feature (e.g., 1D measurements determined by a tailor or obtained from a 3D body scan) can be received to be used as ground truth data. The actual human measurements can be used as valid data and for training the algorithms used by the system. For example, the actual human measurements can be used to minimize an error function or loss function associated with a machine learning algorithm (mean squared error, likelihood loss, log loss, hinge loss, etc.). In one embodiment, the annotation lines from step 205 and the ground truth data from step 206 are used to train the annotation DLN used in Figure 1A step 107 and the sizing ML step 108.
[0131] In one embodiment, human measurements can be received from a user input (e.g., an input into a user device such as a smart phone). In another embodiment, human measurements can be received from a network (e.g., the Internet) via a website, for example. For instance, a tailor can upload one or more measurements to a website, and the system can receive the measurements. As noted, in another embodiment, the actual measurements can be used to train and / or improve the accuracy of the results of an AI-based algorithm (e.g., a deep learning algorithm), which will be discussed below. The manner in which the segmentation DLN and the annotation DLN and the sizing ML module are trained is not a limitation of the present invention.
[0132] In step 208, the segmentation DLN can be trained with body segmentation or body feature extraction. In one embodiment, the annotated human segmentation obtained from step 204 can be used to train the segmentation DLN. For example, labeled data (e.g., an image of a user and the associated actual body segmentation) can be presented to the segmentation DLN, and the segmentation DLN can determine an error function (e.g., according to a loss function as discussed above) based on the results of the segmentation DLN and the actual body segmentation. The segmentation DLN can be trained to reduce the magnitude of this error function.
[0133] In another embodiment, the segmentation DLN can be validated by accuracy estimation techniques such as a holdout method, which can divide data (e.g., all images, including images with corresponding segmentations and images that will have segmentations extracted using the segmentation DLN and do not have corresponding segmentations) into a training set and a test set (conventionally a 2 / 3 training set and 1 / 3 test set design) and can evaluate the performance of the segmentation DLN model based on the test set. In another embodiment, an N-fold cross-validation method can be used, where the method randomly divides the data into k subsets, where k - 1 instances of the data are used to train the segmentation DLN model and the kth instance is used to test the predictive ability of the segmentation DLN model. In addition to the holdout method and the cross-validation method, a bootstrap method can also be used, which repeatedly samples n instances from a data set and can be used to assess the accuracy of the segmentation DLN model.
[0134] In step 210, one or more annotation DLNs for each body feature can be trained, or a single annotation DLN for the entire body can be trained. For example, sixteen annotation DLNs can be trained, one for each of the sixteen different body parts. In one embodiment, the annotations obtained from step 205 can be used to train the annotation DLN. For example, labeled data (e.g., images of a user's body features with line annotations) can be presented to the annotation DLN, and the annotation DLN can determine an error function (e.g., according to a loss function, as discussed above) based on the results of the annotation DLN and the actual annotations. The annotation DLN can be trained to reduce the magnitude of this error function.
[0135] In another embodiment, the annotation DLN can be specifically trained to produce annotation lines for a particular body feature (e.g., a particular body part such as an arm, a leg, a neck, etc.). In another embodiment, the training of the annotation DLNs for each body feature can be performed sequentially (e.g., hierarchically, where groups of related body features are trained one after another) or simultaneously. In another embodiment, different training data sets can be used for different annotation DLNs, which correspond to different body features or body parts. In one embodiment, there can be more or fewer than sixteen DLNs for the sixteen body parts, e.g., depending on the computing resources. In another embodiment, the training of the annotation DLN can be performed at least partially in the cloud, which will be described below.
[0136] In addition, at step 210, one or more sizing ML modules for each body feature can be trained, or a single sizing ML module for the entire body can be trained. In one implementation, the sizing ML module can be trained using the measurements obtained from step 206. For example, labeled data (e.g., the lengths of the annotation lines and the associated actual measurement data) can be presented to the sizing ML module, and the sizing ML module can determine an error function (e.g., according to a loss function as discussed above) based on the results of the sizing ML module and the actual measurement results. The sizing ML module can be trained to reduce the magnitude of this error function.
[0137] In another implementation, the sizing ML module can be specifically trained to extract the measurements of specific body features (e.g., specific body parts such as arms, legs, necks, etc.). In another implementation, the training of the sizing ML modules for each body feature can be performed sequentially (e.g., in a hierarchical manner where groups of related body features are trained one after another) or simultaneously. In another implementation, different training data sets can be used for different sizing ML modules corresponding to different body features or body parts. In one implementation, there can be more or fewer than sixteen sizing ML modules for sixteen body parts, e.g., depending on the computing resources. In another implementation, the training of the sizing ML module can be performed at least partially in the cloud, which will be described below.
[0138] At step 212, the trained segmentation DLN, annotation DLN, and sizing ML modules that will be used in Figure 1A , Figure 1B and Figure 1C can be output. Specifically, the segmentation DLN trained in step 208 is output for use in step 106 in Figure 1A . Similarly, one or more annotation DLNs trained in step 210 are output for use in step 107 in Figure 1A . Finally, the sizing ML module trained in step 210 is output for use in step 108 in Figure 1A .
[0139] Figures 3 to 6 shows a schematic diagram of a graphical user interface (GUI) corresponding to the process steps in Figure 2 for generating training data to train the segmentation DLN and the annotation DLN. Figure 3 shows a schematic diagram of a captured user image for training the segmentation DLN, which shows a human body wearing clothes. Although in Figures 3 to 6A specific user pose, namely, the "A-pose", is shown, but one of ordinary skill in the art will understand that any pose (such as the A-pose, hands at sides, or any other pose) falls within the scope of the present invention. The optimal pose will clearly show the legs and arms separated from the body. One advantage of the present invention is that a person can stand in front of any type of background in almost any reasonable pose. A person does not need to stand in front of a blank background or make special arrangements for the place where the photograph is taken.
[0140] Figure 4 A schematic diagram showing a annotator or operator manually segmenting one or more features of a human body under clothing from a background to train a segmentation DLN. In Figure 4 , the annotator manually annotates the position of the left leg under the clothing. People have rich experience in looking for other people and estimating the shape of their bodies under clothing, and this data is used to train the segmentation DLN to automatically perform similar operations on new photographs of unknown people. Figure 5 A schematic diagram showing the body features of a human body segmented from a background after the annotator has successfully annotated all body features. This data is used to train a human segmentation DLN. Figure 5 Provide the manually annotated data for training the segmentation DLN in Figure 2 Step 208. Subsequently, in Figure 1A Step 106, use the segmentation DLN trained with the data obtained in Figure 2 Step 208 in Figure 5 obtained in.
[0141] Figure 6 A schematic diagram showing an annotator manually drawing annotation lines for training an annotation DLN. This data is used to train the annotation DLN to automatically draw annotation lines for each body feature. Figure 6 Provide the manually annotated data for training the annotation DLN in Figure 2 Step 210. Subsequently, in Figure 1A Step 107 in Figure 2 Step 210 in Figure 6 obtained in for training the annotation DLN.
[0142] Although in Figures 3 to 6Only the front view is shown, but those of ordinary skill in the art will recognize that views in any other orientation (including side views, 45-degree views, top views, etc.) are within the scope of the present invention, depending on the type of anthropometric measurements desired. For example, in one embodiment, a side view of a person is similarly segmented and labeled for estimating the circumference of body parts. As another example, for head measurements for making custom hats, a top-down photo of a person's head will be optimal. Similarly, for face measurements for sizing glasses, optical instruments, etc., only a photo of the front face will be optimal. Close-up photos of the front and back of a human hand can be used for sizing custom gloves, custom PPE (personal protective equipment) for the hands, custom nails, etc.
[0143] In one embodiment, similar to directed supervised learning, a person can also be arranged to assist the deep learning network in the computational process. A human annotator can manually adjust or edit the results from the segmentation DLN and / or the annotation DLN to obtain more accurate sizing results. The adjustment data made by the human annotator to the segmentation map and the annotation map from the deep learning network can be used in the feedback loop interface of the deep learning network to automatically improve the DNL model over time.
[0144] Optional deep learning network (DLN) architectures
[0145] Figure 7 An illustrative client-server diagram for implementing body measurement result extraction according to an embodiment of the present invention is shown. The client side (user) 709 is shown at the top, while the server side 703 is shown at the bottom. The client side starts the process by sending front and side images at 702. After receiving the images, the server checks the images at 704 to determine the correctness of the format and perform other format checks. If at 705 the images are not in the correct format or have other format issues, such as incorrect pose, weak contrast, too far or too close, the subject not in the field of view, the subject partially obscured, etc., then the process returns this information to the client at 701. At 701, in one embodiment, an error message or other communication can be displayed to the user to enable the user to retake the image.
[0146] If the image at 705 is in the correct format and has no other formatting issues, then the image is preprocessed at 706 such that it can be processed by a DLN (Deep Learning Network). As described in more detail previously, the image is then processed by the DLN at 708 to determine dimensions. The dimension results or full body measurement results are returned from the server at 710. At 712, the client checks the dimension results. If, as determined at 713, the dimension results have any form issues, such as out of bounds, too small or too large, etc., then the process returns to 701 and an error message is similarly displayed, or other communication can be shown to the user to enable the user to retake the image. If, as determined at 713, the dimension results have no form issues, then the process ends where the full body measurement results are ready to be used.
[0147] Figure 8 Figure showing an exemplary flowchart for body measurement determination (using separate segmentation DLN, annotation DLN, and dimension setting ML modules) according to an embodiment of the present invention. In one embodiment, a front and side image are received from a user at 802. The image is preprocessed at 804. As previously discussed, in some embodiments, when needed, preprocessing of one or more images of the user can be performed on the front and side photographs, such as perspective correction. For example, the system can use OpenCV, i.e., an open source machine vision library, and can use the features of the head and the height of the user in the front and side photographs as a reference for perspective correction. Various computer vision techniques can be utilized to further preprocess the one or more images. In addition to perspective correction, examples of preprocessing steps can include contrast, illumination, and other image processing techniques to improve the quality of the one or more images before further processing.
[0148] After preprocessing, as previously discussed, the preprocessed image is sent at 806 to a segmentation DLN to generate a segmentation map. At 814, the segmentation map is aggregated with the remaining data. Concurrent with the segmentation, in one embodiment, as previously discussed, the preprocessed image is also sent at 808 to an annotation DLN to generate an annotation measurement line. At 814, the annotation map is aggregated with the remaining data. In one embodiment, as previously discussed, the annotation map is provided to a sizing machine learning (ML) module 810 to generate body feature measurements of each body feature that has been segmented and annotated by measuring each annotation line. At 814, the sizing results are aggregated with the remaining data. At 812, the sizing results are output to one or more external systems for various purposes. Finally, at 1016, all the aggregated and structured data that has been aggregated at 814, (1) preprocessed front and side images, (2) segmentation map, (3) annotation map, and (4) sizing results, are stored in a database for further DLN training.
[0149] Figure 9 FIG. showing another exemplary flowchart for body measurement determination (using a combined segmentation-annotation DLN and a sizing ML module) according to another embodiment of the present invention. As previously discussed, front and side images are received from a user at 902, and the images are preprocessed at 904. Examples of preprocessing steps can include perspective correction, contrast, lighting, and other image processing techniques to improve the quality of the one or more images before further processing.
[0150] After preprocessing, as previously discussed, the preprocessed image is sent directly at 918 to an annotation DLN to generate an annotation map. Instead of first performing body feature segmentation, in this alternative embodiment, annotation lines are drawn directly on the image without the need to use a specially trained combined segmentation-annotation DLN to explicitly segment body features from the background, where the combined segmentation-annotation DLN effectively combines the features of a segmentation DLN and an annotation DLN (shown in the embodiment of Figure 8 into a single annotation DLN as shown in Figure 9 . In effect, body feature segmentation is performed implicitly by the annotation DLN. At 914, the annotation map is aggregated with the remaining data.
[0151] In one embodiment, as previously discussed, the annotated map is provided to the sizing machine learning (ML) module 910 to generate body feature measurements for each body feature that has been annotated by measuring each annotation line. At 914, the sizing results are aggregated with the remaining data. At 912, the sizing results are output to one or more external systems for various uses as described herein. Finally, at 916, all the aggregated and structured data that has been aggregated at 914, (1) the preprocessed front and side images, (2) the annotated map, and (3) the sizing results, are stored in a database for further DLN training.
[0152] Figure 10 FIG. showing another exemplary flowchart for body measurement determination (using a combined sizing DLN) according to another embodiment of the present invention. As previously discussed, at 1002, front and side images are received from a user, and at 1004, the images are preprocessed. Examples of preprocessing steps include perspective correction, contrast, lighting, and other image processing techniques to improve the quality of the one or more images before further processing.
[0153] After preprocessing, as previously discussed, at 1010, the preprocessed images are directly sent to the sizing DLN to generate complete body feature measurements. Instead of first performing body feature segmentation and annotation of measurement lines and then measuring the lines, in this alternative embodiment, body features are directly extracted from the preprocessed images without using a specially trained sizing DLN to explicitly segment body features from the background (and without explicitly drawing annotation lines), and the sizing DLN effectively combines the features of the segmentation DLN, annotation DLN, and measurement machine learning module (shown in the embodiment of Figure 8 into the single sizing DLN shown in Figure 10 . In fact, body feature segmentation and annotation of measurement lines are implicitly performed by the sizing DLN.
[0154] At 1014, the sizing results are aggregated with the remaining data. At 1012, the sizing results are output to one or more external systems for various uses as described herein. Finally, at 1016, all the aggregated and structured data that has been aggregated at 1014, (1) the preprocessed front and side images and (2) the sizing results, are stored in a database for further DLN training.
[0155] 3D Model and Bone / Joint Position Model Embodiments
[0156] Figure 11A diagram showing another exemplary process flow for body measurement determination operations according to an exemplary embodiment of the present disclosure. The process starts at step 1101. At step 1102, user parameters (e.g., height, weight, demographic data, athletic ability, and the like) can be received from the user and / or parameters automatically generated by a phone camera can be received. In additional aspects, the user parameters can be determined automatically (e.g., using computer vision algorithms or mined from one or more databases) or determined according to the user (e.g., user input). In another embodiment, based on these parameters, the body mass index (BMI) can be calculated. As noted, the BMI (or any other parameter determined above) can be used to calibrate weight against height.
[0157] At step 1104, an image of the user (e.g., a first image and a second image representing a front view and a side view of the user's full body) can be received, and an optional third image (e.g., a 45-degree view between the front view and the side view, which can be used to improve the accuracy of subsequent algorithms) can be received. In another embodiment, the image can be obtained from a user device (e.g., a mobile phone, a laptop computer, a tablet computer, etc.). In another embodiment, the image can be determined from a database (e.g., a social media database). In another embodiment, to obtain more accurate results, the user can indicate whether he or she is wearing tight clothes, normal clothes, or loose clothes. In some optional embodiments, as described above, perspective correction can be performed on the front and side view images when needed.
[0158] At step 1106, person segmentation (e.g., extracting the person from the background of the image) can be performed, and a 3D model can be fitted to the extracted person. Additionally, 3D modeling techniques can be used to estimate the 3D shape. In one embodiment, the system can utilize deep learning techniques and / or OpenCV to extract the human body, including the clothes, from the background. Before performing this step on data from real users, the system can first be trained (e.g.) with sample images of people posing with their hands at 45 degrees (“A pose”) wearing different clothes in different environments.
[0159] In step 1108, bone detection can be used to determine the joint positions and postures of a person; alternatively, a pose estimation algorithm (such as OpenPose, i.e., an open-source algorithm) can be used to perform the determination for pose detection. In one embodiment, body pose estimation can include algorithms and systems for recovering the pose of an articulated body, which is composed of joints and rigid parts using image-based observation data. In another embodiment, OpenPose can include a real-time multi-person system to jointly detect human body, hand, face, and foot key points (a total of 135 key points) on a single image. In one embodiment, key points can refer to the estimated parts of a person's pose, such as the nose, right ear, left knee, right foot, etc. The key points contain positions and key point confidence scores. Other aspects of OpenPose functionality include but are not limited to 2D real-time multi-person key point body estimation. The functionality can also include the ability of the algorithm to be a runtime invariant of the number of people detected. Another aspect of its functionality can include but may not be limited to 3D real-time single-person key point detection, including 3D triangulation from multiple single views.
[0160] In step 1110, body size measurements can be determined based on the estimated three-dimensional shape, joint positions, and / or postures. In another embodiment, the system can use inputs of height, weight, and / or other parameters (such as BMI index, age, gender, etc.) to determine the sizes of body parts. In one embodiment, the system can partially use the Virtuoso algorithm, i.e., an algorithm that provides standard DaVinci models of human body parts and the relevant sizes of body parts.
[0161] Additionally, as described above, the system can generate and extract body measurements by using an AI-based algorithm (such as the DLN algorithm), e.g., by drawing measurement lines according to signals obtained from bone points. Specifically, the system can look at one or more bone points, calculate the bones that form an angle with the edge of the user's body, and draw measurement lines in certain orientations or directions. Each measurement line can be different for each body part and can be drawn differently. The system can also use other algorithms, averages, medians, and other resources.
[0162] Additionally, the body measurements can be output to the user device and / or the corresponding server. In one embodiment, the output can be in the form of text messages, emails, written descriptions on a mobile application or website, combinations thereof, and the like.
[0163] In step 1112, the body size measurement result can be updated by using a supervised deep learning algorithm that utilizes training data, which includes manually determined body detections under the clothing. In some aspects, as described above, any suitable deep learning architecture can be used, such as deep neural networks, deep belief networks, and recurrent neural networks. In one implementation, the training data can be obtained from annotator input, as described above, where the annotator input extracts the human body in a given image from the background of the image and removes the clothing. Additionally, at least a portion of such annotated images can be used as training data, which can be fed into the deep learning network so that the GPU can learn from the silhouettes of people wearing clothing in any background. Briefly, in some implementations, the deep learning algorithms described above can be used in combination to improve the accuracy and reliability of the 3D model and the bone / joint position method.
[0164] Hardware, software, and cloud implementations of the present invention
[0165] As discussed, the data described in this disclosure (e.g., images, text descriptions, and the like) can include data stored on a database that is stored or hosted on a cloud computing platform. It should be understood that although this disclosure includes a detailed description of cloud computing below, the implementation of the teachings recited herein is not limited to a cloud computing environment. Instead, embodiments of the present invention are capable of being implemented in conjunction with any other type of computing environment now known or later developed.
[0166] Cloud computing can refer to a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services), which can be rapidly provisioned and released with minimal management effort or interaction with the service provider. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0167] Features can include one or more of the following. On-demand self-service: Cloud consumers can automatically and unilaterally provision computing capabilities, such as server time and network storage, as needed without human interaction with the service provider. Broad network access: Capabilities can be obtained over a network and accessed through standard mechanisms, thus facilitating use of the capabilities by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs). Resource pooling: The provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically assigned and reassigned according to demand. There is a sense of location independence, as consumers generally do not control or know the exact location of the provided resources but can be able to specify location at a higher level of abstraction (e.g., country, state, or data center). Rapid elasticity: Capabilities can be provided quickly and elastically (automatically in some cases) to scale out rapidly and can be released quickly to scale in rapidly. For the consumer, the capabilities available for provisioning generally appear to be infinite and can be purchased in any quantity at any time. Measured service: The cloud system automatically controls and optimizes resource use by leveraging metering capabilities at a certain level of abstraction appropriate to the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource use can be monitored, controlled, and reported, providing transparency for both the provider and consumer of the utilized service.
[0168] In another embodiment, the service model can include one or more of the following. Software as a Service (SaaS): The capabilities provided to the consumer will use the provider's applications running on the cloud infrastructure. The applications can be accessed from various client devices via a thin client interface (such as a web browser) (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application capabilities, although limited user-specific application configuration settings may be an exception.
[0169] Platform as a Service (PaaS): The capabilities provided to the consumer will be deployed onto applications created or acquired by the consumer on the cloud infrastructure using provider-supported programming languages and tools. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but does need to control the deployed applications and possibly control application hosting environment configuration.
[0170] Infrastructure as a Service (IaaS): The ability provided to the consumer is to offer processing, storage, networking, and other fundamental computing resources, where the consumer can deploy and run any software, which may include an operating system and applications. The consumer does not manage or control the underlying cloud infrastructure, but controls the operating system, storage, deployed applications, and may have limited control over selected network components (e.g., host firewall).
[0171] Deployment models can include one or more of the following. Private cloud: The cloud infrastructure works only for a certain organization. The cloud infrastructure can be managed by the organization or a third party and can be deployed on-premises or externally.
[0172] Community cloud: The cloud infrastructure is shared by several organizations and supports a specific community with common concerns (e.g., tasks, security requirements, policies, and compliance considerations). The cloud infrastructure can be managed by the organization or a third party and can be deployed on-premises or externally.
[0173] Public cloud: The cloud infrastructure can be used by the public or large industrial clusters and is owned by the organization selling the cloud services.
[0174] Hybrid cloud: The cloud infrastructure is a combination of two or more clouds (private cloud, community cloud, or public cloud) that remain distinct entities but are combined through standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting that occurs to achieve load balancing between clouds).
[0175] The cloud computing environment is service-oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. The key to cloud computing is the infrastructure of a network including interconnected nodes.
[0176] The cloud computing environment can include one or more cloud computing nodes. Local computing devices used by cloud consumers (such as, for example, personal digital assistants (PDAs) or cellular phones, desktop computers, laptop computers, and / or in-vehicle computer systems) can communicate with the cloud computing nodes. The nodes can communicate with each other. In one or more networks (such as the private cloud, community cloud, public cloud, or hybrid cloud or a combination thereof described above), the nodes can be physically or virtually grouped. This allows the cloud computing environment to provide infrastructure, platform, and / or software as a service, for which the cloud consumer does not need to maintain resources on the local computing device. It should be understood that the types of computing devices are intended to be exemplary only, and the computing nodes and cloud computing environment can communicate with any type of computerized device via any type of network and / or network addressable connection (e.g., using a web browser).
[0177] The present invention can be implemented using server-based hardware and software. Figure 12 An illustrative hardware architecture diagram of a server showing an embodiment for implementing the present invention is presented. Many components of the system, such as network interfaces, etc., are not shown so as not to obscure the present invention. However, those of ordinary skill in the art will understand that the system will necessarily include these components. The user device is hardware that includes at least one processor 1240, and the at least one processor is coupled to a memory 1250. The processor may represent one or more processors (e.g., microprocessors), and the memory may represent a random access memory (RAM) device, including the main storage device of the hardware and any supplementary levels of memory, such as cache memory, non-volatile or backup memory (e.g., programmable or flash memory), read-only memory, etc. Additionally, the memory may be considered to include memory storage devices physically located elsewhere in the hardware, such as any cache memory in the processor, and any storage capacity used as virtual memory, such as that stored on a mass storage device.
[0178] The hardware of the user device also typically receives many inputs 1210 and outputs 1220 to communicate information externally. For the interface with the user, the hardware may include one or more user input devices (e.g., keyboard, mouse, scanner, microphone, web camera, etc.) and a display (e.g., liquid crystal display (LCD) panel). For additional storage, the hardware may also include one or more mass storage devices 1290, such as floppy disk or other removable disk drives, hard disk drives, direct access storage devices (DASD), optical drives (e.g., compact disc (CD) drives, digital versatile disc (DVD) drives, etc.) and / or tape drives, etc. In addition, the hardware may include an interface with one or more external SQL databases 1230 and one or more networks 1280 (e.g., local area network (LAN), wide area network (WAN), wireless network, and / or the Internet, etc.) to permit communication of information with other components coupled to the network. It should be understood that the hardware typically includes suitable analog and / or digital interfaces to communicate with each other.
[0179] The hardware operates under the control of an operating system 1270 and executes various computer software applications 1260, components, programs, code, program libraries, object programs, modules, etc., commonly indicated by reference numerals, to perform the methods, processes, and techniques described above.
[0180] The present invention can be implemented in a client-server environment. Figure 13Illustrates an exemplary system architecture for implementing an embodiment of the present invention in a client-server environment. User device 1310 on the client side may include a smart phone 1312, a laptop computer 1314, a desktop PC 1316, a tablet computer 1318, or other devices. Such user devices 1310 access the services of system server 1330 through some network connection 1320 (such as the Internet).
[0181] In some embodiments of the present invention, in a so-called cloud implementation, the entire system can be implemented and provided to end users and operators via the Internet. Neither software nor hardware needs to be installed locally, and end users and operators will be allowed to directly access the system of the present invention via a web browser or similar software on the client, which can be a desktop computer, a laptop computer, a mobile device, etc. This eliminates any need for custom software installation on the client side and increases the flexibility of service (Software as a Service) provision, and increases user satisfaction and ease of use. Various business models, revenue models, and delivery mechanisms are envisioned for the present invention, and all of the various business models, revenue models, and delivery mechanisms are considered to be within the scope of the present invention.
[0182] Generally speaking, the methods performed to implement the embodiments of the present invention can be implemented as part of an operating system or a specific application program, component, program, object program, module, or sequence of instructions (referred to as a "computer program" or "computer program code"). The computer program generally includes one or more sets of instructions in various memories and storage devices in a computer at various times, and when read and executed by one or more processors in the computer, causes the computer to perform the operations required to execute the elements including various aspects of the present invention. In addition, although the present invention has been described in the context of full-function computers and computer systems, those skilled in the art will understand that the various embodiments of the present invention can be distributed as a program product in various forms, and the present invention is equally applicable regardless of the specific type of machine or computer-readable medium used to actually implement such distribution. Examples of computer-readable media include, but are not limited to, recordable-type media, such as volatile and non-volatile memory devices, floppy disks and other removable disks, hard disk drives, optical disks (e.g., compact disc read-only memory (CD ROM), digital versatile disk (DVD), etc.), and digital and analog communication media.
[0183] Exemplary use cases of the present invention
[0184] Figure 14 Is a schematic diagram of a use case of the present invention, in which a single camera on a mobile device is used to capture anthropometric results, showing a front view of a person wearing ordinary clothes standing in front of a normal background.Figure 14 The mobile device shown includes at least one camera, a processor, a non-transitory storage medium, and a communication link to a server. In one embodiment, one or more images of a user's body are transmitted to the server, which performs the operations described herein. In one embodiment, the one or more images of the user's body are analyzed locally by the processor of the mobile device. The operations performed return one or more body measurement results, which can be stored on the server and presented to the user. Additionally, the body measurement results can subsequently be utilized for a number of purposes, including but not limited to selling one or more custom-made clothing items, custom glasses, custom gloves, custom compression garments, custom PPE (personal protective equipment), custom hats, custom diet groups, custom training, fitness, and exercise routines, etc. Without loss of generality, the body measurement results can be output, transmitted, and / or utilized for any purpose for which they are useful.
[0185] Finally, Figures 15 to 21 An illustrative mobile graphical user interface (GUI) is shown in which some embodiments of the present invention have been implemented. Figure 15 A schematic diagram of a mobile device GUI showing user instructions for capturing a frontal image according to one embodiment of the present invention is shown. Figure 16 A schematic diagram of a mobile device GUI requesting a user to enter their height (and optionally other demographic information such as weight, age, etc.) and select their preferred style type (tight, regular, or loose style) according to one embodiment of the present invention is shown. Figure 17 A schematic diagram of a mobile device GUI for capturing a frontal image according to one embodiment of the present invention is shown. Figure 18 Another schematic diagram of a mobile device GUI for capturing a frontal image with an illustrative A - pose shown in dashed lines according to one embodiment of the present invention is shown. Figure 19 A schematic diagram of a mobile device GUI for capturing a side image according to one embodiment of the present invention is shown. Figure 20 A schematic diagram of a mobile device GUI shown while the system processes the image to extract body measurement results according to one embodiment of the present invention is shown. Finally, Figure 21 A schematic diagram of a mobile device GUI showing a notification screen when the body measurement results have been successfully extracted according to one embodiment of the present invention is shown.
[0186] The present invention has been successfully implemented, resulting in body measurement results with an accuracy difference of less than 1 cm compared to human tailors. The system is capable of achieving comparable accuracy to human tailors using only two photographs. The system does not require the use of any specialized hardware sensors, does not require the user to stand in front of any special background, does not require special lighting, and can be used for photographs taken at any distance and with the user wearing any type of clothing. The result is a body measurement system that works with any mobile device, enabling anyone to easily take their own photographs and benefit from automatic full-body measurement result extraction.
[0187] Those of ordinary skill in the art will appreciate that user cases, structures, diagrams, and flowcharts can be performed in other orders or combinations, but the innovative concepts of the present invention remain without departing from its broader scope. Each implementation can be unique, and the methods / steps can be shortened or lengthened, overlapped with other activities, postponed, delayed, and continued after a certain time interval, such that each user is suitable for implementing the method of the present invention.
[0188] Although the present invention has been described with reference to specific exemplary implementations, it will be apparent that various modifications and changes can be made to these implementations without departing from the broader scope of the present invention. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. It will also be obvious to those skilled in the art that the above-described implementations are specific instances of a single broader invention, which may have a scope greater than any single description taught. There may also be many changes to the description without departing from the scope of the present invention.
Claims
1. A computer-implemented method for generating human body size measurement results, the computer-implemented method being executable by a hardware processor, the method comprising: receiving one or more user parameters; receiving at least one image containing the person wearing clothes and a background; performing body segmentation on the at least one image to identify one or more body features associated with the person from the background, the body segmentation identifying the one or more body features including one or more body parts from under the clothes; performing body feature annotation on the identified body features to generate annotation lines corresponding to body feature measurement results on each body feature, the body feature annotation annotating at least one body feature of the person under the clothes on the at least one image, the body feature annotation using an annotation deep learning network trained with annotated training data, the annotated training data including one or more images of one or more sample body features and annotation lines for each body feature; using a sizing machine learning module based on the annotated body features and the one or more user parameters to generate body feature measurement results of the person under the clothes from the one or more annotated body features; and generating body size measurement results by aggregating the body feature measurement results of each body feature.
2. The method according to claim 1, wherein the body segmentation uses a segmentation deep learning network trained with segmentation training data, and wherein the segmentation training data includes one or more images of one or more sample persons and body feature segmentations of each body feature of the one or more sample persons, and wherein the segmentation training data includes manually measured data, 3D body size measurement results from 3D body scans, and / or segmentation training data automatically generated by one or more algorithms.
3. The method according to claim 2, wherein the body feature segmentations are extracted from under the clothes, and wherein the segmentation training data includes body segmentations of the person under the clothes estimated by an annotator.
4. The method according to claim 1, wherein the annotation lines on each body feature include one or more line segments corresponding to a given body feature measurement result, and wherein generating the body feature measurement results from the one or more annotated body features utilizes the annotation lines on each body feature.
5. The method according to claim 1, wherein the at least one image includes at least a front view image and a side view image of the person, and wherein the method further comprises the following steps after performing the body feature annotation step: calculating at least one perimeter of at least one annotated body feature using the front view image and side view image with line annotations and the height of the person; and using the sizing machine learning module based on the at least one perimeter, the height, and the one or more user parameters to generate the body feature measurement results from the at least one perimeter.
6. The method according to claim 1, wherein the size-setting machine learning module includes a random forest algorithm, and wherein the size-setting machine learning module is trained with real data, the real data including one or more sample body size measurement results of one or more sample persons.
7. The method according to claim 1, wherein the one or more user parameters are selected from the group consisting of height, weight, gender, age, and demographic information.
8. The method according to claim 1, wherein receiving the one or more user parameters includes receiving user input of the one or more user parameters via a user device.
9. The method according to claim 1, wherein receiving the one or more user parameters includes receiving measurements performed by a user device.
10. The method according to claim 1, wherein the at least one image is selected from the group consisting of a front view image of the person and a side view image of the person.
11. The method according to claim 10, wherein the at least one image further includes an additional image of the person taken at an angle of 45 degrees relative to the front view image of the person.
12. The method according to claim 1, wherein performing the body segmentation on the at least one image further includes receiving user input to improve the accuracy of the body segmentation, and wherein the user input includes user selection of one or more parts of the body features, the one or more parts corresponding to a given region of the person's body.
13. The method according to claim 1, wherein the at least one image includes at least one image of a user dressed in full or partially dressed, and wherein generating the body feature measurement results further includes generating the body feature measurement results based on the at least one image of the user dressed in full or the partially dressed user.
14. The method according to claim 1, wherein the body size measurement results include first body size measurement results, wherein the method further includes using a second size-setting machine learning module to generate second body size measurement results, and wherein the accuracy of the second body size measurement results is higher than the accuracy of the first body size measurement results.
15. The method according to claim 1, the method further includes: determining whether a given body feature measurement result of the body features corresponds to a confidence level lower than a predetermined value; and in response to determining that the given body feature measurement result corresponds to a confidence level lower than the predetermined value, performing 3D model matching on the body features using a 3D model matching module to determine a matching 3D model of the person, wherein one or more high-confidence body feature measurement results are used to guide the 3D model matching module, performing body feature measurement based on the matching 3D model, and replacing the given body feature measurement result with the predicted body feature measurement result from the matching 3D model.
16. The method according to claim 1, the method further includes: determining whether a given body feature measurement result of the body features corresponds to a confidence level lower than a predetermined value; and In response to determining that the given body feature measurement corresponds to a confidence level below the predetermined value, perform bone detection on the body feature using a bone detection module to determine the joint positions of the person, wherein one or more body feature measurements with high confidence are used to guide the bone detection module, perform body feature measurement based on the determined joint positions, and replace the given body feature measurement with the predicted body feature measurement from the bone detection module.
17. The method according to claim 1, the method further comprises: preprocess the at least one image of the person and the background before performing the body segmentation.
18. The method according to claim 17, wherein the preprocessing at least comprises perspective correction of the at least one image.
19. The method according to claim 18, wherein the perspective correction is selected from the group consisting of perspective correction using the head of the person, perspective correction using a gyroscope of the user device, and perspective correction using another sensor of the user device.
20. The method according to claim 1, wherein the step of identifying one or more body features comprises: generate a segmentation map of the body features on the person; and crop the one or more identified body features from the person and the background before performing the body feature annotation step, wherein the performing the body feature annotation step utilizes a plurality of annotation deep learning networks that have been trained separately with each body feature.
21. A computer program product for generating body size measurements of a person, the computer program product comprising a non-transitory computer-readable storage medium having program instructions recorded therein, the program instructions being executable by a processor to cause the processor to: receive one or more user parameters; receive at least one image, the at least one image containing the person wearing clothes and a background; perform body segmentation on the at least one image to identify one or more body features associated with the person from the background, the body segmentation identifying the one or more body features including one or more body parts from under the clothes; perform body feature annotation on the extracted body features to generate annotation lines corresponding to body feature measurements on each body feature, the body feature annotation annotating at least one body feature of the person under the clothes on the at least one image, the body feature annotation utilizing an annotation deep learning network trained with annotation training data, wherein the annotation training data includes one or more images of one or more sample body features and annotation lines for each body feature; use a sizing machine learning module based on the annotated body features and the one or more user parameters to generate body feature measurements of the person under the clothes from the one or more annotated body features; and generate a body size measurement by aggregating the body feature measurements of each body feature.
22. The computer program product according to claim 21, wherein the body segmentation utilizes a segmentation deep learning network trained with used segmentation training data, and wherein the segmentation training data includes one or more images of one or more sample persons and body feature segmentations of each body feature of the one or more sample persons, and wherein the segmentation training data includes manually measured data, 3D body size measurements from 3D body scans, and / or segmentation training data automatically generated by one or more algorithms.
23. The computer program product according to claim 22, wherein the body feature segmentation is extracted under the clothing, and wherein the segmentation training data includes body segmentations by annotators estimating the body of the person under the clothing.
24. The computer program product according to claim 21, wherein the annotation lines on each body feature include one or more line segments corresponding to a given body feature measurement, and wherein the body feature measurement is generated from the one or more annotated body features using the annotation lines on each body feature.
25. The computer program product according to claim 21, wherein the at least one image includes at least a front view image and a side view image of the person, and wherein the program instructions further cause the processor to perform the following operations: calculate at least one perimeter of at least one annotated body feature using the front view image and the side view image with line annotations and the height of the person; and generate the body feature measurement from the at least one perimeter, the height, and the one or more user parameters using the size setting machine learning module based on the at least one perimeter, the height, and the one or more user parameters.
26. The computer program product according to claim 21, wherein the size setting machine learning module includes a random forest algorithm, and wherein the size setting machine learning module is trained with real data, the real data including one or more sample body size measurements of one or more sample persons.
27. The computer program product according to claim 21, wherein the one or more user parameters are selected from the group consisting of height, weight, gender, age, and demographic information.
28. The computer program product according to claim 21, wherein the program code for receiving the one or more user parameters includes program code for receiving user input of the one or more user parameters via a user device.
29. The computer program product according to claim 21, wherein the program code for receiving the one or more user parameters includes program code for receiving measurements performed by a user device.
30. The computer program product according to claim 21, wherein the at least one image is selected from the group consisting of a front view image of the person and a side view image of the person.
31. The computer program product according to claim 30, wherein the at least one image further includes an additional image of the person taken at an angle of 45 degrees relative to the front view image of the person.
32. The computer program product according to claim 21, wherein performing the body segmentation on the at least one image further comprises receiving user input to improve the accuracy of the body segmentation, and wherein the user input comprises user selection of one or more portions of the body feature, the one or more portions corresponding to a given region of the person's body.
33. The computer program product according to claim 21, wherein the at least one image comprises at least one image of a user dressed in full or a user dressed partially, and wherein generating the body feature measurement results further comprises generating the body feature measurement results based on the at least one image of the user dressed in full or the user dressed partially.
34. The computer program product according to claim 21, wherein the body size measurement results comprise first body size measurement results, wherein the program code further comprises using a second size-setting machine learning module to generate second body size measurement results, and wherein the accuracy of the second body size measurement results is higher than the accuracy of the first body size measurement results.
35. The computer program product according to claim 21, the computer program product further comprising program code for performing the following operations: determining whether a given body feature measurement result of the body feature corresponds to a confidence level lower than a predetermined value; and in response to determining that the given body feature measurement result corresponds to a confidence level lower than the predetermined value, performing 3D model matching on the body feature using a 3D model matching module to determine a matching 3D model of the person, wherein one or more high-confidence body feature measurement results are used to guide the 3D model matching module, performing body feature measurement based on the matching 3D model, and replacing the given body feature measurement result with an expected body feature measurement result from the matching 3D model.
36. The computer program product according to claim 21, the computer program product further comprising program code for performing the following operations: determining whether a given body feature measurement result of the body feature corresponds to a confidence level lower than a predetermined value; and in response to determining that the given body feature measurement result corresponds to a confidence level lower than the predetermined value, performing bone detection on the body feature using a bone detection module to determine the joint positions of the person, wherein one or more high-confidence body feature measurement results are used to guide the bone detection module, performing body feature measurement based on the determined joint positions, and replacing the given body feature measurement result with an expected body feature measurement result from the bone detection module.
37. The computer program product according to claim 21, the computer program product further comprising program code for performing the following operations: preprocessing the at least one image of the person and the background before performing the body segmentation.
38. The computer program product according to claim 37, wherein the preprocessing at least comprises perspective correction of the at least one image.
39. The computer program product according to claim 38, wherein the perspective correction is selected from the group consisting of perspective correction using the perspective of the person's head, perspective correction using a gyroscope of the user device, and perspective correction using another sensor of the user device.
40. The computer program product according to claim 21, wherein the program code includes additional program code for performing the following operations: generating a segmentation map of the body feature on the person; and cropping the one or more identified body features from the person and the background before performing the body feature annotation, wherein the program code for performing the body feature annotation utilizes a plurality of annotation deep learning networks that have been individually trained for each body feature.
Citation Information
Patent Citations
Fast 3D model fitting and anthropometrics
CN107111833A
Calculating bodysize system using two photos
KR1020180007016A
Body profile coding method and apparatus useful for assisting users to select wearing apparel
US20020178061A1