Ai based body measurement system and method thereof
The AI-based body measurement system addresses environmental and error challenges with real-time guidance and robust correction, enabling precise anthropometric measurements and versatile garment sizing across industries.
Patent Information
- Application Number
- PCT/IB2025/060676
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-21
- Filing Date
- 2025-10-20
- Publication Date
- 2026-04-30
AI Technical Summary
Existing body measurement technologies face challenges in capturing accurate, repeatable measurements in uncontrolled environments, lacking real-time user guidance, robust error correction, and broad industrial applicability, particularly in consumer-grade devices.
An AI-based system integrating real-time user guidance, dual-source avatar generation, predictor-corrector models, and automated size chart parsing, utilizing deep learning and classical computer vision techniques for precise anthropometric measurements across diverse environments and privacy constraints.
Ensures accurate, repeatable body measurements with real-time feedback, multi-level error correction, and broad industrial applicability, supporting privacy-preserving execution modes and diverse garment sizing recommendations.
Smart Images

Figure IB2025060676_30042026_PF_FP_ABST
Abstract
Description
Al BASED BODY MEASUREMENT SYSTEM AND METHOD THEREOFFIELD OF THE INVENTION
[0001] The present invention generally relates to anthropometric measurement systems. More specifically, the present invention is directed to an artificial intelligence-based system and method for extracting accurate body measurements from two-dimensional images combined with minimal user-provided data. The system may optionally generate a three- dimensional body model based on extracted measurements and / or input images and / or processed images. The invention further relates to the integration of these measurements into garment size recommendation engines, industrial design processes, and medical or ergonomic applications.BACKGROUND
[0002] Body measurement is a cornerstone for multiple industries including fashion retail, workwear, healthcare, sports science, defense, and ergonomics. Inaccurate or impractical body measurement techniques lead to sizing mismatches, poor ergonomics, elevated product return rates, and potential safety risks. The anthropometric data of a human body such as height, weight, and circumferences of body parts such as neck, limbs, thorax, abdomen, waist, and hips, helps to determine the body characteristics of a person. The anthropometric data does not include the age of a person but it is often recorded alongside body metrics because age can significantly influence body size, shape, and composition. Body parameters are mainly measured as linear measurements and contour measurements. Linear measurements are taken in a straight line, i.e., the distance between two points on the body is taken, e.g., arm length, knee to ankle length, cervical to feet length, etc. In contour measurements, the curves and shapes of the body which are not in straight lines, are measured. Body measurements are crucial for many applications, e.g., industrial workplaces developing assistive ergonomic designs for optimizing workers’ productivity and improving their safety, e-commerce vendors dealing with clothing and apparel products, healthcare professionals accessing nutritional status and diagnosing health conditions, and various industries such as automotive, aerospace, manufacturing, healthcare requiring digital human models for designing new products, improving existing products, testing of safety features, etc. For instance, body measurements are essential in designing medical workwear such as compression socks. Compression socks are effective for people suffering from lymphedema, lipedema, thrombosis and other medicalissues requiring compression therapy, where the correct fit and right pressure of the socks help to improve blood circulation in the body thereby reducing swelling in the legs.
[0003] The technical challenge lies in capturing accurate, repeatable body measurements with user devices (e.g., smartphones) in unsuitable environments (improper lighting, clutter, apparel occlusion, etc.). Particularly, in the case of human wearable products available for sale on digital markets, both e-commerce retailers as well as buyers face challenges in suggesting and selecting the perfect size respectively, which may vary from person to person due to different body characteristics. The clothing or garment manufacturers rely on standard sizing charts and presenting tables that establish a connection between product sizes and body measurements either in inches or centimeters. Consequently, the buyer either needs to go for conventional customized tailoring or dubiously decide on the right size relying on the available information on e-commerce retailer website. Moreover, when a buyer receives a product of the correct mentioned size but an imperfect fit, it becomes an overhead for both the buyer and the e-commerce retailer. As a result, numerous orders end up being incorrect, and the realization of this discrepancy only occurs after the product has been received and tried. This dilemma leads to elevated return rates, subsequently causing financial losses for both the customer and the retailer.
[0004] The workwear industry represents a domain where precise body measurements are of paramount importance. The workwear is often tailored for comfort, safety and functionality to meet the varying demands of different industries and occupations. The fitting of workwear may vary depending on its purpose, need, and the regulations. Ill-fitting workwear can lead to accidents, reduced productivity, and discomfort for workers. By integrating accurate body measurements into the design and production process, workwear manufacturers can improve the overall quality and safety of their products. Furthermore, body measurements play a vital role in designing uniforms and protective gear for military personnel, firefighters, and law enforcement officers, which can withstand demanding conditions. Therefore, accurate body measurements are essential to ensure that these garments fit correctly, providing adequate protection without hindering movement. Also in sports, accurate body measurements are essential for athletes and trainers to track progress, optimize training routines, and select the right equipment. Athletes’ body measurements can directly impact their performance and help in injury prevention. This can be achieved by adopting custom tailoring where accurate body measurements are manually recorded by tailors having certain limitations, e.g., additional time and costs, physical presence required for recording body measurements, etc.
[0005] Interior designers, architects, and product designers also consider body measurements when designing spaces and products. This ensures that furniture, appliances, and spaces are comfortable and functional for users of various body types. Probably one of the most impactful applications is the development and fitting of medical devices, such as prosthetics and orthopedic braces, which require precise body measurements to ensure comfort, mobility, and effectiveness for patients. Furthermore, precise body measurements can be used for optimizing the calibration process of medical machines such as MRI, where the calibration time can be reduced significantly, thus allowing more people to be scanned in the same timeframe.
[0006] Various techniques are available in the field of anthropometry or body measurements extraction, such as pose estimation, semantic segmentation, depth estimation, etc. Document US20210082180 describes a method for remote clothing selection by determining the anthropometric parameters of a user. The document describes the processing of images or 3D scans post-acquisition without ensuring real-time guidance for optimizing image quality. Moreover, the segmentation discussed therein refers to the segmentation of specific body parts (hands, legs, torso, etc.) based on depth maps or 3D scans, and it does not mention custom semantic segmentation of the body contour or detailed contour refinement techniques for improving accuracy on input consisting of 2D images. The document does not address the issues of problematic images and merely focuses on body measurements for general clothing without offering any separate system for hand and foot measurements. This document does not describe methods for parsing size charts from different formats or calculating conversion coefficients between body measurements and clothing sizes. The document does not feature a predictive model or fallback system for adjusting measurements to mitigate errors. It primarily relies on images for generating body models but does not integrate user input data (e.g., height, weight) as an alternative source for avatar creation.
[0007] Document US 10321728 discloses systems and methods for full body measurements extraction using a mobile device camera. The document primarily focuses on determining body measurements using images captured by a mobile phone camera without providing a separate system for foot or hand measurements and does not mention any real-time feedback or virtual assistant to assist users during the extraction process. It focuses on imagebased body measurements but does not support dual-source avatar generation from both images and manual input data. Also, the document does not mention parsing size charts or calculating conversion coefficients for clothing recommendations. The document utilizes deep learningnetworks for body segmentation but does not emphasize the segmentation refinement process or classical image processing techniques for contour correction, lacking any customized semantic segmentation and refinement. The segmentation model employed in the document estimates the body segmentation contour underneath the clothing, which relies on manually annotated data estimating such contours which induces a high degree of subjectivity. Also, the document does not specifically address problematic image handling or provide fallback mechanisms for error correction while processing images. As the document utilizes machine learning for body feature measurement but lacks a predictor-corrector system or statistical fallback to handle errors or incomplete data.
[0008] Document US20190122424 describes a system for generating measurements of a human body using device-captured images and processed fiducial maps of human bodies. The document does not include a virtual assistant or any real-time guidance mechanism during image capture. It relies on users positioning themselves without assistance. The document focuses on general body measurements for generating 3D models but does not provide a separate system for foot and hand measurements or elaborate on custom body part identification techniques. The document uses fiducial maps and statistical models to estimate measurements but it does not describe a segmentation refinement process to improve accuracy in measurement extraction. Further, the best- fit fiducial map approach utilized by the document lacks a predictor-corrector system for adjusting measurements based on additional factors. Also, the document does not describe any specific methods for error detection or handling problematic images during the image capture process. The document is limited in providing a system for identifying best-fit clothing based on 3D body measurements but does not include a detailed method for size chart parsing or coefficient calculation.
[0009] Document US20160110595 describes the body measurements extraction of a user from a single frontal-view depth map using joint location information. Estimates of body measurements are combined with local geometry features around joint locations to form a robust multi-dimensional feature vector. A fast nearest-neighbor search is performed using the feature vector for the user and the feature vectors for the synthetic models to identify the closest match. The retrieved model is used in various applications such as clothes shopping, virtual reality, online gaming, and others. However, the document does not implement a real-time virtual assistant or any system to guide users during the image acquisition phase. Further, the document uses depth sensor data and joint location information to create body models but does not describe a custom segmentation refinement process to improve accuracy. The documentgenerates 3D models from depth maps but does not include mechanisms for problematic image handling or detection of issues during image acquisition. The body measurements are performed using joint location and depth map data without having any separate mechanism for foot or hand measurements. The document relies on synthetic datasets and pre-generated 3D models for measurement estimation but lacks a predictor-corrector model to refine results based on user inputs. It primarily generates 3D avatars from depth sensor data and does not support manual input for avatar creation. The document provides size matching based on 3D models but does not detail any system for parsing size charts or calculating conversion coefficients for clothing.
[0010] Further, Document US 10657709 discloses systems and methods using fiducial maps and silhouette matching against pre-stored libraries for body measurement extraction. The document relies on matching user images to preprocessed databases of fiducial maps and silhouettes, requiring large annotated databases for operation. The system fails in edge cases and lacks real-time user assistance or fallback pathways when the matching process encounters incomplete or noisy data.
[0011] Document US9928412B2 describes model- selection systems that emphasize matching user images to pre-synthesized models based on attributes like height, weight, and gender. While suitable for virtual try-on applications, these systems do not emphasize finegrained measurement extraction and lack robust error correction mechanisms.
[0012] Document US 11393163 discloses dimensional comparison systems that apply weighted functions to compare garment dimensions and user body parameters. These approaches are rigid and often brand-dependent, lacking explainable recommendations or automated chart parsing capabilities.
[0013] Document US 10321728 describes segmentation and annotation systems which uses deep-learning networks to segment body features and apply annotation lines. While this system is capable of producing measurements, they depend on subjective annotation data and lack predictor-corrector fallbacks.
[0014] Document EP3972239 describes virtual fitting systems that focus on 2D garment transfer via DensePose and CE2P segmentation for photo-realistic try-on applications. While these systems achieve visual garment overlays, they do not address accurate measurement extraction or size recommendation based on precise anthropometric data.
[0015] Across all existing techniques, several persistent limitations remain unaddressed. Manual anthropometric measurement is time-consuming, inconsistent, and requires professional expertise. Automated or semi- automated methods using depth sensors or 3D body scanners are accurate but expensive, require additional resources, and unsuitable for large-scale consumer use.
[0016] There is an absence of real-time user guidance mechanisms such as pose correction and environment feedback during image capture. The prior art shows limited handling of problematic inputs including multiple persons in frame, glare, poor lighting, and loose clothing interference. Existing systems lack multi-level fallback systems for measurement confidence and do not provide dual-source avatar generation capabilities using both image data and manual input. The prior art demonstrates weak integration with garment size charts, particularly in handling heterogeneous formats and ease allowance estimation. Additionally, there is insufficient support for privacy-preserving execution modes such as offline edge devices, and limited applicability across multiple industries beyond fashion retail.
[0017] Therefore, in view of the limitations and challenges associated with the existing body measurements extraction techniques, there is a need for an artificial intelligence based accurate body measurements extraction which utilizes user data and two-dimensional images to generate a three-dimensional model with enhanced accuracy, solving the technical problems associated with the prior arts and available knowledge to extract accurate body measurements and recommend best- fit clothing based thereon. The required techniques should further support privacy-preserving execution modes such as offline edge devices and provide broad industrial applicability extending beyond fashion retail to medical, workwear, defense, sports, and ergonomic applications, thereby addressing the persistent gaps in existing anthropometric measurement technologies.OBJECTIVES OF THE INVENTION
[0018] The primary objective of the present invention is to provide an artificial intelligence-based body measurement system that extracts accurate, repeatable anthropometric measurements using consumer-grade devices while addressing uncontrolled capture environments including variable lighting conditions, background clutter, and apparel occlusion through real-time user guidance and environmental feedback mechanisms.
[0019] Another objective of the present invention is to implement a virtual assistant with real-time guidance capabilities that provide pose correction guidance, environmental feedback,and quality scoring mechanisms to ensure optimal image acquisition conditions and reduce user errors during the measurement extraction process.
[0020] Yet another objective of the present invention is to establish a predictor-corrector model architecture that provide multi-level validation and error correction capabilities through statistical models trained on anthropometric datasets, enabling fallback mechanisms and plausibility references for image-based measurements under uncertainty conditions.
[0021] Yet another objective of the present invention is to provide dual-source avatar generation capabilities that enable three-dimensional body model creation from either obtained images or manual input data alone, ensuring system functionality across diverse operational requirements and privacy constraints.
[0022] Yet another objective of the present invention is to deliver custom semantic segmentation with contour refinement by integrating deep learning architectures with classical computer vision techniques including morphological operations, edge detection, and histogram equalization to improve measurement accuracy under challenging acquisition conditions.
[0023] Yet another objective of the present invention is to implement automated size chart parsing and normalization capabilities that ingest heterogeneous size chart formats, compute body-to-garment conversion coefficients, and provide explainable size recommendations with transparent fit scoring and rationale.
[0024] Yet another objective of the present invention is to provide dedicated measurement modules for hand and foot dimensions using planar reference objects, extending the system’s applicability beyond body measurements to footwear and handwear sizing applications including gloves, rings, jewelry, and other hand-worn items.
[0025] Yet another objective of the present invention is to ensure privacy-preserving execution modes including offline edge device deployment and GDPR-compliant data handling for enterprise and medical contexts requiring data isolation and regulatory compliance.
[0026] Yet another objective of the present invention is to support industrial-grade calibration protocols that enable brand- specific tuning through empirical measurement sessions and adaptive learning mechanisms for improved collection- specific accuracy.
[0027] Yet another objective of the present invention is to provide broad industrial applicability extending beyond fashion retail to protective workwear, healthcare applications,sports equipment optimization, defense uniform provisioning, and ergonomic design processes across multiple industry verticals.SUMMARY
[0028] The present invention is directed to a comprehensive Al-based body measurement extraction system that accurately determines body measurements of a human body for recommending suitable clothing sizes across multiple industries. The system integrates advanced artificial intelligence techniques with classical computer vision methods to provide robust, accurate, and privacy-preserving anthropometric measurement capabilities. The invention addresses significant limitations in existing body measurement technologies by providing real-time user guidance, multi-level error correction, dual-source avatar generation, automated size chart integration, and comprehensive industrial applicability extending beyond fashion retail to medical, defense, sports, and ergonomic applications.
[0029] In some embodiments, the system implements a dual-source avatar generation capability, wherein three-dimensional models may be generated from either captured images or manual input data alone. When images are available and of sufficient quality, the system utilizes segmentation masks from frontal and lateral views combined with depth estimation to create image-driven avatars. Alternatively, when image capture is not feasible due to privacy restrictions, poor lighting conditions, or user preference, the system may generate avatars using only manual input parameters such as height, weight, age, and gender through statistical body models trained on anthropometric datasets.
[0030] The system may incorporate a predictor-corrector model architecture that provides multi-level validation and error correction capabilities. The predictor component utilizes statistical models trained on large anthropometric datasets to estimate body measurements based on user metadata, serving as both a fallback mechanism and a plausibility reference for image-based measurements. The corrector component cross-validates extracted measurements against anthropometric ranges, detects anomalies, and applies corrections or substitutions when confidence levels fall below predetermined thresholds.
[0031] The system implements comprehensive confidence scoring throughout the measurement pipeline, with each intermediate output including segmentation masks, keypoint detections, scale factors, and final measurements assigned confidence scores between 0 and 1. These confidence scores propagate through the system using Bayesian fusion methods, enabling generation of measurement values with associated confidence intervals anduncertainty estimates. The system includes a hierarchical fallback system comprising multiple levels of redundancy: when primary measurement extraction methods encounter difficulties, the system invokes redundant geometric computation methods, statistical fit predictors, exemplar matching against validated body shape libraries, skeleton-based proxy estimation, or user-guided recapture workflows.
[0032] In some embodiments, the system incorporates automated size chart parsing and normalization capabilities. The system may ingest size charts from heterogeneous formats including structured digital files, semi- structured documents, and unstructured image sources. Parsing methods may include schema detection algorithms, table recognition, optical character recognition combined with natural language processing, and human-in-the-loop validation for ambiguous cases. Parsed data may be normalized into standardized schemas with unit conversion, tolerance factor adjustment, and brand metadata preservation.
[0033] The system may compute body-to-garment conversion coefficients to address the gap between body dimensions and the garment dimensions listed in size charts required for proper fit. These coefficients may be derived through collection-level regression analysis from calibration sessions, garment- type specific domain knowledge, and style modifiers for intended fit preferences. The coefficients ensure that size recommendations are based on true fit logic rather than raw chart values.
[0034] The system supports multiple execution modes including cloud-based processing with elastic scalability through stateless microservices, edge device deployment utilizing optimized models on specialized edge computing hardware such as but not limited to NVIDIA Jetson devices or mobile devices like tablets, and fully offline operation ensuring complete data isolation with all processing occurring locally. This multi-mode architecture supports privacy- sensitive applications in medical, defense, or enterprise contexts while enabling multiindustry applications extending beyond fashion retail to workwear and personal protective equipment sizing, medical device fitting, sports equipment optimization, defense uniform provisioning, and ergonomic design for furniture and workspace optimization.
[0035] In some embodiments, the system implements explainable recommendation logic that provides transparent rationale behind size suggestions. Instead of opaque recommendations, the system may output detailed explanations including dimensional deviations, fit scores, alternative options, and confidence assessments. This transparency may improve user trust and adoption while enabling informed decision-making.BRIEF DESCRIPTION OF DRAWINGS
[0036] The present invention will be better understood after reading the following detailed description of the presently preferred aspects thereof with reference to the appended drawings, in which the features, other aspects and advantages of certain exemplary embodiments of the invention will be more apparent from the accompanying drawing in which:
[0037] Figures 1A-1B illustrate a three-dimensional model based on the input images and body statistics data of a user.
[0038] Figures 2A-2B illustrate the processing of two-dimensional frontal and lateral images.
[0039] Figures 3A-3C illustrate the processing of foot and hand images.
[0040] Figures 4A-B illustrate an example of input images.
[0041] Figures 5A-B illustrate the corresponding segmentation contours extracted from the input images.
[0042] Figures 6A-B illustrate the armpits extracted from the segmentation contours.
[0043] Figures 6C, 6E illustrate the broader chest-armpit-arm region, that can be used for determining the right level for chest measurements.
[0044] Figures 6D, 6F illustrate the thorax contour line direction.
[0045] Figures 7A-B illustrate the elliptical axis for the ankle circumference.
[0046] Figures 7C-D illustrate the elliptical axis for the biceps circumference.
[0047] Figures 7E-F illustrate the minor and major elliptical axis for bottom circumference.
[0048] Figures 7G-H illustrate the minor and major elliptical axis for calf circumference.
[0049] Figures 7I-J illustrate the minor and major elliptical axis for chest circumference.
[0050] Figures 7K-L illustrate the minor and major elliptical axis for forearm circumference.
[0051] Figures 7M, 70 illustrate the arm length.
[0052] Figures 7N, 7P illustrate the outer arm contour.
[0053] Figures 7Q illustrate the leg length.
[0054] Figures 7R-W illustrate the minor and major elliptical axis for lower waist, upper waist and middle waist circumference.
[0055] Figures 7X-Y illustrate the minor and major elliptical axis for wrist circumference.
[0056] Figures 7Z illustrate the elliptical axis for the neck circumference alongside with the general upper and lower limits of the neck.DETAILED DESCRIPTION
[0057] The following description describes various features and functions of the disclosed system with reference to the accompanying figures. In the figures, similar symbols identify similar components, unless context dictates otherwise. The illustrative aspects described herein are not meant to be limiting. It may be readily understood that certain aspects of the disclosed system can be arranged and combined in a wide variety of different configurations, all of which have not been contemplated herein.
[0058] Accordingly, those of ordinary skill in the art will recognize that various changes and modifications of the embodiments described herein can be made without departing from the scope of the invention. In addition, descriptions of well-known functions and constructions are omitted for clarity and conciseness.
[0059] Features that are described and / or illustrated with respect to one embodiment may be used in the same way or in a similar way in one or more other embodiments and / or in combination with or instead of the features of the other embodiments.
[0060] The terms and words used in the following description are not limited to the bibliographical meanings, but, are merely used to enable a clear and consistent understanding of the invention. Accordingly, it should be apparent to those skilled in the art that the following description of exemplary embodiments of the present invention are provided for illustrative purposes only and not to limit the invention.
[0061] It is to be understood that the singular forms “a”, “an” and “the” include plural referents unless the context clearly dictates otherwise.
[0062] It should be emphasized that the term “comprises / comprising” when used in this specification is taken to specify the presence of stated features, integers, steps or components but does not preclude the presence or addition of one or more other features, steps, components or groups thereof.
[0063] The artificial intelligence based body measurements of a human body may be extracted using an ensemble model containing Computer Vision (CV) techniques, Deep Learning (DL) models, classic Machine Learning (ML) and Statistical Methods, which extract body measurements with an error of maximum 5 mms, once all the optimal measuring conditions are met. The body measurements are extracted only by using some basic user information i.e., weight, height, age, and gender, and by capturing frontal and lateral images. It may be appreciated that an ensemble model combines several individual models to produce more accurate predictions.
[0064] The ensemble model consists of two principal solutions i.e., one based on Machine Learning and Statistical Methods, and the other based on Deep Learning and Computer Vision techniques. The Machine Learning and Statistical Methods solution consists of a double-input Neural Network, that is fed raw data and a Multivariate Gaussian’s predictions on the raw data and that outputs an array of more than 50 predicted body dimensions. This solution does not require pictures and the predictions are mean to serve as anomaly detectors, correctors, and failsafe for the Computer Vision solution. Moreover, this solution has been deployed as a standalone fit predictor - sizing recommender based on predicted body measurements.
[0065] The Computer Vision solution consists of a complex combination of different models, one for pose estimation, one for segmentation, one for depth estimation and another one for extracting the actual dimensions based on the keypoints, body binary mask, weight, and height. The pose estimation model may be based on PoseNet or MediaPipe, developed by Google as open-source frameworks. The semantic segmentation model may have U2-Net as a backbone, trained specifically on a dataset containing only images with humans. The depthestimation model may be trained on a dataset containing images with rooms, thus predictions on a close picture with a person would have a relatively high accuracy. The final model may receive two original images as input, the two resulted masks, the two depth-estimated images, the key -points for both images, height, and weight. The output is an array of over 100 body dimensions. In one embodiment, a 3D model of the user may be generated using segmentation masks. In another embodiment, a 3D model may be generated directly from the manual input of the user.
[0066] In some embodiments, the present invention discloses a comprehensive Al-driven system for extracting human body measurements and generating corresponding three- dimensional avatars, capable of providing accurate garment size recommendations. The system comprises a plurality of components consolidated to perform a continuous and streamlinedprocess. In some embodiments, the plurality of components comprises at least a Virtual Assistant, a measurement extraction engine, a predictor-corrector model, a three-dimensional (3D) avatar generation model, a garment size recommendation engine, and a plurality of specialized submodules.
[0067] The Virtual Assistant is a software-driven module running on client devices (smartphone, tablet, or desktop), providing real-time interactive guidance to ensure proper user pose, framing, and environment through real-time feedback during image capture. The Virtual Assistant is implemented as a machine learning model for detecting key body points and enforcing predefined pose templates optimized for anthropometric analysis, addressing one of the most significant shortcomings of prior art such as poor capture quality and uncontrolled environments. The visual assistant enforces quality standards before data capture and operates directly on the client device, guaranteeing that images sent to the processing stack meet minimum quality thresholds.
[0068] The virtual assistant may continuously evaluate environmental conditions through multiple validation mechanisms. Lighting analysis may utilize histogram and dynamic range analysis to detect overexposure, underexposure, or glare conditions, with prompts such as “increase lighting” or “avoid backlight” being issued when problematic conditions are detected. Background clutter detection may employ object detection models to check for extraneous items, advising users to reposition when clutter is excessive. Multi-person detection algorithms may ensure a single primary subject, blocking capture with guidance when multiple individuals appear in view. Mirror and reflective surface detection algorithms may identify duplicate body contours caused by reflections, prompting repositioning to avoid measurement errors.
[0069] The virtual assistant may implement quality scoring mechanisms that assign a quality score to every frame based on multiple metrics. The pose alignment score may measure deviation of detected keypoints from predefined templates, while the sharpness score may utilize Laplacian variance to detect blur conditions. The exposure score may evaluate the balance of pixel intensities across the histogram to identify overexposure or underexposure, and the centering score may assess bounding-box position relative to the frame center. The composite quality score may be compared against a predefined threshold, with only frames above the threshold being eligible for automatic capture or user-triggered capture, ensuring that suboptimal images are filtered out before processing.
[0070] The virtual assistant may implement a structured capture workflow that proceeds through multiple sequential stages. The initialization stage may involve user entry of metadata including height, weight, age, and gender, while the guidance stage may display silhouette overlays and instruct users to assume proper poses. The validation stage may check environment and pose conditions in real-time, followed by a quality scoring stage that continuously evaluates capture readiness. When all conditions are satisfied, the capture stage may automatically capture images or enable manual override, with an optional confirmation stage where users may view captured images with acceptance or retry options.
[0071] The virtual assistant may enforce multi- view protocols that ensure correct sequence of capture views for reliable measurement synthesis. The system may require a minimum front- to-side sequence as the basic requirement for reliable circumference synthesis, with optional oblique views captured last when enabled in advanced configurations. Hand and foot capture may utilize separate widget sessions with planar reference alignment protocols. The virtual assistant may prevent incomplete sequences by blocking premature submission until all required views are acquired, ensuring that the measurement extraction engine receives complete datasets for processing.
[0072] The virtual assistant may implement error handling and recovery mechanisms that classify suboptimal conditions into recoverable and irrecoverable categories. Recoverable errors may include poor lighting, background clutter, incorrect pose, or minor occlusion, which may trigger corrective prompts such as repositioning guidance or lighting adjustment instructions. Irrecoverable errors may include multiple people centered in frame, severe occlusion, or incomplete subject visibility, which may trigger rejection and mandatory recapture workflows. The system may provide context-specific prompts including distance adjustment instructions, pose correction guidance, visibility requirements, and lighting improvement suggestions to guide users toward successful capture. Non-limiting examples of prompts include: “Step further back until your entire body is visible”, “Raise your left hand to match the overlay”, “Ensure no one else is visible in the frame”, “Improve lighting before continuing”.
[0073] The virtual assistant may enforce security and privacy measures at the capture stage through secure local handling of raw data. Raw images may be stored only in volatile memory until transfer completion, with transmission occurring via encrypted channels such as but not limited to TLS 1.3 or higher protocols. User consent may be explicitly requested before capture begins, ensuring compliance with privacy regulations. In all types of deployments includingonline and offline modes, derived images that contain no personal data or data that may be used to identify a person may be stored, such as segmentation maps, contour images, depth maps, and other non-identifiable derivatives. In offline deployments, raw images may be processed immediately and discarded thereafter, retaining only derived measurement data to minimize privacy exposure while maintaining system functionality.
[0074] The virtual assistant may provide multiple advantages by ensuring that only high- quality, standardized inputs reach the measurement extraction engine. The system may reduce measurement errors caused by poor lighting conditions or incorrect poses, while lowering rejection and recapture rates compared to unguided measurement methods. The virtual assistant may improve user experience by providing step-by-step support for non-expert users, enabling successful measurement capture without specialized training. The system may demonstrate greater robustness to environmental variability across consumer and industrial use cases, ensuring consistent performance across diverse deployment scenarios.
[0075] In some embodiments, the present invention implements a measurement extraction engine. The measurement engine may implement multiple segmentation network architectures to accommodate varying computational resources and accuracy requirements. The measurement engine is a processing pipeline combining deep learning models such as pose estimation, segmentation, depth estimation, with classical image processing refinements such as edge detection, morphological operations, histogram equalization. The measurement engine extracts circumferences, lengths, and contours from frontal and lateral images and correlates them with user-provided inputs such as height, weight, age, gender.
[0076] Semantic segmentation and contour refinement may represent critical stages of the measurement extraction engine that define the precise body outline required for accurate measurement extraction. Unlike existing approaches that rely solely on raw segmentation outputs, the present invention may implement a hybrid refinement process that leverages both modem deep learning architectures and classical image processing techniques to ensure robustness against clothing occlusion, environmental noise, and poor lighting conditions. The hybrid approach may provide superior accuracy and reliability compared to systems that depend exclusively on neural network outputs without post-processing refinement.
[0077] The measurement extraction engine may implement custom semantic segmentation using deep learning architectures. The segmentation models may be trained on curated datasets with clothing variability, environmental diversity, and pose variations to ensure robustperformance across diverse capture conditions. The segmentation module implements multiple deep neural network architectures trained specifically for human body parsing, including Convolutional Neural Networks (U2-Net, UNet++, DeepLabV3+), Transformer-Based Models (Vision Transformers, Swin Transformer, SegFormer), and Hybrid Models combining CNN encoders with transformer decoders. The invention supports architecture pluralism to allow adaptation to available datasets, computational resources, and industrial needs. Training data includes clothing variability (loose garments, sportswear, uniforms, PPE), environmental variability (indoor / outdoor settings, various lighting conditions), and pose variability to ensure robust performance across diverse capture conditions.
[0078] The segmentation models may be trained on curated datasets comprising full-body human images with comprehensive augmentation strategies. Training data may include clothing variability encompassing loose garments, sportswear, uniforms, and personal protective equipment to ensure robust performance across diverse apparel types. Environmental variability may incorporate indoor and outdoor settings, cluttered and plain backgrounds, and various lighting conditions including low light scenarios. Pose variability may include standing positions, slight rotations, and minor limb deviations to accommodate natural user positioning variations during capture sessions.
[0079] The training labels include precise human contours and body part divisions such as torso, limbs, head. The manual annotation quality control ensures contour accuracy, particularly around occlusion zones such as the underarm or between legs. Also, optional semisupervised training may also be applied, leveraging large unlabeled datasets combined with a smaller annotated subset. Active learning prioritizes uncertain cases for human annotation.
[0080] The segmentation stage may produce multiple types of outputs that are passed downstream for contour analysis and refinement. Binary body masks may provide foreground human versus background classification, while part-level masks may offer optional outputs for specific body regions including head, torso, arms, and legs. Soft probability maps may generate pixel-level confidence scores that are utilized in later refinement stages to identify areas of uncertainty and guide correction processes. These diverse output formats may enable flexible downstream processing and improve the robustness of subsequent measurement extraction steps.
[0081] The system applies contour refinement techniques combining deep learning outputs with classical image processing methods including but not limited to morphologicaloperations, edge detection using Canny or Sobel operators, histogram equalization, Fourier transform analysis, and color space conversion. These refinement techniques act as predictorcorrector layers to adjust contours where segmentation models exhibit uncertainty or boundary inaccuracies. The refinement process addresses challenging scenarios including loose clothing compensation using statistical priors, hair and accessory handling through anomaly detection, background clutter mitigation via edge consistency checks, and mirror / reflection handling through shape redundancy detection.
[0082] While deep learning segmentation may provide strong baseline results, fine details may often require correction through classical computer vision techniques. The refinement pipeline may apply morphological operations to close small gaps or smooth irregularities, edge detection using Canny or Sobel operators to sharpen boundaries in regions of uncertainty, histogram equalization to improve contrast in low-light captures, Fourier transform analysis to detect and correct oscillatory edge noise, and color space conversion using HSV or LAB formats to normalize background-foreground contrast. These classical techniques may act as predictor-corrector layers, adjusting contours where segmentation models exhibit uncertainty and ensuring robust performance across diverse capture conditions.
[0083] The contour refinement process may address challenging scenarios through specialized handling mechanisms. Loose clothing compensation may utilize statistical priors on human body shape to suppress large outward deviations caused by baggy garments, with under-clothing priors trained on datasets containing paired clothed and unclothed silhouettes that have been ethically sourced and properly anonymized to protect individual privacy. Hair and accessory handling may detect anomalies such as long hair or hats and exclude them from contour propagation to prevent measurement distortion. Background clutter mitigation may employ edge consistency checks to filter out spurious contours overlapping with background objects. Mirror and reflection handling may identify duplicate contours through shape redundancy detection and exclude them from measurement calculations.
[0084] In some embodiments, the system may apply under-clothing priors comprising learned shape models that approximate true body contours beneath clothing layers. These priors may reduce errors introduced by loose outer garments by providing statistical constraints on plausible body shapes. The under-clothing priors may be trained on datasets containing paired clothed and unclothed silhouettes that have been ethically sourced and properly anonymized to protect individual privacy, enabling the system to estimate underlying body contours even when obscured by baggy or loose-fitting garments. This capability may beparticularly valuable for accurate measurement extraction in scenarios involving heavy clothing or protective equipment.
[0085] Every contour may be associated with a confidence metric based on multiple validation sources including neural network softmax output probabilities, agreement between binary mask and edge detector boundaries, and plausibility relative to anthropometric priors. The confidence propagation system may enable the identification of low-confidence segments that trigger correction workflows or fallback systems to ensure measurement reliability. Confidence scores may be propagated through the processing pipeline to provide transparency regarding measurement uncertainty and enable informed decision-making in downstream applications such as size recommendation and avatar generation.
[0086] The combination of deep learning and classical refinement may provide multiple advantages including superior robustness across diverse environments and apparel types, finegrained accuracy in high-variance areas such as shoulders, hips, and ankles, reduced false positives from clutter and reflections, and enhanced explainability as corrections can be traced to explicit image processing operations. The hybrid refinement process may demonstrate improved performance compared to purely neural network-based approaches, particularly in challenging scenarios involving poor lighting, complex backgrounds, or unusual clothing configurations. This comprehensive approach may ensure reliable measurement extraction across the wide range of conditions encountered in real-world deployment scenarios.
[0087] The system may implement a predictor-corrector model that provides multi-level validation and error correction capabilities even when some measurements remain unreliable due to poor capture conditions, clothing interference, or missing inputs. The predictor-corrector fallback hierarchy may ensure that a complete set of anthropometric values is always generated, even when some sources are compromised. This dual-track architecture may address scenarios where geometric fallback is insufficient by invoking statistical models and exemplar matching to maintain measurement continuity and accuracy across diverse operational conditions.
[0088] The system may implement redundant geometry re-estimation as the first level of fallback through redundant geometric computation methods. When ellipse fitting fails, contour integration methods may be applied as alternatives, while alternative views may be used when keypoint detection fails on one view. Partial estimates may be extrapolated from available views when one image is corrupted, and this redundancy may ensure resilience to single -point failures in the measurement extraction process.
[0089] When geometric fallback is insufficient, the system may invoke a statistical fit predictor implemented as a machine learning regression model such as neural networks or Gaussian processes. The predictor may utilize user metadata including height, weight, gender, and age, and available partial measurements as inputs to generate predicted values as output for missing or unreliable measurements. This predictor may be trained on large anthropometric datasets and calibrated against population distributions, serving as both a backup mechanism and a plausibility reference for anomaly detection in measurement validation.
[0090] The system may include an exemplar matching module for additional robustness that maintains a library of validated body shapes and measurement profiles. For each new user, the system may identify nearest-neighbor exemplars using feature vectors derived from partial measurements and use exemplar values to fill gaps or validate statistical predictions. This approach may be particularly effective for outlier body types not well represented by regression models, ensuring comprehensive coverage across diverse anthropometric variations.
[0091] When limb or joint measurements are missing due to occlusion, the system may apply skeleton-based proxy estimation using skeletal proportions derived from anthropometric priors. Proxies may be scaled to user-specific parameters such as total height and limb span, with skeletal proportions including femur-to-tibia ratio and humerus -to -radius ratio being applied to ensure measurement continuity. This approach may ensure reliable operation even when lower-limb or arm segments are poorly captured during the image acquisition process.
[0092] The system may implement capture retry logic when confidence remains below threshold despite fallback methods, escalating to user-guided recapture with context-aware guidance. The system may provide specific instructions based on the failed measurement, such as “Please retake the side photo, ensuring your full body is visible” for full body visibility or adjusting lighting when knees are not clearly visible. Also, context-aware guidance for specific instructions based on the failed measurement may be provided, such as “Your knees are not clearly visible, adjust lighting”. This user involvement in final recovery may ensure successful measurement completion where automated correction proves insufficient while maintaining user engagement in the process.
[0093] The fallback system may be structured as a hierarchical decision flow proceeding through multiple levels starting from a primary path direct extraction from segmentation, keypoints, and scale, followed by five stages, namely Level 1 Fallback to Level 5 Fallback. Level 1 Fallback includes redundant geometric methods, Level 2 Fallback includes statisticalfit predictor, Level 3 Fallback includes exemplar matching, Level 4 Fallback includes skeletonbased proxies, and Level 5 Fallback includes user recapture request as successive fallback levels. At each stage, confidence scores may be updated and results tagged with the fallback level used, ensuring traceability and transparency in the measurement derivation process.
[0094] The predictor-corrector system may provide multiple advantages including measurement continuity that ensures a complete set of measurements regardless of input quality, accuracy preservation through statistical and exemplar validation to prevent error propagation, user-centric recovery that invokes recapture only as a last resort to reduce friction, and explainability where each measurement can be traced back to the fallback pathway used. This comprehensive approach may improve transparency and reliability while maintaining system robustness across diverse capture conditions and user scenarios.
[0095] The system may further implement a dual-source 3D avatar generation capability that creates watertight digital models of users through image-driven reconstruction from segmentation masks and depth cues, or input-only avatar generation leveraging metadata such as height, weight, age, and gender when reliable image data cannot be obtained. While measurements provide numerical anthropometric values, many industrial, retail, and healthcare applications may require a visual 3D representation of the user’s body. This dual-source approach may ensure versatility and robustness across diverse use cases by providing multiple pathways for avatar creation.
[0096] The system may fit a parametric body model such as SMPL-like mesh to the extracted measurements through a structured process involving (i) initialization where the base mesh is scaled using user height and weight, (ii) constraint application where landmark positions including shoulders, hips, knees, and ankles are mapped to corresponding vertices, (iii) optimization through iterative fitting using gradient descent to minimize error between extracted measurements and model dimensions, and (iv) regularization where anthropometric priors ensure anatomically plausible body shapes while preventing distortions.
[0097] When images are available, the system may perform image-driven reconstruction by: (i) projecting segmentation masks from frontal and lateral views onto 3D space; (ii) performing depth estimation by utilizing monocular depth networks to provide relative depth cues; (iii) applying multi-view fusion to intersect silhouettes and approximate body volume; and (iv) performing mesh fitting to align the parametric body model with the fused volumewhile refining it with extracted measurements. This process may yield a watertight, metrically accurate 3D mesh that accurately represents the user’s body geometry.
[0098] In scenarios where image capture is not possible due to privacy restrictions, incomplete data, or enterprise preferences, the system may generate a 3D avatar using only manual input data including height, weight, age, and gender. This process may rely on statistical body models trained on large anthropometric datasets, predictor-corrector mechanisms to ensure that generated avatars conform to population-level distributions, and customization options where users may input preferences such as fitness level to refine the model. Although less personalized than image-driven avatars, input-only avatars may remain useful for baseline fitting and ergonomic simulations.
[0099] The system may implement dual-source mode as a key novelty that integrates both image-driven and input-only approaches: (i) if image-based results are high confidence the primary avatar is image-driven; (ii) if image quality is poor the system falls back to input-only avatar generation; and (iii) if both sources are available the outputs can be cross-validated for consistency. This dual-source capability may ensure that the system can always deliver a usable 3D model regardless of input conditions or constraints.
[0100] Further, steps of mesh outputs and exports are performed. Generated avatars may include watertight mesh suitable for visualization and simulation, landmarked mesh with annotated vertices for key anthropometric points, and optional surface texturing with neutral skin-tone textures for visualization purposes. Export formats may include OBJ, FBX, and GLTF formats, enabling integration into CAD tools, apparel design software, virtual try-on platforms, and gaming engines to support diverse industrial applications and workflows.
[0101] The 3D avatar generation system may support multiple industrial applications including apparel applications for virtual try-on, size validation, and return reduction, medical applications for prosthetic fitting, compression garment tailoring, and rehabilitation monitoring, workwear and PPE applications for ergonomic uniform design and safety equipment fitting, sports applications for performance gear optimization and biomechanics analysis, and defense applications for rapid provisioning of uniforms and mission-specific gear.
[0102] The dual-source 3D avatar generation system may provide multiple advantages including versatility that operates in both image-rich and image-limited environments, accuracy that combines geometric fitting with statistical priors, continuity that ensures an avatar is always generated even under suboptimal conditions, and compatibility where meshexports integrate with industry-standard 3D software. This comprehensive approach may enable reliable avatar generation across diverse deployment scenarios while maintaining high quality and usability standards.
[0103] The system may further incorporate a garment size recommendation engine comprising a software component capable of parsing heterogeneous size charts from formats including but not limited to CSV, PDF, JSON, and image sources, normalizing them into a consistent schema, and computing body-to-garment conversion coefficients. The recommendation engine may produce ranked size recommendations with fit scores, explanations, and adjustments for style preferences including tight, regular, and relaxed fits. This comprehensive approach may transform body measurements and normalized garment size charts into actionable outputs that provide accurate and explainable size recommendations across diverse garment categories and brands.
[0104] Garment size charts may represent a cornerstone of apparel fitting, yet they vary widely in format, structure, and semantics across brands and collections. The system may introduce a unified parsing and normalization system that ingests heterogeneous size charts, standardizes them into a consistent schema, and links them with body measurements. This step may be essential for accurate, explainable, and scalable size recommendations across diverse retail and industrial applications.
[0105] Size charts may be provided in numerous input formats including: (i) structured digital files such as CSV, Excel, JSON, and XML; (ii) semi- structured documents such as PDFs and HTML tables; (iii) unstructured sources such as images of printed size charts and catalog scans; and (iv) hybrid formats including e-commerce APIs returning partial sizing data. The system may employ both automated and semi- automated ingestion pathways depending on format complexity and data structure.
[0106] The system may implement multiple parsing methods including structured formats such as schema detection algorithms, semi-structured documents, unstructured images, and Human-in-the-Loop formats. The schema detection algorithms that automatically map known fields for structured formats, table recognition and key-value extraction algorithms for processing PDFs and HTML documents, optical character recognition combined with natural language processing for extracting measurement terms and units from unstructured images, and human-in-the-loop validation for ambiguous cases with results cached for future automation.
[0107] Further, in normalization pipeline step, extracted data may be normalized into a standardized schema including: (i) garment type classification such as shirt, jacket, trousers, dress, footwear, gloves; (ii) measurement field standardization encompassing a comprehensive range of body and garment dimensions including but not limited to chest circumference, waist circumference, hip circumference, inseam, sleeve length, and various other anthropometric and garment- specific measurements; (iii) unit conversion to metric with inch equivalents preserved; (iv) tolerance and ease factor adjustments based on garment type and style; and (v) brand metadata preservation including brand name, collection identifier, season, and regional sizing systems such as US, EU, JP, etc., sizing systems.
[0108] The system may address the gap between body dimensions and garment dimensions by computing body-to-garment coefficients through: (i) collection-level coefficients derived by regression from calibration sessions; (ii) garment-type coefficients derived from domain knowledge such as trousers require waist allowance +4-6 cm; (iii) and style modifiers adjusted for intended fit preferences such as slim fit vs. relaxed fit. These coefficients may ensure that recommendations are based on true fit logic rather than raw chart values, accounting for design ease and manufacturing tolerances. For example, a jacket labeled "chest 100 cm" may actually fit a person with a 92 cm chest due to design ease.
[0109] To prevent errors in parsed data, the system may apply unit consistency checks to detect anomalies, range checks to ensure values fall within garment-type norms, cross-field validation to confirm proportionality between related fields, and historical cross-validation to compare parsed results against prior versions of the same brand or collection. These validation mechanisms may ensure data integrity and reliability throughout the normalization process.
[0110] Normalized size charts may be stored in a centralized garment database with indexing by brand, collection, and garment type, APIs for querying that enable downstream modules to retrieve normalized charts in real time, reusability features where once a brand’s chart has been parsed subsequent collections require minimal incremental effort, and versioning capabilities that track updates to size charts with version history for audit and rollback purposes.
[0111] The size chart parsing and normalization system may provide multiple advantages including scalability that handles heterogeneous inputs across hundreds of brands and collections, accuracy through computed coefficients bridging the gap between body and garment dimensions, explain- ability that provides traceable mappings from raw chart values tonormalized schema, and integration-ready structured outputs that feed directly into the recommendation engine for seamless processing.
[0112] In some embodiments, the present invention discloses a recommendation engine. The recommendation engine may transform body measurements and normalized garment size charts into actionable outputs through garment size recommendations ranked by fit quality. Unlike prior art that simply maps a single body dimension to a garment size, the system may employ multi-parameter fit scoring, style adjustment, and user preference modeling, resulting in more accurate and explainable recommendations across diverse garment categories and user preferences.
[0113] The engine may compute a fit score representing how well a garment size corresponds to a user’s body measurements through: (i) vector representation of user body dimensions and garment size chart values; (ii) weighted distance functions that aggregate differences between body and garment vectors using weights that vary by garment type and importance of each dimension; (iii) ease adjustment for comfort and style allowances, e.g., +4 cm ease for trousers, +2 cm for jackets).; and (iv) fit score output normalized between 0 for poor fit and 1 for perfect fit with transparent scoring methodology. The fit score may be computed using the formula: FitScore = 1 - (S wi * IBi - Gil) / (S wi * Tolerancei), Where Bi = body dimension, Gi = garment dimension, wi = weight factor, Tolerancei = acceptable deviation for that dimension.
[0114] The ranking and recommendation is determined for each garment by analyzing the candidate sizes may be scored individually and ranked by descending fit score, with sizes exceeding threshold values being ordered by preference rules such as smaller size first for slim fit or larger for relaxed fit. If multiple sizes exceed the fit threshold, the system may provide ranked alternatives, while if no size fits within tolerance, the system may recommend custom sizing or notify that the garment is unsuitable for the user’s measurements.
[0115] Personalization and preference modeling is based on the user preferences which may be explicitly incorporated through fit preference input where users specify desired fit including tight, regular, or relaxed options, preference weights that adjust ease allowances accordingly, learning from history where if users repeatedly return garments or favor certain fits the model adapts weights dynamically, and context-aware adjustment where workwear, sportswear, or medical garments prioritize mobility and safety over style considerations.
[0116] The system may provide explainable recommendation logic that outputs detailed reasoning instead of opaque recommendations, including specific dimensional deviations, fit quality assessments, alternative size options, and transparent rationale such as size recommendations with waist and chest measurements relative to user body dimensions. This transparency may improve user trust and adoption while enabling informed decision-making in garment selection and purchase processes. For example, instead of opaque “Size M recommended,” the engine outputs reasoning: “Size 52 recommended: waist +3 cm, chest +2 cm, hips -1 cm relative to your body.” “This size provides a regular fit according to brand sizing chart.” “Next best alternative: size 54 (waist +5 cm, relaxed fit).”
[0117] The recommendation engine outputs may be consumed by e-commerce platforms through plug-ins that display recommended sizes on product pages, enterprise procurement systems for workwear orders automatically mapped to correct sizes, medical systems for compression garments prescribed with precise sizing, and sports applications where athletes receive tailored equipment recommendations. The system may support multiple integration pathways and API formats for seamless deployment across diverse platforms.
[0118] The present invention performs error handling by providing low-confidence recommendations. If confidence in body measurements or parsed size charts is low, fit score thresholds may be lowered and recommendations flagged as tentative, with the system suggesting manual confirmation for higher accuracy, e.g., “Please verify your waist measurement for higher accuracy”, or providing fallback defaults in enterprise contexts, e.g., median size from calibration dataset.
[0119] The recommendation engine may provide multiple advantages including high accuracy through multi-parameter fit scoring, flexibility supporting diverse garment categories from casual apparel to PPE and medical garments, personalization incorporating explicit and implicit user preferences, explainability with transparent and traceable fit recommendations, and enterprise readiness with direct integration capabilities for retail, procurement, and medical workflows.
[0120] The system may implement specialized submodules comprising widgets dedicated to extracting hand and foot measurements using planar references such as A4 sheets, ensuring applicability to footwear and glove sizing applications. While most prior art focuses exclusively on torso and limb measurements, the system may extend coverage to foot and hand measurements, addressing footwear and glove sizing which represent critical categories inapparel, sports, healthcare, and PPE industries. Dedicated modules may ensure high accuracy by leveraging planar references and specialized acquisition workflows.
[0121] The system implements capture protocol with planar reference. The foot and hand measurement module may operate through a dedicated widget integrated into the client capture application, where users place their foot or hand on or near a planar object of known dimensions, typically an A4 sheet measuring 210 x 297 mm or similar reference. On-screen rectangular overlays may guide users to align the reference object properly, while computer vision algorithms detect the planar reference and confirm correct alignment and scaling. Once alignment is verified, images may be captured automatically to ensure accuracy, with users also being prompted to take photos manually when needed, providing a consistent scale reference independent of user-reported height.
[0122] Further, feature extraction is performed from the captured images. The system may extract multiple measurements through: (i) contour segmentation networks that isolate the foot or hand region from background; (ii) landmark detection that identifies keypoints such as heel, toe tips, ball of foot for feet and wrist crease, fingertips for hands; (iii) dimension calculation including foot length from heel to toe, width at ball of foot, instep girth, ankle circumference for feet, and palm length, palm breadth, finger lengths, wrist circumference for hands; and (iv) curvature analysis which may capture non-linear features including arch height proxies for feet and finger spread for hands to provide comprehensive measurement coverage.
[0123] The system performs mapping to footwear and glove size charts. Extracted dimensions may be mapped to dedicated size charts including footwear charts for EU, US, UK sizing systems plus brand- specific conversions, glove charts with standardized hand circumference-to-size mappings, and coefficient adjustments for brand and category- specific ease allowances such as sports shoes versus formal shoes or disposable gloves versus protective gloves. When multiple size options are available, the engine may provide ranked recommendations with explanations such as specific size recommendations with alternative relaxed fit options. For example, size 43 EU is recommended if size 44 EU as an alternative relaxed fit.
[0124] The specialized measurement modules may support multiple use cases including medical applications for orthopedic insoles, compression socks, and rehabilitation gloves, sports applications for performance footwear, goalkeeper gloves, and climbing shoes, workwear and PPE applications for safety boots, protective gloves, and chemical-resistantfootwear, and fashion applications for everyday shoes and custom glove tailoring, rings, jewelry, and other hand-worn items. This broad applicability may ensure the system addresses diverse industry requirements across multiple sectors.
[0125] The system may implement comprehensive error handling mechanisms including: (i) misalignment detection where if reference sheets are skewed or partially out of frame the system prompts recapture; (ii) occlusion handling where if toes or fingers are obscured by socks or jewelry capture is rejected; and (iii) confidence scoring where each dimension is tagged with a confidence score enabling fallback estimates via population priors. These error handling capabilities may ensure measurement reliability and accuracy across diverse capture conditions.
[0126] The specialized submodules may provide multiple advantages including extension beyond body measurement to enable footwear and glove recommendations, leveraging reference objects for improved scaling accuracy, supporting multiple industries including fashion, sports, healthcare, and PPE, and integrating seamlessly with the broader measurement and recommendation system. This comprehensive approach may ensure that the system addresses the complete spectrum of sizing requirements across diverse application domains while maintaining consistency with the overall measurement framework.
[0127] In some implementations, the system may support multiple execution modes including cloud-based execution and offline edge-device deployment such as NVIDIA Jetson devices or mobile devices such as tablets, with offline operation ensuring privacy, GDPR compliance, and resilience against connectivity issues. Given that body images and measurements constitute sensitive personal data, the system may incorporate a comprehensive framework for privacy preservation, secure data handling, and GDPR compliance. These measures may ensure trust for end-users, enterprises, and regulators across healthcare, retail, and defense contexts.
[0128] The system implements comprehensive privacy-preserving measures adhering to data minimization principles by capturing only essential data, prioritizing derived representations over raw images, and processing raw images ephemerally with automatic deletion after processing. The system retains only non-identifiable derivatives including segmentation masks, depth maps, measurement arrays, and confidence scores. Comprehensive encryption includes end-to-end encryption with TLS 1.3 for data in transit, AES -256 encryption for data at rest, and rotating key management with hardware security moduleintegration. User rights enforcement includes explicit opt-in consent, right to erasure, right to access, and right to anonymization. For environments requiring strict isolation, the system supports full offline mode with local processing, no internet connectivity, immediate raw image disposal, and encrypted local storage of derivative results only.
[0129] The system may be configured for enterprise and regulatory integration by integrating with enterprise data protection policies, generating compliance logs and audit trails, supporting Data Protection Impact Assessments, and enforcing regional restrictions such as EU-only data processing. The privacy and security framework may provide multiple advantages including GDPR compliance aligned with EU regulations and similar global frameworks, privacy-first design that minimizes storage of identifiable data, user empowerment with full control over consent, access, deletion, and anonymization, and enterprise readiness with audit trails, logging, and offline operation for sensitive deployments.
[0130] In some embodiments, the present invention may implement a calibration and continuous improvement framework for onboarding new brands and collections, including calibration protocols such as measuring users per collection with ground truth try-ons, active learning loops, and brand- specific adapters. Garment fit may vary significantly across brands, collections, and even product lines, and the system may incorporate a structured calibration and onboarding protocol that adapts the measurement-to-size mapping for each collection. This approach may ensure that recommendations achieve consistently high accuracy despite brandspecific variations.
[0131] When a new collection is introduced, the system may provide a calibration day service where participants covering a range of body types and sizes relevant to the collection are recruited, each participant is measured by the body measurement system under optimal conditions, participants try on garments from the new collection with actual fit outcomes recorded, and both measurement outputs and try-on outcomes are stored as calibration data. This process may establish an empirical mapping between body dimensions and garment sizes specific to the collection.
[0132] Calibration data may be processed to compute regression models linking measurement vectors to garment fit outcomes through linear or nonlinear regression that maps body circumferences and lengths to garment size labels, coefficient estimation that refines brand- specific body-to-garment ease coefficients, and fit thresholds that identify boundariesbetween sizes. These mappings may override default size chart interpretations, significantly improving recommendation accuracy for specific collections and brands.
[0133] The system may support brand adapters comprising modular components that encapsulate calibration results for each brand or collection, including collection-level adapters specific to one garment collection, brand-level adapters aggregated across multiple collections for general brand tendencies, and regional variants supporting region- specific sizing systems. Adapters may be stored in the garment database and automatically applied when users select garments from calibrated collections.
[0134] Calibration may not end with the initial session but continue through feedback integration where post-launch user return data is anonymized and fed back into adapters, adaptive learning where machine learning models update coefficients dynamically as new data accumulates, and incremental calibration where for large enterprises calibration may be repeated quarterly to reflect design adjustments or manufacturing changes. This continuous improvement approach may ensure sustained accuracy over time.
[0135] Calibration protocols may integrate seamlessly into enterprise workflows including fashion retailers ensuring accurate sizing for seasonal collections, workwear suppliers validating PPE sizing against safety standards, medical providers confirming compression garment sizing for therapeutic effectiveness, and defense or uniform procurement ensuring mission-critical uniforms fit diverse body types. This broad applicability may ensure the system addresses diverse industry requirements.
[0136] The calibration and continuous improvement framework may provide multiple advantages including collection- specific accuracy that adapts recommendations to unique garment sizing logic, data-driven mapping that uses empirical measurements rather than assumptions, scalability that works for both small collections and large enterprise catalogs, and dynamic updating that continuously improves with user feedback and additional calibration events. This comprehensive approach may ensure sustained high performance across diverse brands and collections while maintaining adaptability to changing requirements.
[0137] In some aspects, the system may implement keypoint detection using machine learning models trained on human keypoint datasets to detect anatomical landmarks including shoulders, hips, knees, ankles, and wrists. Pose estimation networks may include PoseNet, MediaPipe, OpenPose, HRNet, and other architectures, with custom fine-tuning performed on proprietary datasets that emphasize anthropometric consistency to ensure shoulder-to-elbowratios and other body proportions remain plausible. Multi-view keypoint fusion may be applied when both frontal and lateral images are available, with detected keypoints being cross- referenced to resolve ambiguities and improve depth consistency. Each keypoint may be assigned a confidence score and optionally an uncertainty ellipse indicating positional variance, with keypoints falling below confidence thresholds triggering fallback estimation procedures.
[0138] The system may derive anatomical axes from detected landmarks to provide structural framework for measurement extraction. A torso axis may be computed as a line connecting neck and hip centers to provide vertical reference for body alignment. Limb axes may be calculated through shoulder-elbow-wrist and hip-knee-ankle triplets to provide arm and leg orientation references. A body symmetry axis may be computed from bilateral landmarks including shoulders, hips, knees, and ankles to establish median body alignment. Local reference frames may be defined for each body region, with chest frames originating at midpoint estimations aligned with torso and body symmetry axes, waist frames located at narrowest torso cross-sections perpendicular to torso axis, hip frames defined at maximum lateral extension of pelvis, and limb frames established at proximal and distal joint centers.
[0139] In some embodiments, the system may implement multi-view Bayesian fusion for measurement synthesis. When multiple images are captured, the system may apply likelihood functions to generate independent measurement estimates with confidence intervals from each view. Posterior distributions may be computed by weighting estimates according to viewspecific confidence levels, with final measurements derived from the mean of the posterior distribution to mitigate perspective distortions and occlusion errors.
[0140] Scale estimation may be implemented through multiple redundant methods to ensure accurate pixel-to-metric conversion. Height- anchored metric scaling may detect the user’s full vertical extent in images via segmentation masks and keypoints from top of head to sole of feet, calculate pixel distance between head and feet, and map the pixel span to provided height values with scaling factors applied uniformly to all pixel-based distances. Deviceassisted scaling may utilize integrated depth sensors including LiDAR on smartphones or Intel RealSense cameras to provide direct pixel-to-metric scaling through fusion of depth data with segmentation masks. Reference object scaling may employ planar objects of known dimensions such as A4 sheets, credit cards, or calibration marker sheets, with users aligning reference objects with on-screen overlays and scaling factors derived from detected reference object contours.
[0141] Perspective and gyroscope compensation may be applied to mitigate camera- induced distortions. Gyroscope and accelerometer data from mobile devices may be utilized to estimate camera tilt and roll angles, with homography transformations applied to correct perspective distortion. Pose constraints such as vertical torso axis alignment may be used to realign subjects and ensure scaling factors are not skewed by suboptimal device angles. Multisource fusion may combine multiple scaling cues through Bayesian fusion frameworks, with primary scaling from user height, secondary validation from reference object detection, and optional refinement from device depth sensors being weighted according to confidence levels.
[0142] The system may incorporate multiple scaling methods including height-anchored metric scaling using user-provided height as reference, device-assisted scaling utilizing integrated depth sensors, reference object scaling using planar objects of known dimensions, and perspective compensation using gyroscope and accelerometer data. Multi-source fusion may combine scaling cues using Bayesian frameworks weighted by confidence levels to improve scaling accuracy and reliability.
[0143] Circumference estimation may be performed through a hybrid method combining cross-sectional plane identification, contour extraction, ellipse fitting, and circumference calculation. Cross-sectional planes may be identified based on anatomical landmarks and local reference frames, with segmentation contours intersected with the cross-sectional plane to extract relevant contour points. Minimum enclosing ellipses may be fitted to cross-sectional points, with circumferences estimated from ellipse parameters and corrected using postprocessing methods that refine circumference measurements beyond minimum enclosing ellipse accuracy, which may include geometric models, neural networks, regression models, or other computational approaches that account for depth and pose variations. Circumferences computed may include neck, chest, waist, hips, thigh, calf, biceps, forearm, wrist, ankle, and other circumferences.
[0144] Linear and contour measurements may be computed as direct distances between landmarks with appropriate scaling applied. Vertical measurements may include total height from head to feet, torso length from cervical to pelvis, and inseam from groin to ankle. Horizontal measurements may include shoulder width, inter- acromial distance, and pelvic breadth. Limb length measurements may include arm length, forearm length, upper leg length, and lower leg length. Contour-based measurements may capture lengths along curved surfaces including spine curvature length, arm contour length, and leg contour length to provide comprehensive anthropometric coverage.
[0145] In some aspects, the system may generate comprehensive measurement catalogs including core measurements such as height, chest, waist, hips, inseam, and arm length, extended measurements including neck circumference, limb segments, and girths, and derived indices such as body mass index, waist-to-hip ratio, and shoulder-to-hip ratio. Each measurement entry may include values in multiple units, confidence scores, source identification, and correction flags indicating predictor-corrector adjustments.
[0146] Now referring to Figs. 1A-1B, the present invention is directed to a body measurements extraction system to accurately determine body measurements of a human body to generate a three-dimensional model or avatar from two-dimensional images and to recommend a suitable clothing size. The system comprises a virtual assistant, a measurement extraction engine, a recommendation system, and a foot and hand measurements system. The virtual assistant is provided to capture two-dimensional image data including frontal and lateral images, and a manual input including body statistics data such as height, weight, gender, and age, and preferences of a user in real-time. Preferences of the user may depend on the choices such as but not limited to comfort e.g., loose or tight clothing, work, activities, etc. The captured data including the image data and the manual input is forwarded to a remote server for further processing by the measurement extraction engine. The captured image data is identified for any error or problematic background scenes. Further, body contours are extracted using a customized semantic segmentation model, and said body contours are rectified using a customized semantic contour correction model based on computer vision and classic image processing techniques. The automated size chart parsing and coefficient calculation is performed on the rectified body contours, whereas refining and measurement accuracy thereof is ensured by a predictor-corrector model, and a three-dimensional (3D) model or avatar is generated. Based on the generated 3D model, the recommendation system recommends an accurate clothing size to the user. The foot and hand measurements system is configured to extract the foot and hand measurements.
[0147] In some embodiments, the virtual assistant is configured to receive body statistics and preference details of a user. The virtual assistant may also be configured to guide the user to capture at least two optimal images, i.e., one frontal view image and one lateral view image. It may be appreciated that the virtual assistant may be accessed on a customized mobile application or a website. The virtual assistant may be directly accessed using a link or by scanning a code, such as a QR code, thereby enabling a user to use the virtual assistant on web browser of a desktop / laptop computer or smartphone / tablet without having any requirementof logging into one’s personal account. The virtual assistant may be configured to run on a portable computing device equipped with at least a front camera, gyroscope, and wireless internet connectivity, such as but not limited to a smartphone, tablet, etc. The virtual assistant is implemented using a Machine Learning model to identify key body points and capture optimal frontal and lateral two-dimensional (2D) images by precisely comprehending the optimal pose and position of the user. It should be noted that the 2D image is a single image captured by a camera and monocular depth estimation may be applied on each such image to predict the depth information thereof. The two images, i.e., the frontal image and the lateral image of a user are processed independently for their respective depth information. It may be appreciated that 2D image may be a single channeled grayscale image or a multi channeled (e.g., RGB) color image. The key body points information is utilized to guide the user to the correct pose and position using visual and audio instructions e.g., “check your phone inclination”, “make sure you are completely visible on the screen”, “step further away from your phone”, “come closer to your phone”, “raise your left hand”, “bring your legs slightly closer, etc. The images are automatically collected when the virtual assistant allows the data acquisition and are securely sent to a remote server for processing. Thus, the virtual assistant ensures the optimal conditions for extracting the most accurate measurements of the user.
[0148] Referring to Figs. 2A-2B, the body measurements extraction is performed on the front and lateral images and the body statistics data as received on the remote server. A measurement extraction engine may be provided to process said images and data, wherein the engine is customized to accurately determine a perfect fit for the user. In some non-limiting embodiments, the perfect fit determination may include processes of custom semantic segmentation, body keypoint detection, segmentation contour refinement, problematic image analysis and correction, body parts identification on the segmentation contour, computing minor and major axis sizes, and measurements extraction.
[0149] In custom semantic segmentation, the remote server receives the data from the widget, a Machine Learning based custom semantic segmentation model is run on the frontal and lateral images to extract the body contour therefrom. This model may be trained from the beginning on a customized dataset. During the keypoint detection process, the key body points are detected using the Machine Learning model. These are used in combination with the segmentations (body contours) to determine specific body parts for extracting measurements. While implementing the segmentation contour refinement, the segmentation is performed using the Machine Learning model. Sometimes, the contour might have some minor errors,thus, classic image processing techniques, e.g., edge detection, morphological operations, filtering, thresholding, histogram equalization, Fourier Transform, color space conversion, etc., may be implemented to validate and correct the segmentation contour i.e., predictor-corrector.
[0150] The error or problematic image analysis and correction process is necessary for capturing the frontal and lateral views of the user. Besides the segmentation contour correction, information related to problematic image scenes such as background scenes with overexposed light conditions (glares, light spots etc.), multiple people in the image (at least in the center of the images where the user is positioned), dark images, problematic backgrounds (which are harder to segment), the presence of windows and mirrors, loose clothes, big beard / big hair not tied, etc., may be identified. Such information is relevant for extracting the most accurate measurements. In some cases, problematic image scenes (e.g., light / dark spots) may be corrected and measurements may be determined. While in other cases (e.g., multiple people in the image center), the user may be guided to measure themselves again in the appropriate clear background scenes. The body parts identification on the segmentation contour is performed to identify key body points. For each measurement, a specific body part based on the refined segmentation contour and on the key body points is identified. For each body part, a specific custom algorithm may be used based on classic image processing techniques as described above.
[0151] In case if position of a specific body part, e.g., shoulder, chest, groin, etc., is known, minor and major axes sizes are computed. Said axes are the minor and major axes of a minimum enclosing ellipse. The computation is performed for circumferences only. Based on the minor and major axes, a customized Machine Learning model may be implemented to compute the circumference for each specific body part. For non-circumference measurements, classic image processing techniques may be utilized to measure straight line and contour measurements.
[0152] The body measurements are determined using one or more statistical models, implemented as a fit predictor. In some embodiments, the fit predictor may include a statistical model for estimating body measurements and a predictor corrector. The statistical model may be configured to predict a reduced number of body measurements, e.g., 10-20 body measurements instead of 100+ body measurements, based only on the user input such as height, weight, age and gender of the user. The predictor corrector algorithm along with the customized Machine Learning model is implemented to adjust the body measurements by correlating the statistical model with a couple of questions, e.g., a questionnaire related to the fitness data ofthe user. This may help in estimating the correct fit for the user, and said fitting details may further be processed along with the body measurements data to accurately extract body measurements of the user.
[0153] Further, a fallback system is disclosed. In some embodiments, the fallback system includes sanity checks, body shape checks, and body contour integrity checks. Sanity checks are used for obvious erroneous measurements during measurements extraction, e.g., a circumference of 2 meters is recorded for an average person. During sanity checks, both the statistical model and the received data indicating a probable range for each body part for each height, weight, age, gender combination, are used. Body shape checks are used to check the plausibility of the body contour to ensure that the body contour is of a human body and not something else. Body contour integrity checks are used whether the contour is continuous, and in case of any discontinuity such as holes or voids, the same is covered.
[0154] The measurement engine may be configured to generate a three dimensional (3D) digital model or avatar either from the captured images or the fit predictor, or both. A 3D model may be generated from images, wherein a pipeline with an input of segmentation masks as derived from the segmentation contours of the frontal and lateral images extracted and an output of one 3D object e.g., in .obj format. The pipeline consists of a customized Machine Learning model trained on personalized data and algorithms such as Principal Component Analysis (PCA). Further, a 3D model or avatar may be generated from the fit predictor by using user input data such as height, weight, age and gender, and by implementing a customized Machine Learning model trained on personalized data. In some embodiments, the 3D object is generated in .obj format.
[0155] The present invention discloses a recommendation system. In some embodiments, the recommendation system may include a size chart parsing and integration of the size charts into a suitable format. It may be appreciated that size charts for each piece of clothing are required to recommend correct clothing sizes. If the size charts do not exist, they can be created from physical clothes. Depending on the format of the size charts, available parsing methods may be utilized to convert the size charts to a suitable format which may be recommended for use in the present disclosure. It should be noted that the size charts may be in different formats, e.g., excel, csv, pdf, image, Json, etc. For each of such formats, automatic and semiautomatic parsing methods may be used.
[0156] In some embodiments, the recommendation system may be configured to compute body-size chart coefficients. Usually, the size charts represent the physical dimensions of the pieces of clothing and not the dimensions of the body that would fit in them. For example, a jacket with a chest size of 144 cm in the size chart would fit a person with a chest circumference of 124 cm. This happens a lot, thus requiring computation of coefficients which translate from the body size to the size chart size and vice versa. The coefficients are computed using a custom statistics model trained on personalized data. Once the coefficients are calculated, relevant body sizes are compared to the corresponding body sizes in the size chart using a suitable algorithm, thereby recommending the clothing size which fits the best size for the user. In case, if there is no such size, then it notifies the user that it requires a custom size.
[0157] Figs. 3A-3C show extraction of foot and hand measurements. In some embodiments, a widget may be implemented along with other embodiments disclosed hereinabove, wherein the widget is the data acquisition system which prompts the user to take at least one image of the foot or one image of the hand, depending on the user requirement. It should be noted that the widget for foot and hand measurements may be different from the widget for the full body measurement. In an embodiment, the user is required to place his / her foot or hand on or near a plane surface, e.g., A4 paper sheet, which may be used as a reference. The process of capturing the image requires matching the rectangular overlay of the widget with the contour of the plane surface, e.g., A4 paper sheet. Once the matching is correctly performed, the image is captured either automatically or manually by the user on the widget, which is further verified by the remote server. The engine for extracting foot and hand measurements is implemented through the same stages as the body measurement engine. Further, the recommendation system is similar to the one for clothes, however different size charts and coefficients are utilized for foot or hand measurements recommendations.
[0158] Referring to Figs. 3A-3C, the specialized foot and hand measurement modules extract precise dimensions using planar reference objects for scale-independent accuracy. Fig. 3 A illustrates the foot measurement capture setup, wherein a user’s foot is positioned on an A4 paper reference sheet (210 x 297 mm) that provides standardized scale reference for accurate pixel-to-metric conversion independent of camera distance or device specifications for footwear sizing applications. Fig. 3B demonstrates the segmentation processing results showing the refined foot contour after applying the customized semantic segmentation model and image processing refinement techniques, with precise identification of anatomical landmarks including heel, toe tips, ball of foot, and arch regions for extracting key dimensionsincluding foot length, width, and instep girth. Fig. 3C illustrates the hand measurement capture setup with fingers extended and palm positioned on or near the A4 reference sheet, with visible markings on the hand that assist in landmark identification, ensuring clear visibility of anatomical features including fingertips, palm boundaries, and wrist crease for subsequent extraction of palm length, palm breadth, individual finger lengths, and wrist circumference measurements essential for glove, ring, and handwear sizing across medical, industrial, and fashion applications.
[0159] Referring to Figs. 4A-4B, the input image capture process demonstrates the virtual assistant’s real-time guidance system ensuring optimal measurement conditions. Fig. 4A shows a frontal view image with the user positioned in optimal pose with proper alignment and lighting conditions, displaying the full body silhouette against a clear background that meets the quality standards enforced by the virtual assistant for reliable segmentation and keypoint detection. Fig. 4B illustrates the corresponding lateral view image captured in sequence, showing the user’s side profile with clear visibility of body contours and anatomical landmarks necessary for comprehensive measurement extraction, with both images meeting the quality standards enforced by the virtual assistant for reliable processing by the measurement extraction engine.
[0160] Referring to Figs. 5A-5B, the semantic segmentation and contour refinement process demonstrates the hybrid approach combining deep learning with classical image processing techniques. Fig. 5A shows the corresponding segmentation results from processing the frontal view image of Fig. 4A, displaying the precise body contour extracted by the customized semantic segmentation model with clear delineation of the human silhouette from the background, including detailed boundary definition around complex regions such as arms, legs, and torso areas. Fig. 5B illustrates the refined corresponding segmentation contour from the lateral view image of Fig. 4B, showing the side profile outline after applying contour correction techniques using classical methods including morphological operations, edge detection, and histogram equalization. Together, these segmentation masks from Figs. 5A-5B represent the processed output from the raw input images of Figs. 4A-4B, providing the foundation for subsequent body part identification and measurement extraction processes.
[0161] Referring to Figs. 6A-6F, the armpit and chest region analysis demonstrates the system’s sophisticated approach to anatomical landmark identification and measurement plane establishment. Figs. 6A-6B show the precise armpit region extraction from segmentation contours, critical for determining accurate chest measurement levels and avoiding errors causedby arm positioning variations. Figs. 6C and 6E illustrate the comprehensive chest-armpit-arm region analysis that identifies optimal horizontal levels for chest circumference measurements by analyzing anatomical landmarks and contour features. Figs. 6D and 6F demonstrate the thorax contour line direction analysis, showing the computational approach for establishing measurement planes and orientation vectors that ensure consistent circumference calculations across different body poses. This progressive analysis from Figs. 4A-4B through 5A-5B to 6A- 6F demonstrates the sequential processing pipeline from raw image capture to specific body region measurement extraction. This exemplifies the system’s ability to handle complex anatomical regions where traditional measurement methods often fail, supporting applications in medical compression garments, protective workwear, and precision-fit apparel where chest measurements are critical for safety and functionality.
[0162] Referring to Figs. 7A-7Z, the comprehensive measurement extraction techniques demonstrate the system’s capability to extract over different body dimensions using minimum enclosing ellipse fitting combined with post-processing refinement models. The figures display distinctive black and white markings on different body parts of the user, showing how measurement techniques are applied to processed segmentation data. The system may comprehensively address lower extremity measurements including ankle, calf, and leg dimensions, upper extremity measurements encompassing biceps, forearm, wrist, and arm measurements, torso measurements covering chest, hip, and waist circumferences, and neck measurements for collar and medical applications.
[0163] Lower Extremity Measurements: Figs. 7A-7B demonstrate ankle circumference elliptical axis determination, while Figs. 7G-7H show calf circumference analysis with visual indicators of the elliptical fitting process, and Fig. 7Q illustrates leg length measurement methodology.
[0164] Upper Extremity Measurements: Figs. 7C-7D show biceps circumference analysis using minimum enclosing ellipse fitting, Figs. 7K-7L illustrate forearm circumference analysis with markings for left and right measurements, Figs. 7X-7Y demonstrate wrist circumference analysis, and Figs. 7M-7P show arm length and outer arm contour analysis techniques.
[0165] Torso Measurements: Figs. 7I-7J demonstrate chest circumference elliptical axis computation, Figs. 7E-7F illustrate hip circumference measurements for garment fitting, and Figs. 7R-7W provide comprehensive waist analysis including lower, upper, and middle waist circumferences with frontal and lateral view integration using multi-view Bayesian fusion.
[0166] Neck Measurements: Fig. 7Z illustrates neck circumference measurement with multiple reference points ensuring consistent placement across users, essential for collar sizing and medical applications.
[0167] This comprehensive measurement analysis supports diverse industrial applications from fashion retail and workwear to medical device fitting and sports equipment optimization, with each measurement integrated into the predictor-corrector system for validation and fallback processing when confidence levels fall below predetermined thresholds.
[0168] In some embodiments, the present invention discloses an Al based body measurements extraction system, comprises a virtual assistant with real-time guidance that provides visual and audio cues to guide the user for optimal pose and position during data acquisition (e.g., instructions such as “raise your left hand”, “step further away from your phone”, etc.). Further, the system implements customized semantic segmentation and contour refinement using a customized segmentation contour correction model, wherein the customized semantic segmentation contour correction model is trained from scratch, followed by contour refinement using classical image processing techniques such as edge detection, morphological operations, thresholding, etc. The system performs an error and problematic image analysis, wherein specific methods, such as fallback mechanism, for detecting problematic image conditions like glare, dark scenes, multiple people in the image, loose clothing, etc., are implemented.
[0169] In some embodiments, the body measurements are extracted by computing the circumference of specific body parts such as limbs, neck, hips, etc., based on major and minor axes of the Minimum Enclosing Ellipse. Custom algorithms may be implemented for the identification of specific body parts (e.g., shoulders, chest) based on segmentation and keypoint data. These algorithms help in the accurate extraction of circumferential and linear measurements. A statistical fallback system is implemented and sanity checks are performed to ensure plausible body measurements (e.g., filtering out unlikely values like “2m chest circumference for a 1.7m tall, 75kg person”). The invention also discloses a custom size chart parsing and coefficient calculation, wherein the system parses size charts from various formats (e.g., Excel, PDF, JSON, etc.), and how it computes coefficients to translate from clothing sizes to body measurements. The present invention handles errors using classical image processing techniques (e.g., histogram equalization, Fourier transform, and color space conversion) for correcting segmentation issues in problematic images (like lighting problems or cluttered backgrounds). Statistical fit predictor and predictor-corrector models may be implementedwherein the predictor-corrector model refines measurement estimates based on user responses about fitness level, job, and other characteristics, alongside fallback systems.
[0170] Three-dimensional (3D) models or avatars are generated from multiple sources using two approaches i.e., one from images using segmentation masks, and the other based solely on user input data (height, weight, age, gender), making the present invention a dualsource avatar generation system. The real-time feedback system for image acquisition is provided where a virtual assistant provides ongoing adjustments during the image acquisition process, ensuring the user meets specific pose and environmental conditions. The present invention is designed to handle complex backgrounds and clothing interferences (like loose clothes), or other elements (such as mirrors) that might hinder accurate measurement extraction. Also, the present invention is not limited to extracting the body measurements, but it also provides a separate widget for foot and hand measurements.
[0171] The present invention provides multiple advantages over the existing body measurement extraction techniques.
[0172] The present invention uses a customized trained Machine Learning model for estimating depth to identify some body-keypoints or landmarks which improves body measurements and eliminates the requirement of depth camera devices. For example, a two- dimensional (2D) input image is processed as an input using customized deep-learning models to estimate the depth. This is not an absolute depth, but a relative one. The 2D image utilized in the present invention is a single image processed to predict the depth information using monocular depth estimation, and each 2D image is processed independently for depth information. It may be appreciated that 2D image may be a single channeled grayscale image or a multi channeled (e.g., RGB) color image. The exact distance from any point to the camera may not be known, but some body-keypoints are known which are further or closer away than other points.
[0173] The present invention may be implemented to provide an offline body measurements extraction solution, which does not rely on the cloud at all, but rather runs all the code on a portable physical device. This device is a custom-made computer in a custom- made case, which uses special-purpose computing devices such as but not limited to Nvidia Jetson Orin Nano, and other edge devices. Also, the present invention also works on non-edge devices like desktop, laptop, server, etc., configured with motherboards containing CPUs and / or GPUs as per the user requirement. The offline solution also removes any device relatedissues for the cloud solution, where some devices may be low compute, may have bugs related to camera, gyroscope or other device element, may have a bad internet connection and so on. Also, it solves all the privacy and GDPR worries stated by businesses and users.
[0174] The present invention centralizes all the data and can provide further statistical analysis on the measurements data. Also, it allows the users to use additional data (like materials that can stretch a lot, clothes that shrink after first few washes and so on) to recommend again the sizes for people already measured for other collections or when new data arrives.
[0175] System Architecture: In some non-limiting embodiments, the present invention may be structured as a modular pipeline combining client-side acquisition, server-side processing, and integration interfaces with external systems such as e-commerce platforms, enterprise databases, or healthcare information systems. The architecture may ensure scalability, robustness, and compliance with both technical and regulatory requirements across diverse deployment scenarios.
[0176] At a high level, the system may comprise a client capture widget running on mobile or web applications for user guidance, a processing stack with cloud or edge device execution for measurement extraction and avatar generation, and an integration layer with APIs and data exchange protocols. The architecture may support both cloud-hosted deployments for scalability and offline deployments for privacy-sensitive contexts, with modular pipeline structure ensuring robustness, extensibility, and adaptability across industries.
[0177] Client Capture Widget: The client capture widget may be a lightweight application component designed to run on consumer devices including iOS and Android smartphones, tablets, and desktops with webcams. The widget may provide user guidance through visual overlays including silhouette templates, text instructions, and audio cues to ensure proper pose, distance, and orientation during image capture.
[0178] The widget may perform environment checks to detect background clutter, lighting conditions, and presence of multiple individuals, prompting corrective actions when problematic conditions are detected. Quality scoring may continuously evaluate capture quality using metrics such as sharpness, pose alignment, exposure, and subject centering, with images below threshold being automatically rejected. Upon acceptance, images and metadata may be encrypted and transmitted securely to the processing module.
[0179] Processing Stack - Cloud / Server Mode: In cloud or centralized server deployments, the processing stack may consist of multiple specialized modules. A preprocessing module may normalize input resolution and apply geometric corrections including perspective rectification and gyroscope-based alignment. The semantic segmentation module may utilize deep learning models trained on human- specific datasets to generate binary and multi-class masks for torso, limbs, and head regions.
[0180] The contour refinement module may apply classical image processing techniques including morphological operations, edge enhancement, and histogram equalization to correct segmentation artifacts. Keypoint detection modules may detect anatomical landmarks including shoulders, hips, knees, ankles, and wrists with associated confidence scores. Depth and scale estimation modules may compute pixel-to-metric conversion using user height input, monocular depth estimation networks, or device-assisted depth sensors.
[0181] The measurement extraction engine may calculate circumferences through minimum enclosing ellipses, linear distances, and surface contours, with each measurement tagged with confidence scores. The predictor-corrector system may cross-validate extracted values against statistical anthropometric models, applying redundant geometry or exemplar substitution when anomalies are detected. Avatar generation modules may fit parametric 3D models to the measurements, producing watertight avatar meshes, while recommendation engines may parse garment size charts and compute body-to-garment coefficients to return ranked recommendations with fit scores.
[0182] Edge / Offline Mode: For environments where cloud connectivity is unavailable or where privacy is paramount, the invention may include an offline appliance comprising a self- contained computational device such as but not limited to NVIDIA Jetson devices, industrial PCs or mobile devices such as tablets preloaded with the full processing pipeline. The offline mode may ensure no cloud dependence, with all image capture, processing, and recommendation computation occurring locally.
[0183] Edge deployments may be suitable for factories, clinics, defense facilities, or retail stores requiring data isolation and GDPR compliance. Models may be quantized and optimized for real-time inference on GPU-equipped edge devices, with secure update channels ensuring regulatory compliance while maintaining performance optimization for specialized hardware configurations.
[0184] Integration Layer: The integration layer may expose standardized interfaces for interoperability including REST and GraphQL endpoints for submitting user data, retrieving measurements, and fetching recommendations. Webhook support may provide real-time push notifications to partner systems, while standardized data formats including JSON schemas for measurements and OBJ / GLTF exports for avatars may ensure compatibility across platforms.
[0185] Pre-built e-commerce plugins may enable integration into popular retail platforms, while enterprise connectors may provide adaptors for ERP systems, enabling workwear and PPE procurement workflows. The integration layer may support batch processing for enterprise clients and event-driven integration through webhook mechanisms to ensure seamless deployment across diverse platforms and use cases.
[0186] Data Flow: The architecture may operate as a modular data pipeline proceeding through sequential stages. The client widget may capture user images and metadata, which are securely transferred to server or edge devices for preprocessing including image normalization and environment correction. Segmentation and keypoint detection may extract body contours and anatomical landmarks, followed by measurement extraction calculating circumferences, lengths, and other dimensions.
[0187] The predictor-corrector system may ensure accuracy and apply fallback mechanisms when needed, while 3D avatar generation may create parametric meshes with anatomical landmarks. Recommendation engines may compute garment sizes and generate ranked lists, with results returned to client interfaces, APIs, or enterprise systems. This modular flow may ensure robustness, extensibility, and adaptability across diverse industrial applications and deployment scenarios.
[0188] Training Data and Annotation: The performance of segmentation networks, keypoint detectors, and predictor-corrector models may depend on the quality and diversity of training data. The invention may incorporate a comprehensive strategy for dataset collection, annotation, augmentation, and refinement, ensuring robust performance across diverse populations, clothing types, and environments.
[0189] Data sources may include proprietary collections comprising curated datasets of full-body images, foot / hand captures, and associated ground truth measurements, partnership data collected through collaborations with industry partners such as apparel brands and medical institutions subject to consent and anonymization agreements, public datasets augmented with open-source resources such as COCO, MPII, and Human3.6M for keypoint and segmentationpretraining, and calibration sessions comprising onboarding events where users are measured with ground truth methods such as tape measure and 3D scanner and compared against system outputs.
[0190] The annotation process may include segmentation masks comprising pixel-level labeling of body contours and part-level divisions including head, torso, arms, and legs, landmark labels comprising manual placement of keypoints including shoulders, hips, knees, and wrists with inter- annotator consistency checks, foot / hand labels comprising contours and reference object alignment annotations, and size chart parsing ground truth comprising validation datasets of parsed size charts against brand-provided standards. Annotation may be performed by trained staff, with double-blind verification to ensure accuracy, and disagreements may be resolved through majority voting or expert adjudication.
[0191] Quality control measures may include inter-annotator agreement monitored continuously with thresholds greater than 90% required, annotation audits comprising random samples reviewed weekly, error taxonomy comprising mislabeling, missing labels, or drift identified and corrected, and feedback loops wherein the system automatically flags hard cases with low-confidence outputs for re-annotation.
[0192] To improve generalization, images may undergo augmentation including lighting variations comprising brightness, contrast, and glare simulation, background variations comprising random replacements with cluttered or plain environments, clothing variations comprising synthetic overlays simulating loose garments, uniforms, or PPE, occlusion simulation comprising artificial occluders such as bags, hands, and tools introduced, and geometric transformations comprising rotation, scaling, and cropping to mimic imperfect captures.
[0193] Active learning and hard-case mining may include uncertainty sampling wherein images where models exhibit high uncertainty are prioritized for annotation, error-driven refinement wherein cases failing sanity checks or fallback recovery are flagged, and domainspecific mining wherein special emphasis is placed on industrial clothing, medical wear, and PPE datasets. Ethical and privacy considerations may include consent frameworks wherein all proprietary data is collected with informed consent, anonymization wherein faces are blurred and metadata stripped before training, and regional compliance wherein data handling is aligned with GDPR and equivalent standards.
[0194] Evaluation Methodology and Benchmarks: Evaluation may be critical for validating the accuracy, robustness, and reliability of the invention. The system’s performance may be measured across datasets, conditions, and use cases, with benchmarks ensuring compliance with industrial expectations and providing transparency for clients and regulators.
[0195] The system may be evaluated on a combination of proprietary datasets comprising curated image collections with ground-truth measurements, calibration sets comprising data from calibration days including ground-truth try-on results, public benchmarks comprising datasets such as CAESAR, Human3.6M, or SizeUSA where relevant, and in-the-wild captures comprising user-submitted data under uncontrolled conditions. Evaluation protocols may use train / validation / test splits with no subject overlap to prevent bias.
[0196] Performance may be reported in terms of measurement error tolerance including large circumferences and linear dimensions with ±1% tolerance such as torso length, chest, hips, and thigh length, and small circumferences and linear dimensions with ±0.5% tolerance such as wrist, ankle, and neck. For example, for a 100 cm chest circumference, system error may be < 1 cm, and for a 20 cm wrist circumference, system error may be < 1 mm.
[0197] The recommendation engine may be evaluated separately using key performance indicators including accuracy comprising percentage of users for whom the recommended size matched the actually fitting size, top-2 recall comprising percentage of cases where the correct size appeared in the top two recommendations, collection-level accuracy comprising performance by brand / collection typically 90-95% under controlled conditions, and uncontrolled conditions comprising expected accuracy of 80-85% when tutorial guidance is ignored or environment is suboptimal.
[0198] Errors may be classified to support system improvement including acquisition errors comprising poor lighting, occlusion, and incorrect pose, segmentation errors comprising boundary leakage and background confusion, keypoint errors comprising misaligned or missing landmarks, scaling errors comprising incorrect user height or reference object misalignment, and recommendation errors comprising size chart anomalies and uncalibrated brand collections. This taxonomy may enable targeted retraining and software updates.
[0199] Ablation studies may be performed to demonstrate the contribution of each component including without contour refinement comprising impact of removing CV refinements, without predictor-corrector comprising effect of disabling fallback hierarchy, without multi-view fusion comprising using single-view only, and without calibrationcomprising using raw size charts without brand- specific adapters. Results may consistently show that each module improves accuracy, validating the modular design.
[0200] Deployment and Integration: To maximize adoption across industries, the invention may be designed for flexible deployment and seamless integration into existing ecosystems. It may support web-based retail environments, mobile applications, offline kiosks, and enterprise-level procurement systems, with deployment options tailored to meet the scalability, privacy, and interoperability requirements of different clients.
[0201] Deployment options may include webSDK comprising lightweight JavaScript library embeddable in e-commerce websites providing virtual assistant interface, measurement capture, and API calls for processing, mobile SDK comprising native iOS / Android libraries enabling direct integration into mobile apps including optimized camera workflows, offline preview, and fallback storage, offline kiosk comprising standalone installation on dedicated hardware such as edge devices for clinics, factories, or retail stores operating fully without internet connectivity, and hybrid mode comprising split processing between device for segmentation and pose estimation and cloud for avatar generation and recommendation for balance of performance and privacy.
[0202] APIs and data exchange may include REST / GraphQL APIs exposing endpoints for submitting image data, retrieving measurements, and fetching recommendations, webhook support enabling event-driven integration such as notification when a measurement session is complete, data schemas comprising standardized JSON structures for body measurements, size recommendations, and confidence metadata, and batch processing enabling enterprise clients to upload bulk datasets such as thousands of employees for overnight processing.
[0203] To ensure reliability, the system may include built-in observability comprising telemetry wherein logs capture quality scores, error categories, fallback usage, and latency metrics, dashboards wherein clients can monitor accuracy, adoption rates, and failure causes via a secure dashboard, and alerts comprising automated notifications if system accuracy drops below SLA thresholds.
[0204] The invention may support robust machine learning operations including model versioning wherein each deployment is tied to a specific model version with rollback capability, A / B testing wherein clients can test updated models against baselines in live environments, active learning loop wherein hard cases are automatically flagged for retraining,and continuous deployment wherein updates are delivered securely through encrypted OTA channels or physical media for offline systems.
[0205] Security and compliance may include authentication comprising OAuth2 and API key-based access control, audit logs wherein all API calls are logged for traceability, and regional processing wherein data is routed to region- specific servers to comply with local regulations such as EU-only processing for GDPR.
[0206] Industrial Embodiments: The system may support multiple industrial applications across diverse sectors. In e-commerce and fashion retail, customers may capture body images using the virtual assistant on their smartphone or desktop, measurements may be extracted and mapped to the brand’s size chart, size recommendations may be displayed on product pages reducing uncertainty and minimizing returns, and brands may analyze anonymized measurement distributions to refine sizing strategies, providing advantages including lower return rates, improved customer satisfaction, and reduced environmental impact from logistics.
[0207] In workwear and PPE applications, employees may be measured during onboarding or safety training sessions, the system may ensure PPE including helmets, gloves, safety boots, and overalls fits according to regulatory standards, and enterprise connectors may integrate directly with ERP systems for procurement automation, providing advantages including enhanced worker safety, reduced injuries, and compliance with safety regulations.
[0208] In medical and healthcare applications, the system may support compression garments for accurate sizing for lymphedema or venous insufficiency patients, prosthetics and orthotics for measurement-driven design of personalized medical devices, rehabilitation monitoring for longitudinal tracking of body dimensions over time such as swelling reduction and muscle atrophy recovery, and clinical kiosks comprising offline appliances deployed in clinics to ensure GDPR compliance and fast turnaround, providing advantages including improved patient outcomes, reduced trial- and-error fitting, and lower clinical workload.
[0209] In sports and performance applications, athletes and trainers may benefit from precise body measurement tracking including baseline body dimensions recorded pre-season, measurements updated over training cycles to monitor muscle development or weight changes, and sports equipment such as wetsuits and protective gear fitted with high accuracy, providing advantages including optimized performance, personalized training regimens, and injury prevention through proper equipment fit.
[0210] In defense and uniform provisioning, military and law enforcement organizations may require accurate, rapid sizing for large numbers of personnel including the system deployed in secure offline kiosks for batch measurement of recruits, uniforms, armor, and mission-specific equipment sized automatically, and data anonymized but aggregated for logistics and inventory planning, providing advantages including faster provisioning, reduced logistical errors, and enhanced operational readiness.
[0211] In interior design and ergonomics applications, body measurements may support ergonomics -driven design including furniture dimensions customized to target populations, workplace layouts optimized for safety and productivity, and automotive and aerospace seating tailored for comfort and safety, providing advantages including improved user comfort, reduced risk of injury, and better product-market fit.
[0212] The present invention may be implemented in various alternative configurations and embodiments that provide flexibility and adaptability across different operational requirements and deployment scenarios.
[0213] Single-Image Mode: While the preferred embodiment relies on frontal and lateral images, the invention also supports a single-image mode that uses advanced monocular depth estimation to infer 3D structure from one image. Anthropometric priors constrain plausible body shapes, while the predictor-corrector module compensates for reduced accuracy by refining outputs with statistical models. This mode is particularly useful in low-friction retail contexts where user compliance may be limited.
[0214] No Height Input: In cases where users cannot or will not provide their height, the system estimates scale using alternative methods including reference object detection such as A4 paper or credit card, device depth sensors where available, and statistical priors based on ratios of visible features such as head-to-body ratio.
[0215] Multi-Person Scenes: The virtual assistant enforces single- subject capture, but alternative embodiments can detect multiple subjects in frame, segment individuals separately generating measurements for each, and prioritize a selected subject through bounding box selection in the app. This extension supports family use cases such as parents measuring children or group calibration in enterprise contexts.
[0216] Occlusion Handling: For scenarios involving partial occlusion such as bag across torso or hair covering shoulders, missing contours are reconstructed using symmetryassumptions, landmark regression fills in missing points, and confidence is reduced proportionally to occlusion severity.
[0217] Loose or Heavy Clothing: An alternative embodiment enhances robustness to bulky clothing through under-clothing priors trained on paired datasets with clothed versus unclothed silhouettes, probabilistic adjustment of contours inward where clothing exceeds typical allowances, and user prompts requesting recapture if error margins exceed thresholds.
[0218] Different Segmentation Architectures: Although U2-Net is provided as an example, alternative segmentation backbones may be employed including CNN-based architectures such as DeepLabV3+ and UNet++, transformer-based models such as Swin Transformer and SegFormer, and hybrid models combining CNN encoder with Transformer decoder. This flexibility ensures that the system remains state-of-the-art as architectures evolve.
[0219] Alternative Predictor-Corrector Models: The fallback system may employ different statistical approaches including multivariate Gaussian models, random forests or gradient boosting regressors, and deep neural networks trained on anthropometric distributions.
[0220] Offline-Only Mode: For maximum privacy, the system may be deployed exclusively as an offline appliance with no internet connectivity, raw data processed and discarded locally with only non-identifiable derived data retained, and suitability for defense, medical, or highly regulated environments.
[0221] Hybrid Device-Cloud Workflows: Another embodiment splits processing with on- device capture, segmentation, and pose estimation, while cloud processing handles measurement synthesis, avatar generation, and recommendation. This balances performance, privacy, and resource constraints.
[0222] The present invention may be further extended and varied in numerous ways to enhance functionality, broaden applicability, and provide additional value across diverse use cases and industry requirements.
[0223] Returns Risk Prediction: An extension of the system includes returns risk modeling, where historical purchase and return data is combined with body measurements through feature vectors that incorporate measurement deviations, garment type, and brandspecific fit tendencies. Predictive models using machine learning classifiers predict probability of return for each recommended size, with output suggesting alternative size or alerting retailer if risk exceeds threshold. This reduces costly returns and improves sustainability.
[0224] Shape Categories and Body Typing: Beyond individual measurements, the system can classify users into body shape categories including common typologies such as pear, apple, rectangle, hourglass, and inverted triangle, and data-driven clusters derived from unsupervised clustering of large measurement datasets. Applications include brands using categories for marketing insights while users receive shape-aware recommendations, providing intuitive, non-technical feedback to consumers.
[0225] Visual Try-On Integration: While the primary function is measurement extraction, the system can integrate with virtual try-on systems through 2D overlay rendering garments on extracted body silhouette, 3D simulation fitting garment meshes onto generated avatar, and optional physics integration with cloth simulation for realistic draping. This enhances user engagement and confidence in purchase decisions.
[0226] Re-Measurement and Longitudinal Tracking: The system supports repeat measurement sessions for tracking changes over time including fitness applications to monitor muscle growth or fat loss, medical monitoring to track swelling reduction in patients with edema, and industrial ergonomics to assess workforce anthropometric changes over time. This enables longitudinal analysis and dynamic recommendations.
[0227] Advanced Anthropometric Indices: The system may compute advanced indices beyond basic measurements including Body Surface Area, fat-free mass estimates via regression on circumferences, and posture indices such as scoliosis angle proxies from spine contour. This extends applicability into healthcare, ergonomics, and sports science.
[0228] AI-Driven Style Matching: The recommendation engine can be extended to include style-based suggestions by combining body shape classification with garment style metadata, suggesting styles likely to flatter or suit functional needs. For example, for pear-shaped body, recommend A-line dresses over pencil skirts. This creates a bridge between measurement accuracy and user satisfaction.
[0229] Retailer and Manufacturer Analytics: The anonymized measurement data collected can be used for insights including size distribution analysis to identify gaps in size coverage, regional trends to detect regional body shape variations, and product development to adjust garment grading rules based on real-world body data. This aligns product design with actual consumer needs.
[0230] The following examples illustrate practical applications and deployment scenarios of the present invention across various industries and use cases.
[0231] In retail webshop sizing, customers activate measurement widgets, receive size recommendations with explanations (e.g., “Size 52 recommended - chest +2 cm, waist +3 cm”), reducing return rates and improving confidence. For PPE factory rollouts, 500 employees measured via offline kiosks generate automatic ERP orders for pre-sized safety equipment, ensuring compliance and simplified logistics. In athlete monitoring, baseline measurements track seasonal changes for optimized gear fit and personalized training. Medical clinic kiosks provide GDPR-compliant patient measurements for compression garment prescriptions with ±0.5% accuracy. Defense uniform provisioning uses secure offline appliances for rapid recruit measurement and automated sizing. Ergonomic furniture design leverages aggregated anonymized data to optimize product dimensions for target populations.
[0232] The following implementation considerations address key technical and operational aspects necessary for successful deployment and operation of the present invention across diverse environments and use cases.
[0233] Performance Targets: The system is designed for both consumer-facing responsiveness and enterprise-grade throughput with capture-to-result latency target less than 10 seconds in cloud deployments and less than 20 seconds in offline / edge mode, accuracy of ±0.5-1% error tolerance depending on measurement type, and recommendation latency with size recommendation computed in less than 1 second once measurements are available.
[0234] Scalability and Performance: The architecture scales through cloud mode with stateless microservices enabling elastic scaling across Kubernetes clusters, edge mode with optimized models for real-time inference on such as but not limited to NVIDIA Jetson devices, and batch processing for enterprise deployments handling thousands of measurements in parallel. Model optimization includes quantization to INT8 / FP16 precision, pruning of redundant weights, distillation of lightweight student models, and caching of common computations. Security includes OAuth2.0 authentication, role-based access control, comprehensive audit logging, and threat mitigation against injection attacks. Reliability ensures multi-region hosting, automatic failover, encrypted data replication, and continuous health monitoring. Device compatibility supports iOS 14± / Android 9+ smartphones, Chromium browsers with WebRTC, and Linux-based embedded systems with GPU acceleration.
[0235] Regulatory Compliance Beyond GDPR: In addition to GDPR, the invention can be configured for HIPAA in US with safeguards for handling patient health information, CCPAin California with user rights for data access and deletion, and ISO / IEC 27001 with alignment with enterprise security standards.
[0236] Maintenance and Updates: The system provides version control with clear versioning for models, APIs, and SDKs, update mechanisms with OTA updates for cloud and edge deployments, and backward compatibility with APIs maintained across versions to prevent client disruption.
Claims
AMENDED CLAIMS received by the International Bureau on 20 February 2026 (20.02.2026)1. A system for accurately extracting body measurements of human body parts using artificial intelligence-based techniques with real-time guidance and multi-level error correction, comprising one or more processors and a memory storing programming instructions, the memory is coupled to the processor wherein the processor is configured to execute the instructions to cause the system to: receive a plurality of user information and capture one or more images of the user in an optimal position by a virtual assistant providing real-time interactive guidance, wherein the virtual assistant continuously evaluates environmental conditions and is a machine learning model for detecting a plurality of key body points and enforcing predefined pose templates to guide the user to pose in the optimal position with environmental validation including lighting analysis and background clutter detection, the images and input of the user, the information and the images are provided to a remote server; identify one or more errors in the images including overexposed light conditions, underexposed light conditions, multiple people in the image, hazy backgrounds, multiple objects in the image, mirrors and reflective surfaces, and correct said errors by an image analysis and correction model with problematic scene classification; extract a body contour of the user from the images by a customized semantic segmentation model trained specifically on human body parsing datasets with clothing variability and environmental diversity; identify one or more errors in the body contour and refine said errors by a segmentation contour correction model combining deep learning outputs with classical image processing techniques including morphological operations, edge detection, histogram equalization, and color space conversion; and process the refined body contour and the key body points to identify one or more body parts and extract body measurements by an image processing model with confidence scoring and multi-view Bayesian fusion for measurement synthesis, wherein the system is further configured to generate a three-dimensional (3D) body model of the user using a multi-source avatar generation architecture, the three-dimensional body model being generated from one or more of:(i) captured images through segmentation mask projection, landmark mapping, depth estimation, and multi-view reconstruction techniques;(ii) direct three-dimensional reconstruction from one or more RGB images using deep learning-based geometry inference models configured to generate a full-resolution three- dimensional mesh representation of the user's body with accurate surface geometry;57(iii) statistical inference based solely on user information including height, weight, age, gender, and optionally additional user-provided metadata; wherein the 3D body model is generated through parametric body model fitting, implicit neural representation, mesh regression, or other computational body modeling techniques with constraint-based mapping of detected landmarks or inferred parameters to corresponding model vertices, enabling derivation of additional measurements from the generated 3D body model, wherein the system further comprises a statistical body modeling and measurement prediction module configured to: estimate one or more body measurements directly from user information including height, weight, age, gender, and optionally additional user-provided metadata without requiring image input; generate a three-dimensional body avatar based solely on statistical modeling of anthropometric datasets; and optionally refine the generated body avatar through user-provided preferences including fitness level or other user-specified parameters configured to modify underlying body shape parameters; wherein the statistical body modeling and measurement prediction module is trained on anthropometric datasets and is configured to operate as:(i) an independent measurement extraction flow,(ii) an independent 3D avatar generation flow, and(iii) a fallback measurement and modeling mechanism when image-based confidence levels fall below predetermined thresholds. and wherein the system further comprises dedicated measurement modules for foot and hand measurement extraction, each module configured with a separate widget from the full-body measurement widget, and configured to: capture one or more images of a foot or a hand positioned relative to a planar reference object; perform contour segmentation of the foot or hand region; detect anatomical keypoints specific to the foot or hand; and compute dimensional measurements including lengths, widths, girths, circumferences, arch parameters, and arch height proxies;2. The system of claim 1, wherein the image processing model is configured to compute one or more minor and major axes of the body parts using minimum enclosing ellipse fitting as an initial geometric approximation, followed by post-processing refinement using models58selected from geometric models, neural networks, regression models, or other computational approaches that refine the circumference to better accuracy than a minimum enclosing ellipse to determine corresponding circumferences of the body parts with circumferences measurements including limbs, thorax, abdomen, waist, and hips.
3. The system of claim 1, wherein the image processing model is configured to compute lines and contours of the body parts with non-circumference measurements including linear distances between anatomical landmarks and surface contour measurements along curved body surfaces.
4. The system of claim 1, wherein the system supports one or more input modes including:(i) image-based mode comprising a two-dimensional image, wherein the system supports both single-image mode with frontal view only and multi-view mode with frontal and lateral views of the user captured in sequence according to multi-view protocols; and / or(ii) non-image-based mode comprising user information including height, weight, age, gender, and optionally additional metadata without requiring captured images; and wherein the system is configured to generate a three-dimensional (3D) model of the user through one or more of:(a) parametric body model fitting based on detected landmarks or inferred body parameters;(b) statistical body modeling based on anthropometric datasets; or(c) direct three-dimensional geometric reconstruction configured to generate a metrically accurate mesh representation of the user's body.
5. The system of claim 1, wherein the system is configured to extract foot and hand measurements using dedicated measurement modules with planar reference objects and recommends a correct clothing, footwear and handwear size to the user through automated size chart parsing and body-to-garment coefficient computation.
6. The system of claim 1, wherein the user information includes height, weight, age, gender, and preference of the user, and wherein the system implements multiple scaling methods including height-anchored metric scaling, device-assisted scaling using depth sensors, and reference object scaling for pixel -to-metric conversion, and wherein the system utilizes monocular depth estimation to predict depth information from two-dimensional images without requiring specialized depth sensors.
7. The system of claim 1, wherein the system implements a predictor-corrector model architecture providing multi-level validation and error correction capabilities through statistical models trained on anthropometric datasets, enabling hierarchical fallback mechanisms including redundant geometric computation methods, statistical fit predictors,59exemplar matching against validated body shape libraries, skeleton-based proxy estimation, and userguided recapture workflows.
8. The system of claim 1, wherein the system implements multi-source avatar generation capabilities enabling three-dimensional body model creation from one or more of:(i) manual or user-provided input data including height, weight, age, gender, and optionally additional metadata without requiring captured images;(ii) one or more captured images processed through segmentation mask projection, landmark detection, and depth estimation techniques;(iii) one or more captured RGB images processed directly using deep learning-based three- dimensional reconstruction models configured to infer full body geometry from original color image data without requiring explicit segmentation masks; wherein each three-dimensional body model is generated using parametric body modeling architectures or other deep learning-based body modeling techniques, and wherein the three- dimensional body model is configured to provide a realistic avatar representation suitable for visualization, simulation, fitting, or digital interaction applications.9.The system of claim 1, wherein the system supports multiple execution modes including:(i) cloud-based processing with elastic scalability;(ii) edge device deployment on specialized hardware;(iii) fully offline operation performed entirely on a local computing device without requiring communication with a remote server; and(iv) software execution on mobile computing devices including smartphones and tablets, wherein the system is configured to operate as a native application, embedded software module, browser-based application, or other executable software capable of running directly on the mobile device hardware; wherein each execution mode includes ephemeral processing of raw images and derivative- only storage for privacy-preserving execution in medical, defense, enterprise, or consumer contexts.
10. The system of claim 1, wherein the system incorporates automated size chart parsing and normalization capabilities that ingest heterogeneous size chart formats including structured digital files, semi -structured documents, and unstructured image sources, compute body-to- garment conversion coefficients through collection-level regression analysis, and provide explainable size recommendations with transparent fit scoring and rationale including dimensional deviations and alternative options.6011. A method for accurately extracting body measurements of human body parts using artificial intelligence-based techniques with real-time guidance and multi-level error correction on a user device connected to a remote server, comprising the steps of: receiving a plurality of user information and capturing one or more images of the user in an optimal position by a virtual assistant providing real-time interactive guidance, wherein the virtual assistant is a machine learning model for detecting a plurality of key body points and enforcing predefined pose templates to guide the user to pose in the optimal position with environmental validation including lighting analysis and background clutter detection, the images and input of the user, the information and the images are provided to a remote server; identifying one or more errors in the images including overexposed light conditions, underexposed light conditions, multiple people in the image, hazy backgrounds, multiple objects in the image, mirrors and reflective surfaces, and correcting said errors by an image analysis and correction model with problematic scene classification; extracting a body contour of the user from the images by a customized semantic segmentation model trained specifically on human body parsing datasets with clothing variability and environmental diversity; identifying one or more errors in the body contour and refining said errors by a segmentation contour correction model combining deep learning outputs with classical image processing techniques including morphological operations, edge detection, histogram equalization, and color space conversion; processing the refined body contour and the key body points to identify one or more body parts by an image processing model with confidence scoring and multi-view Bayesian fusion for measurement synthesis; and extracting body measurements of the user using geometric approximation techniques including minimum enclosing ellipse fitting for circumferences as an initial estimation step, and further refining the estimated circumferences through one or more post-processing mechanisms including statistical correction models, regression-based adjustment, machine learning refinement, multi-view fusion, confidence-weighted optimization, or other computational enhancement techniques; and performing linear distance calculations for non-circumference measurements including distances between anatomical landmarks and contour-following measurements along curved body surfaces; and based on said body measurements recommending a suitable size through automated size chart parsing and body-to-garment coefficient computation.
12. The method of claim 11, further comprising implementing a predictor-corrector model architecture that provides multi-level validation and error correction through statistical models trained on anthropometric datasets, and applying hierarchical fallback mechanisms61including redundant geometric computation methods, statistical fit predictors, exemplar matching against validated body shape libraries, skeleton-based proxy estimation, and user- guided recapture workflows when measurement confidence falls below predetermined thresholds.
13. The method of claim 11, further comprising generating a three-dimensional body model through multi-source avatar generation, wherein the three-dimensional body model is created from one or more of(i) manual or user-provided input data including height, weight, age, gender, and optionally additional metadata without requiring captured images;(ii) one or more captured images processed through segmentation mask projection, landmark detection, and depth estimation techniques;(iii) one or more captured RGB images processed directly using deep learning-based three- dimensional reconstruction models configured to infer full body geometry from original color image data without requiring explicit segmentation masks; wherein the three-dimensional body model is generated using one or more deep learningbased body modeling architectures including parametric body model fitting based on learned body parameters and constraint-based landmark mapping; and wherein the generated three-dimensional body model is configured to provide a realistic avatar representation suitable for visualization, simulation, fitting, or digital interaction applications.
14. The method of claim 11, further comprising extracting foot and hand measurements using dedicated measurement modules with planar reference objects, wherein the planar reference objects provide scale references independent of user-reported height, and wherein the method includes contour segmentation of foot or hand regions, landmark detection of anatomical keypoints, and dimension calculation including length, width, and circumference measurements.
15. The method of claim 11, further comprising executing the method in multiple execution modes including cloud-based processing with elastic scalability, edge device deployment on specialized hardware, and fully offline operation with complete data isolation, wherein each execution mode includes ephemeral processing of raw images, derivative-only storage, and privacy-preserving execution, and wherein the offline operation is distinguished by performing all processing locally without requiring a remote server for medical, defense, and enterprise applications.
Citation Information
Patent Citations
Method and apparatus for virtual fitting
EP3972239A1
Systems and methods for full body measurements extraction
US10321728B1
Generation of body models and measurements
US10657709B2
Method and system for remote clothing selection
US11393163B2
Fast 3D model fitting and anthropometrics using synthetic data
US20160110595A1