A computer-implemented method for estimating body dimensions and providing clothing recommendations

The method uses deep neural networks and depth detection to calculate body dimensions from two images and estimate clothing size without a reference scale, addressing accuracy and complexity issues in existing virtual fitting technologies.

WO2025181627A1PCT designated stage Publication Date: 2025-09-04HOSSEINI SEYEDMASOUD +4

Patent Information

Application Number
PCT/IB2025/051806
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing methods for calculating body dimensions from images suffer from low accuracy, require multiple images, rely on complex 3D modeling, and lack a reliable system for precise garment labeling and feature recognition, while estimating garment sizes is challenging without a reference scale.

Method used

A computer-implemented method using deep neural networks and depth detection models to calculate body dimensions from two images (frontal and side view), automatically tag clothing, and estimate clothing size from various angles without a reference scale, employing a virtual fitting room without 3D models.

Benefits of technology

Achieves high accuracy in body size calculation, reduces user inconvenience by eliminating 3D modeling, enhances versatility in detecting clothing dimensions, and provides efficient virtual fitting experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025051806_04092025_PF_FP_ABST
    Figure IB2025051806_04092025_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method estimates body dimensions and provides clothing recommendations using frontal and side images from a user device. The images undergo validation via a deep neural network, ensuring correct posture, visibility, and lighting. A convolutional neural network (CNN) extracts geometric features, while segmentation techniques remove background noise. A simplified 3D body model is generated, and key body dimensions—such as shoulders, chest, waist, and legs— are calculated using object detection models and a standardized height reference. The method then suggests clothing sizes by matching the user's measurements with predefined size charts and enables a virtual fitting room for trying on clothes. An auto-tagging system identifies clothing attributes and links them to an image search engine for recommendations. All images are encrypted and securely stored, with an option for permanent deletion.The method can be integrated into third-party applications via an SDK or API, allowing real-time body measurement and virtual fitting functionality.
Need to check novelty before this filing date? Find Prior Art

Description

DescriptionTitle of Invention :A computer-implemented method for estimating body dimensions and providing clothing recommendationsTechnical Field

[0001] The present invention relates to the field of garment customization, and in particular, to a garment matching customization management system and method.Background Art

[0002] Existing platforms for calculating body size from images often suffer from low accuracy or require multiple images for reliable results. Clothing labeling systems are typically restricted to vendor-provided information or manual tagging, limiting their effectiveness. Virtual fitting rooms depend on complex and costly 3D modeling processes, while estimating garment dimensions generally requires a reference scale, such as a ruler or predefined product information. This invention introduces an innovative approach to overcome these limitations, offering a more accurate, efficient, and scalable solution.

[0003] Some methods use statistical shape models to estimate body shape based on image data or range maps, as described in US 2010 / 0111370. These models capture population-wide variations using a limited number of parameters, but they often introduce errors when applied to individuals, particularly in applications requiring high accuracy.

[0004] WO2011033258A1 presents an alternative approach for body modeling using image processing techniques. It describes a system that generates images of foreground objects by manipulating electromagnetic radiation emission to differentiate between overlapping objects. This process helps create an alpha matte for improved segmentation. Additionally, it discloses a method for generating a body model by mapping control points from a standard model to a subject’s body, allowing for a more individualized representation.

[0005] While these methods provide advancements in body measurement and modeling, they still face challenges in achieving the necessary accuracy for practical applications such as personalized clothing fittings.Summary of Invention

[0006] This platform utilizes deep neural networks and depth detection models to address existing issues by calculating body dimensions with high accuracy using two images (frontal and side view), automatically tagging clothing in images and linking them to an image search engine for similar recommendations, designing a virtual fitting room without needing a 3D model, and employing a depth-based model to estimate clothing size from various angles without a reference scale. The steps involve users uploading two images of themselves (frontal and side view) using their smartphone, the Al calculating the body dimensions and providing a suggested size, users viewing and filtering features of their desired clothing, and in the virtual fitting room, users uploading their image and trying on their chosen clothing. If there is a need to check the length or size of the clothing, the depth detection model estimates it.

[0007] Calculating Body Dimensions Using Two Images (Frontal and Side View) with High Accuracy: This system offers an intelligent and quick solution for measuring body dimensions, requiring only two simple images (frontal and side view) from the user to calculate various body dimensions with high precision. Unlike other methods that need images from multiple angles or 3D body scans, this system simplifies the process, making it easier for users. Key features of this service include:

[0008] No Special Equipment Required: Images can be taken with a smartphone camera.

[0009] Body Pose Estimation: Using advanced object detection models, the system identifies key body points and extracts measurements from these points.

[0010] Pixel-to-Real Unit Conversion: The model converts pixels to centimeters or inches using the user's height as the primary reference.

[0011] Accurate Measurement of Key Body Parts: Dimensions include chest circumference, waist circumference, hip circumference, upper body length, lowerbody length, arm length, shoulder width, neck circumference, underbust circumference, and even Bra Size and Cup Size.

[0012] High Processing Speed: Utilizing efficient architectures like MobileNet or EfficientNet, results are displayed to the user in the shortest possible time.

[0013] Environmental Flexibility: Algorithms are optimized for accurate recognition even in varying lighting conditions or images of medium quality.

[0014] Automatic Clothing Tagging in Images and Connecting to an Image Search Engine for Similar Recommendations: This service extracts detailed features of clothing from images and connects to an image search engine to offer similar recommendations as part of 52 of step 50. Unlike conventional methods requiring multiple images and angles, this system performs exceptionally well with just a single 2D image of the clothing. Key features include:

[0015] The system identifies features like color, collar style (round, V-neck, standing), sleeve type (long, short, sleeveless), clothing length (long, short, midi), pattern (plain, checkered, floral), and other details. The system utilizes neural networks like ResNet and Inception to extract precise features. Models perform well even with minimal training data using Data Augmentation and Transfer Learning strategies. Results are processed and displayed very fast. The search engine uses data from stores and other websites, providing a wide range of options to the user. Users can filter results by features like color, price, brand, or type of clothing. The engine can easily integrate with new stores or data sources.

[0016] This service allows users to try on clothing virtually by combining their image with the desired clothing image. Unlike common methods that require 3D models, this system only needs a full-body image of the user and 2D images of the clothing. Key features include:

[0017] No 3D Model Required: Clothing in stores can be used directly without conversion to 3D models.

[0018] Fast Processing: Using Image Overlay and Segmentation algorithms, the simulation process is quickly executed.

[0019] High Accuracy: Mask R-CNN technology is used to separate different parts of the image and naturally place the clothing on the user's image.

[0020] Ease of Use: Users only need to take an image with open hands in the standard frame of the application or SDK and try on the desired clothing. This technology plays a key role in improving the online shopping experience and reducing return rates.

[0021] Using a Depth-Based Model to Estimate Clothing Size from Various Angles Without a Reference Scale: One of the most advanced services of this platform is estimating the actual size of clothing from images without the need for a reference scale or objects of known size. Key features include:

[0022] Unlike other methods, the model does not require scaling objects (like rulers or size references) and operates solely based on image data. The model can consistently identify clothing dimensions even with images from different angles. This technology is useful for analyzing clothing in social networks or stores that do not provide dimensional information.

[0023] Steps:

[0024] The user uploads two images (frontal and side view) using their smartphone.

[0025] The Al calculates body dimensions and provides a suggested size.

[0026] Users can view and filter features of their desired clothing.

[0027] In the virtual fitting room, users upload their image and try on their chosen clothing.

[0028] If there is a need to check the length or size of the clothing, the depth detection model estimates it.Technical Problem

[0029] The invention aims to address several challenges in online shopping and virtual garment fitting, including the difficulty of quickly and accurately obtaining body measurements, the absence of a reliable system for precise garment labeling and feature recognition in photos, and the high costs and timeconsuming nature of 3D model-dependent virtual fitting rooms. Additionally, the lack of a reference scale makes estimating garment sizes challenging. To overcome these issues, the proposed solution introduces a simple yet precise method for calculating body dimensions using images, an intelligent system for garment labeling and visual similarity searches, a cost-effective virtual fittingapproach that eliminates the need for complex 3D modeling, and a model capable of accurately estimating garment dimensions without relying on a reference scale.Advantageous Effects of Invention

[0030] Higher accuracy in body size calculation using two images reduces common errors. The simplicity of eliminating the need for a 3D model for virtual fitting increases user convenience. Flexibility in detecting clothing dimensions without requiring reference information adds to the service's versatility. The efficiency of providing intelligent and quick suggestions for similar clothing enhances user experience. Innovation is demonstrated by the depth detection model for dimension estimation, which has fewer competitors in the market.Brief Description of Drawings

[0031] [Fig.1 shows the user’s guide in capturing properly positioned full-body photos, validated by an Inception-based deep learning mod.

[0032] Fig. 2 shows half side stance guide in front of the camera.

[0033] Fig. 3 shows the step of separating the clothing part of the image from the background.

[0034] Fig. 4 shows the process of identifying the position of a person's clothing in the recorded image.

[0035] Fig. 5 shows the use of a statistical algorithm to identify the dominant color of clothing.

[0036] Fig. 6 shows the step of identifying clothing that is specific to a particular gender as well as clothing that can be used jointly by both genders.

[0037] Fig.7 shows the step of recognizing upper body garment features.

[0038] Fig. 8 shows the step for determining whether a garment is cropped (above the person's navel).

[0039] Fig. 9 shows the step of recognizing lower body clothing features.

[0040] Fig.10 shows how to access the service, where the user connects to the system directly with the SDK or data is collected through another user interface (front-end) and the analysis results are returned to the same interface.

[0041] Fig.11 shows the steps of user images processing for virtual fitting, ensuring data privacy and providing size recommendations. ]Description of Embodiments

[0042] This platform utilizes deep neural networks and depth detection models to address existing issues:

[0043] Calculating body dimensions with high accuracy using two images (frontal and side view) in user input step(10).

[0044] Automatically tagging clothing in images and linking them to an image search engine for similar recommendations as part of 52 of step 50 in Fig. 11 .

[0045] Designing a virtual fitting room in step 40 without needing a 3D model, using existing store images.

[0046] Employing a depth-based model to estimate clothing size from various angles without a reference scale.

[0047] Steps:

[0048] Users upload two images of themselves (frontal and side view) using their smartphone.

[0049] The Al calculates the body dimensions and provides a suggested size.

[0050] Users can view and filter features of their desired clothing.

[0051] In the virtual fitting room, users upload their image and try on their chosen clothing.

[0052] If there is a need to check the length or size of the clothing, the depth detection model estimates it.

[0053] Body Size Measurement System

[0054] Accurate body size estimation using images is a critical area of research in computer vision and deep learning, with broad applications in medical diagnostics, apparel design, and human modeling. This invention introduces a novel system capable of non-contact estimation of various body dimensions using two user-provided images, leveraging advanced deep learning techniques. The system employs the user's height as a reference scale to convert pixel-based measurements into precise centimeter values.

[0055] To analyze input images, the system utilizes a convolutional neural network (CNN) architecture designed to extract key geometric features of the body as shown in step 30 of Fig.1 1 . These extracted features are integrated into a simplified three-dimensional model, enabling an accurate approximation of the user’s body structure. Additionally, the proposed algorithm is designed to function independently of specialized hardware, making it compatible with widely available consumer devices such as smartphones. This capability enhances accessibility while reducing human intervention, thereby improving both measurement accuracy and operational efficiency.

[0056] The subsequent sections detail the system's network architecture, data preprocessing methods, and estimation algorithms, followed by an evaluation of its performance through quantitative and qualitative experiments.

[0057] Standardization of User-Provided Height Data

[0058] Ensuring the consistency and accuracy of user-provided height data is a crucial step in maintaining system integrity and enhancing measurement precision. The system allows users to input their height in either centimeters or inches.

[0059] To standardize this input, as shown as part of 32 of step 30 in Fig.1 1 , the system automatically detects the unit of measurement. If the user enters their height in inches, the system converts the value to centimeters by multiplying it by 2.54. If the input is already in centimeters, it is stored without modification. This automated conversion eliminates discrepancies arising from regional variations in measurement systems, ensuring seamless data processing.

[0060] By standardizing height data, the system optimizes the pixel-to-centimeter ratio calculation, leading to more precise body dimension estimations. Additionally, this feature enhances user convenience, allowing for effortless interaction regardless of the measurement unit commonly used in their region. As part of the system’s intelligent design, this process ensures that all user input data is processed accurately, maintaining consistency in the final results.

[0061] Capturing a Full-Body Image for Analysis

[0062] The acquisition of a high-quality full-body image is a fundamental step in the body measurement process, as the accuracy of this image directly affects thesystem's final output. To ensure consistency and precision, the system requires the user to capture a full-body image using their smartphone camera, following predefined standards.

[0063] Automated Image Capture Process

[0064] Phone Angle Adjustment:

[0065] The system verifies that the smartphone is positioned at an optimal angle (between 85° and 95°) using the device’s gyroscope sensor.

[0066] This feature is implemented within a React FvJS-based interface.

[0067] Users receive real-time text and voice prompts to adjust their phone’s angle correctly, enhancing the overall experience and ensuring optimal image quality.

[0068] User Positioning:

[0069] The system guides the user through voice and text instructions to ensure correct body alignment within the camera’s frame.

[0070] The user is given some seconds to position themselves without requiring manual interaction with the device, ensuring a steady and well-aligned image capture.

[0071] Automated Image Capture:

[0072] Once the optimal position and angle are achieved, the system automatically captures the full-body image without requiring the user to press any buttons.

[0073] This automation minimizes user effort, reduces potential errors, and enhances measurement accuracy.

[0074] By integrating automated positioning and image capture functionalities, the system significantly improves user experience, reduces manual errors, and ensu Validation of Captured Images

[0075] Once the image is captured, it undergoes a validation process(26) using a deep neural network model. The validation system employs a pre-trained Inception network, which has been specifically optimized . The model has been trained on an extensive dataset of sample images, allowing it to recognize and validate images with high accuracy. To ensure continuous improvement, the system incorporates an automated training mechanism that updates the modelbased on newly validated data, enhancing both its precision and efficiency over time.

[0076] During the validation process, the image is assessed based on the following predefined criteria:

[0077] User Posture: The individual must be standing completely straight and upright.

[0078] Full-Body Visibility: The user’s entire body must be within the image frame.

[0079] Frontal Orientation: The person must be facing the camera directly.

[0080] Optimal Distance from Camera: The user must be positioned at a distance that is neither too close nor too far to maintain measurement accuracy.

[0081] Sufficient Illumination: The lighting conditions must be adequate to ensure proper image analysis.

[0082] Image Quality: The resolution and clarity must be sufficient for detecting necessary details.

[0083] If the image meets all these criteria, it is validated and the system proceeds to the profile image capture step. Otherwise, the system automatically provides realtime guidance, instructing the user on necessary adjustments before recapturing the image.

[0084] The validation process integrates advanced Al-driven technologies, including mobile device sensors, computer vision, and deep learning algorithms. Key innovations include gyroscope-assisted angle detection, which ensures proper camera alignment before image capture, Al-powered image validation that uses deep neural network models to confirm image compliance with system standards, and intelligent user guidance that provides real-time voice and text instructions to assist users in correctly positioning themselves. By combining these cutting-edge technologies, the system enhances measurement accuracy while offering a seamless, user-friendly experience.

[0085] Profile Image Acquisition

[0086] The profile image capture step complements the full-frontal image, providing additional data required for precise body dimension estimation. This step isimplemented with a combination of advanced Al techniques and user-centered design principles to optimize accuracy and ease of use.

[0087] Automated Profile Image Capture Process

[0088] User Positioning and Guidance:

[0089] The system provides step-by-step voice and text instructions to help the user position themselves correctly.

[0090] These instructions include details on stance, distance from the camera, and precise side-angle alignment.

[0091] This ensures users can effortlessly achieve the correct posture without confusion.

[0092] Automatic Image Capture:

[0093] The system assigns sufficient time for the user to adjust their position before starting the capture process.

[0094] No manual user interaction is required, reducing potential errors.

[0095] Validation of Profile Image

[0096] Once the profile image is captured, it undergoes a comprehensive validation process using the Inception deep neural network model, optimized for profile image analysis. The model continuously improves through automated learning mechanisms, incorporating newly validated data to enhance accuracy and performance.

[0097] The captured image is evaluated based on the following criteria:

[0098] Posture Accuracy: The user must be standing straight and upright to prevent measurement distortions.

[0099] Full-Body Visibility: The entire body must be visible within the frame.

[0100] Precise Profile Orientation: The user must be positioned at an exact 90- degree angle relative to the camera.

[0101] Optimal Distance from Camera: The user must be neither too close nor too far, ensuring accurate dimensional extraction.

[0102] Adequate Lighting Conditions: Uniform illumination is required to prevent shadows and loss of detail.

[0103] Image Resolution and Clarity: The captured image must be of sufficient quality to allow precise processing.

[0104] If the image meets all validation requirements, it is approved for further processing. Otherwise, the system automatically guides the user through corrective steps, repeating the capture and validation cycle until a satisfactory image is obtained.

[0105] Background Removal and Noise Reduction

[0106] The background removal and noise reduction process is a crucial step in enhancing image accuracy and quality, directly impacting the performance of computer vision and image processing algorithms. This process is essential for applications such as object recognition, body detection, and 3D modeling, ensuring that only the relevant portions of an image are analyzed.

[0107] Background Removal

[0108] Background removal is a fundamental image preprocessing technique that eliminates unnecessary elements from an image while preserving the subject of interest (e.g., the user's body). This operation enhances data efficiency, processing speed, and recognition accuracy by allowing algorithms to focus solely on key image details.

[0109] The process involves:

[0110] Image Segmentation Techniques: Foreground and background separation is achieved through advanced segmentation algorithms. Tools such as OpenCV and DeepLab are commonly used for this purpose.

[0111] Graph-Based Methods: OpenCV's GrabCut algorithm, a graph-cut-based technique, enables high-precision background removal by classifying pixels into foreground and background based on color, texture, and edge information.

[0112] Deep Learning-Based Background Removal: In cases of complex or noisy backgrounds, pre-trained deep neural networks (DNNs) are employed for more accurate and automated background extraction. These models analyze largedatasets to refine segmentation boundaries and improve the precision of the extracted subject.

[0113] By removing the background, the system reduces unnecessary data volume, increases processing efficiency, and enhances algorithmic focus on the subject.

[0114] Noise Reduction

[0115] Image noise, which can arise from factors such as low camera quality, inadequate lighting, or environmental interference, often degrades image clarity and reduces processing accuracy. To mitigate this, noise reduction techniques are applied, ensuring high-fidelity image analysis in subsequent stages.

[0116] Filter-Based Noise Reduction:

[0117] Gaussian Blur & Median Blur: These filters suppress high-frequency noise by averaging pixel values, effectively reducing random distortions while preserving edges.

[0118] Non-Local Means Denoising: This advanced algorithm analyzes pixel similarity across the image, effectively reducing noise while maintaining image details.

[0119] Super-Resolution Techniques:

[0120] For low-resolution images, deep learning-based Super-Resolution models enhance image clarity by reconstructing higher-resolution versions. These models use convolutional neural networks (CNNs) to generate sharper, more detailed outputs.

[0121] Image Storage and Data Protection

[0122] To safeguard user privacy and ensure comprehensive data security, stringent data protection measures have been implemented for image processing and storage. These measures are designed to minimize unauthorized access, prevent data misuse, and comply with global privacy regulations such as the General Data Protection Regulation (GDPR) and other applicable data protection laws.

[0123] The image storage framework follows a three-tier security approach:

[0124] 1. Facial Anonymization (Face Removal)

[0125] To protect individual privacy, the system utilizes advanced object detection algorithms to accurately identify the facial region in an image. This step ensures that facial data is never stored or processed beyond its immediate recognition.

[0126] Automated Face Detection: Al-powered computer vision models detect and isolate the facial area.

[0127] Face Removal & Redaction: Using OpenCV, Pillow, and other image processing frameworks, the detected face is completely removed from the image. This ensures that no personally identifiable facial features remain, rendering the image anonymous and non-traceable.

[0128] Compliance & Privacy Assurance: This anonymization step adheres to legal and ethical privacy standards, reducing the risk of biometric data misuse.

[0129] 2. Image Encryption & Secure Storage

[0130] To prevent unauthorized access, all processed images are encrypted before being stored. This encryption mechanism ensures that images remain unreadable and unusable without proper decryption keys.

[0131] End-to-End Encryption: Images are secured using AES-256 or an equivalent cryptographic standard, ensuring military-grade data protection.

[0132] Access Control Policies: Only authorized personnel or systems with strict access permissions can decrypt and retrieve stored images.

[0133] Data Integrity & Cybersecurity Compliance: These encryption practices align with ISO / IEC 27001 standards for information security management.

[0134] 3. Data Deletion & User-Controlled Erasure (Optional)

[0135] Users have full control over their stored data and can request permanent deletion of their images at any time. Upon user request, stored images undergo a secure deletion process to prevent recovery.

[0136] Complete & Irreversible Data Erasure: Deletion follows secure wiping protocols, ensuring that data is permanently removed from all storage locations and cannot be reconstructed.

[0137] User Consent & T ransparency: This feature empowers users to manage their data privacy according to their preferences.

[0138] Legal Compliance: The deletion process is aligned with right-to-be- forgotten principles under GDPR and similar regulations, guaranteeing user data is removed upon request.

[0139] Object Detection Process for Height Measurement and Scale Calculation

[0140] To accurately estimate height and calculate scale (pixels to centimeters), the system employs advanced object detection algorithms. This process ensures precision in body measurement applications while adhering to data privacy, biometric processing, and Al compliance regulations. The procedure follows these key steps:

[0141] 1. User Body Detection

[0142] The system utilizes state-of-the-art object detection models, such as YOLO (You Only Look Once) or Faster R-CNN (Region-based Convolutional Neural Networks), to detect the user's body position in the image. These models allow for precise localization of the user's body, particularly the vertical segments essential for height estimation.

[0143] 2. Bounding Box Generation

[0144] Upon detecting the body, the system defines a bounding rectangle that extends from the top of the head to the bottom of the heels. This bounding box serves as a reference for height calculation, ensuring that only the relevant body area is considered.

[0145] 3. Extraction of Height in Pixels

[0146] The system determines the vertical length of the bounding box in pixels, denoted as P_height. This value represents the total pixel count corresponding to the user's height in the image.

[0147] 4. Calibration: Comparing Pixel Length to Actual Height

[0148] To convert pixel measurements into real-world dimensions, the system requires a reference actual height (H_cm):

[0149] User-Provided Height: The user may enter their height manually.

[0150] Pre-Defined Standards: The system may use standard height measurements from medical or biometric databases.

[0151] This real-world height reference ensures that pixel-based measurements are accurately translated into centimeter-based dimensions.

[0152] 5. Scale Calculation (Pixels to Centimeters Ratio)

[0153] The conversion scale (centimeters per pixel) is derived using the following formula:

[0154] Scale=Hcm / Pheight

[0155] Where:

[0156] Scale = Centimeters per pixel

[0157] Hem = User's actual height (in cm)

[0158] Pheight = Height in pixels from the image

[0159] This scaling factor is then used for all subsequent body dimension calculations.

[0160] 6. Applying Scale for Other Body Measurements

[0161] Once the scale is established, the system automatically converts other body measurements (such as neck circumference, chest circumference, shoulder width, and waist circumference) using the same process.

[0162] For each body part:

[0163] The system detects and calculates the pixel length of the body part (Pbodypart).

[0164] The scale is applied to convert the pixel measurement into centimeters:

[0165] Lem =Scale x Pbodypart

[0166] Where:

[0167] Lem = Actual body part length (in cm)

[0168] Pbodypart = Number of pixels corresponding to the body part in the image

[0169] Scale = Conversion factor derived from height measurement

[0170] This automated process ensures high-precision, real-world body measurements using only image-based detection, making it ideal for medical applications, virtual fittings, fitness tracking, and biometric verification.

[0171] Legal and Compliance Considerations

[0172] The implementation of this system is designed to comply with global privacy regulations and biometric data protection standards, including:

[0173] GDPR (General Data Protection Regulation): Ensuring user consent and data minimization principles.

[0174] ISO / IEC 27701 (Privacy Information Management System): Standardizing biometric data handling.

[0175] HIPAA (Health Insurance Portability and Accountability Act), if applicable: Protecting user health-related measurements in medical applications.

[0176] Fair Al & Bias Reduction: Ensuring model fairness, accuracy, and mitigation of bias in height and scale detection.

[0177] Body Part Detection

[0178] To ensure accurate identification of body parts in full-face and profile images, advanced object detection algorithms are employed. YOLO (You Only Look Once) models, trained on diverse datasets, effectively detect and segment different body parts, including:

[0179] Head, Neck, Shoulders, Elbows, Chest, Waist, Legs.

[0180] Using deep learning and image processing techniques, these models can accurately locate key anatomical landmarks within an image. This process plays a crucial role in applications such as virtual clothing fitting, biometric verification, and augmented reality (AR) simulations.

[0181] Pixel Measurement for Body Parts

[0182] After detecting each body part, the number of pixels corresponding to that region is calculated through object detection algorithms. The workflow follows these steps:

[0183] The model first detects and isolates the body part (e.g., shoulders).

[0184] The number of pixels within the detected region is measured.

[0185] The pixel count is converted to real-world dimensions using a predefined scale factor (calibrated from height or reference objects).

[0186] This process enhances measurement accuracy for 3D modeling, digital health assessments, and customized virtual experiences.

[0187] YOLO Network Architecture & Functionality

[0188] The YOLO object detection model follows an end-to-end deep learning approach, enabling real-time body part detection. It operates as follows:

[0189] 1 . Grid-Based Object Detection

[0190] YOLO divides the input image into multiple grids and assigns bounding boxes to detected objects. Each grid cell generates a vector representation with the following structure:

[0191] (Pc,bx,by,bw,bh,c1 ,c2,... ,cn)

[0192] Where:

[0193] Pc Probability of an object being present in the cell.

[0194] (bx, by) Coordinates of the bounding box center relative to the grid.

[0195] (bw, bh) Width and height of the bounding box relative to the full image.

[0196] (c1 , c2, ..., cn) Probabilities of belonging to specific object classes (e.g., head, shoulders, waist).

[0197] 2. Processing Workflow

[0198] Image Preprocessing: The input image I is resized to a standard resolution (e.g., 416x416 pixels).

[0199] Feature Extraction: A CNN (Convolutional Neural Network) processes the image, extracting key features.

[0200] Bounding Box Prediction: Using fully connected layers, YOLO generates bounding boxes and confidence scores, enabling real-time detection.

[0201] Body Shape Recognition for Circumference Calculation

[0202] To ensure accurate circumference measurement of various body parts (e.g., waist, chest, thighs), it is essential to precisely identify and analyze the shape ofeach body segment. Since human body shapes vary significantly, failure to account for individual variations can result in measurement inaccuracies.

[0203] A robust body shape recognition system enables personalized and adaptive measurement based on individual body morphology, real-time shape analysis for enhanced accuracy in circumference calculations, and Al-driven corrections to ensure reliable and consistent results. Different body shapes require adaptive measurement models to cater to body shape variability and measurement adjustments. For instance, ectomorphic (thin) individuals typically have an oval waist shape, where the horizontal diameter (width) is greater than the vertical diameter (height). In contrast, endomorphic (larger abdomen) individuals may have a circular or even square waist shape, necessitating dynamic formula adjustments to ensure accurate measurements.

[0204] Challenges in standard measurement approaches include errors from fixed geometric models (e.g., assuming all waists are elliptical), traditional circumference formulas not accounting for variations in body contours, and subjective inconsistencies from manual measurements. The solution involves Al- based adaptive measurement models that adjust circumference calculations based on real-time body shape detection. To overcome these challenges, state- of-the-art computer vision and Al techniques are deployed, such as shape recognition via deep learning with pre-trained object recognition models (e.g., YOLO, Faster R-CNN) to detect key body regions, and body segmentation algorithms to isolate waist, chest, thighs, etc. Feature extraction networks analyze contours and geometry to determine the exact body shape. For adaptive circumference calculation, waist circumference is approximated as an ellipse or circle based on real-world geometric proportions, with Al dynamically modifying formulas to reflect observed body morphology.

[0205] The system can estimate body measurements using mathematical formulas or image processing.

[0206] Accurate Calculation of Body Part Dimensions

[0207] In the final stage of body measurement processing, precise dimensions of individual body parts are computed using pixel data and detected geometricshapes. This process ensures high accuracy in size estimation for various applications, including virtual fitting, health monitoring, and ergonomic analysis.

[0208] The calculation involves:

[0209] Extracting pixel-based dimensions from object recognition models.

[0210] Applying geometric shape analysis to account for variations in body morphology.

[0211] Using a pre-determined pixel-to-centimeter scale (derived from height) to obtain actual measurements.

[0212] For instance, to compute shoulder length, the system:

[0213] Identifies shoulder key points in the image. Measures the pixel distance between the key points.Applies the pixel-to-centimeter conversion scale to derive the real-world measurement.

[0214] This automated process ensures that shoulder length, chest circumference, waist circumference, hip circumference, knee circumference, and arm length are calculated with high precision and presented in real-world units.

[0215] The Measurement Service offers body measurement capabilities through two integration methods: a Software Development Kit (SDK) for direct in-app usage and an Application Programming Interface (API) for server-side processing. Both options provide scalable, accurate, and seamless body measurement solutions tailored to different business and technical needs.

[0216] SDK Integration: The Measurement Service SDK enables direct measurement processing within an application’s user interface, making it ideal for applications that require real-time feedback and a seamless user experience. Features and benefits include direct image processing where users submit images directly within the app, on-device analysis that reduces latency and enhances privacy, interactive results display providing real-time feedback on measurements, and customizable integration easily embedded into mobile, desktop, or web applications. The implementation process involves installing and initializing the SDK within the application, capturing or uploading an image of the user, sending the image to the Measurement Service’s module for analysis, receiving real-time measurement results (e.g., shoulder width, waist circumference), and dynamicallydisplaying the results within the app’s III. This approach is ideal for fashion, fitness, and health apps requiring instant user feedback.

[0217] API Integration: The Measurement Service API is designed for server-side processing, making it suitable for large-scale systems, enterprise applications, and data-intensive platforms. Features and benefits include batch processing capabilities for high-volume requests, centralized server control where measurements are processed externally, reducing app storage and processing overhead, advanced pre-processing techniques including background removal, noise reduction, and feature extraction, and structured data output in formats such as JSON, ensuring easy integration with backend systems. The implementation process involves developers sending user images via the API interface, the Measurement Service performing background removal, feature extraction, and body segmentation, computing measurements, and returning structured data (e.g., JSON output). Developers then integrate the results into their platform (e.g., e-commerce size recommendation system, fitness tracking app). This approach is ideal for web platforms, large-scale e-commerce, medical applications, and any system requiring batch processing and data storage.

[0218] Auto-Tagging

[0219] The purpose of this service is to organize and manage products. One of the primary challenges in this field is the process of accurate and efficient product tagging, which can play a significant role in facilitating search, personalized recommendations, and enhancing product display. Auto-Tagging , an advanced technology based on artificial intelligence and machine learning, has been introduced to address this challenge.

[0220] Auto-Tagging(50) is a process that uses image and text analysis related to products to automatically identify and tag key features, including type, color, pattern, material, and size. This technology leverages advanced image processing algorithms and deep learning to achieve higher accuracy and speed compared to traditional methods, minimizing the need for manual tagging.

[0221] The development and implementation of Auto-Tagging systems not only reduce costs and increase efficiency but also enable more personalized services for customers. For instance, faster and more accurate product searches,recommending similar items based on customer preferences, and improving inventory management systems are some of the benefits of this technology. Auto-Tagging improves the SEO of online stores and can increase the conversion rate of visitors to buyers.

[0222] Uploading a clothing image into the system, as the first step in the AutoTagging process(50), plays a crucial role in the automatic analysis and tagging of products. At this stage, users upload images of clothing, which can include a complete view of the clothing on a person or a raw image of the clothing itself. The primary goal of this stage is to prepare input data for subsequent advanced processing. These images can be collected from various sources:

[0223] User-taken images: Typically include images captured via smartphone cameras or other devices.

[0224] Promotional or catalog images: These images usually have high quality and are prepared directly for online sales or advertising.

[0225] Images taken in complex environments: Include photos with diverse backgrounds that may contain noise and unnecessary information.

[0226] Background Removal : In this stage, the system separates the clothing item from the background using advanced image processing algorithms, leading to noise reduction and improved image quality for subsequent stages. Background removal is an important stage in image processing, especially applicable to AutoTagging clothing in stores. This process specifically helps in isolating special areas of the image from non-essential parts (like backgrounds), allowing algorithms to focus more on primary features like color, model, and material of the clothing. Background removal not only enhances processing accuracy and speed but also reduces the volume of processed data, enabling systems to operate more efficiently.

[0227] Various image segmentation techniques are generally used in this process to separate the background from the foreground. This segmentation can be done manually or automatically. Tools like OpenCV and advanced models such as DeepLab or U-Net are employed to identify and separate different areas of the image. For instance, the GrabCut algorithm in OpenCV uses its graph-basedapproach to accurately separate the background from the foreground and shows good performance even when the background is complex.

[0228] Background removal ensures that only important areas like clothing remain, increasing the accuracy of machine learning models. This is especially useful in Auto-Tagging clothing systems, as after background removal, models can more accurately identify clothing features and categorize them accordingly. Overall, this technique is used in online stores to create cleaner and more accurate product images, in augmented reality systems for better clothing simulation, and in image analysis systems to extract specific clothing features.

[0229] Clothing Position Identification in Images

[0230] Identifying the position of clothing in images is a fundamental step in image analysis and processing, widely used in many artificial intelligence applications, particularly in Auto-Tagging systems for clothing, augmented reality (AR), and online shopping. The objective of this process is to detect and separate different parts of the clothing in an image, such as the upper body, lower body, and shoes. This allows machine learning models to identify various features of clothing more accurately and categorize them accordingly.

[0231] In this process, image processing algorithms use geometric and visual features such as color, texture, shape, and position in the image to identify different parts of the clothing. Generally, models and algorithms employed in this stage automatically separate the various parts of the clothing from other parts of the body and backgrounds. This stage is crucial for improving the accuracy of recognition and Auto-Tagging systems.

[0232] One of the most important tools used at this stage is the OpenCV library. OpenCV (Open-Source Computer Vision Library) is a collection of functions and algorithms for image processing that is widely used in image processing and computer vision projects. This library provides numerous tools for identifying and cropping different parts of an image. For example, algorithms available in OpenCV can easily identify various regions of an image based on color and geometric features and separate them from the main image.

[0233] It is essential to comply with data privacy and protection regulations when using image processing technologies. Ensure that any images used are collectedand processed with proper consent and that any data storage and handling comply with relevant laws and regulations, such as the GDPR (General Data Protection Regulation) in the European Union. Unauthorized use of images or personal data may lead to legal consequences.

[0234] Clothing Feature Detection

[0235] The primary objective of this stage is to extract precise and complex features from clothing items to aid in identifying and categorizing various types of clothing. To achieve this, Convolutional Neural Networks (CNNs) are utilized, which are among the most powerful models for image processing and analysis. These models are specifically trained to identify complex visual features such as color, texture, gender, pattern, type of clothing, and similar characteristics. Due to their high capability in analyzing spatial features of images, CNNs are an excellent choice for processing images of clothing. These models use convolutional layers to identify and extract various image features. Each convolutional layer identifies specific features of the image, such as edges, textures, and patterns. Typically, these networks include multiple convolutional layers and fully connected layers designed to recognize more complex features from images. In the process of clothing feature detection, CNN models are usually trained with large sets of training images where clothing items are accurately cropped and their features pre-identified. These images are typically provided to the model in a cropped format to allow the system to focus more precisely on each clothing feature. Below is a detailed explanation of each feature and the model used for it:

[0236] Color Detection: Identifying color is one of the fundamental features in clothing detection. To initiate the color detection process, the original image, which includes the person wearing the clothing, must be accurately examined. At this stage, a real-time object detection system as the YOLO (You Only Look Once) network is used, which can identify and distinguish various objects in images. The YOLO network is specifically employed for object detection in images and can separate different parts of the body, such as clothing, hands, or other body parts. This process is particularly useful when we want to identify only clothing and exclude other parts like skin and the individual's body from the image.

[0237] Once the clothing is cropped from the original image and extracted separately, the next step is to detect the color of each pixel of the clothing. To do this, image processing libraries like OpenCV are used. OpenCV provides a set of advanced functions and tools that allow for identifying the color of each pixel in the image. This library offers extensive capabilities for converting the color space of the image from different models, such as RGB (Red, Green, Blue) to other models like HSV (Hue, Saturation, Value) or HSL (Hue, Saturation, Lightness). These changes in color space enable the model to analyze colors more accurately and identify the dominant color in the image precisely.

[0238] In the next step, as shown in Fig. 5 a statistical algorithm is used to identify the dominant color. Generally, the goal is to select the color that has the highest frequency among the clothing pixels as the final color of the clothing. This process involves analyzing all pixels for their colors and counting the occurrences of each color in the image. Then, the color with the highest number of pixels is identified as the dominant color of the clothing.

[0239] Gender Classification: This stage is particularly important for identifying clothing that is specific to a particular gender and also for clothing that can be used by both genders. Accurate gender classification can help users make better choices and also improve clothing recommendation systems. To classify the gender of clothing, deep neural networks (DNNs) are used, which have high capabilities in processing and analyzing complex image data. These models can identify various features of clothing and determine whether the clothing is intended for women, men, or can be used by both genders. Initially, images of clothing for gender classification are provided to neural networks. Generally, the data is divided into three main classes:

[0240] Women's clothing, men's clothing and unisex clothing

[0241] These three classes include a large number of training data, each representing clothing specifically designed for one gender or both genders. These data typically include images of clothing from different angles, in various sizes, and under different lighting conditions, allowing the model to accurately identify common and distinguishing gender features of the clothing.

[0242] Deep neural networks, especially the Inception ResNetV2 model, are well- suited for this purpose due to their ability to extract complex visual features from images. This network comprises different layers, each identifying specific features such as edges, colors, textures, and patterns of the clothing. For example, women's clothing often has unique designs, lighter or floral fabrics, while men's clothing may exhibit simpler patterns and sturdier materials.

[0243] Lower Garment Feature Detection: Identifying features of lower garments, similar to upper garments, is crucial in automated clothing tagging systems. These features are particularly helpful in online clothing shopping, fashion recommendations, and recommender systems. In this section, two main features can be extracted for lower garments: type of clothing and clothing pattern.

[0244] The first feature that can be identified for lower garments is the type of clothing. At this stage, the system should recognize the type of lower garment, such as pants, shorts, or skirts. To achieve this, the Inception ResNetV2 neural network is employed. This network has been trained on three different types of lower garments and can distinguish each category with high accuracy.

[0245] The second feature that can be extracted for lower garments is the clothing pattern. Similar to how the ResNet152V2 network is used for identifying various patterns in upper garments, this network is also utilized for recognizing lower garment patterns. The patterns that this network can identify include animal prints, floral, gingham, graphics, simple, polka dots, stripes, and zebra. Due to its complex architecture, the ResNet152V2 network can accurately identify and distinguish various pattern features.

[0246] The ResNet152V2 neural network is specifically trained to identify different types of shoes. This network can accurately classify various shoes into some different categories as boots, sneakers, men's dress shoes, high heels and women's flats.

[0247] Using the complex features extracted from shoe images, this neural network can accurately identify the type of shoe in the image and categorize it into one of these classes.

[0248] Accessing the Service:

[0249] SDK: Users connect directly to the system using the SDK, and the analysis results are displayed within the same user interface.

[0250] API: In this method, the end user does not directly interact with the service. Instead, data is collected through another user interface (front-end), and the analysis results are returned to the same interface.

[0251] The method involves a comprehensive process for generating accurate body measurements and providing personalized clothing size recommendations. The step(10) of validating frontal and side images includes detecting and correcting image distortions caused by factors such as camera angle misalignment, lens warping, or user movement. This ensures that the captured images are free from artifacts that could compromise the accuracy of subsequent measurements (As shown in Fig. 1 .

[0252] To remove background elements in step (20) as processing module, the method employs deep learning-based segmentation models, which are trained to distinguish the user's body from the background with high precision. The part (24) of step 20 ensures that external objects do not interfere with body measurement calculations, thereby enhancing the reliability of the results.

[0253] In step(30)as body measurement system ,the convolutional neural network (CNN) (33) used in the process is pre-trained on a large-scale dataset of human body images and further fine-tuned with custom datasets to improve its performance and adaptability to specific use cases.

[0254] The method also includes detecting asymmetry in the user’s body, where leftright differences in body part dimensions, such as shoulder height imbalance or uneven waist curves, are identified and factored into the clothing size recommendations. This feature ensures that the recommendations account for individual anatomical variations, providing a more personalized fit.

[0255] The three-dimensional model of the user’s body is generated using advanced multi-view reconstruction techniques, which combine data from multiple angles to create a detailed and accurate representation of the user’s physique. This model serves as the foundation for all subsequent calculations and simulations.

[0256] To accommodate regional or brand-specific sizing standards, the method adapts clothing size recommendations by automatically converting the user’sbody dimensions into different sizing systems, such as US, EU, UK, and Asian sizing charts. This ensures compatibility with a wide range of clothing brands and markets.

[0257] The virtual fitting room interface leverages augmented reality (AR) overlays to dynamically simulate fabric draping, stretching, and fitting behavior. This provides the user with a realistic preview of how clothing items will fit on their body, enhancing the overall shopping experience. Additionally, the system provides real-time feedback on the user’s posture and positioning during image capture, guiding the user to stand straight, maintain an optimal distance, and ensure proper lighting for the most accurate measurements.

[0258] The clothing size chart matching part(62) in size estimation(60) involves using the garment sizing information derived from the depth-based model to find the most appropriate size from standard or retailer-specific clothing size charts. This might involve fuzzy matching algorithms to account for variations in sizing conventions.

[0259] A critical aspect covering the security of user data and compliance with privacy regulations.

[0260] Techniques such as blurring or pixelation to obscure the user's face, protecting their identity in face anonymization (72) .

[0261] Encrypting data from the user's device all the way to the storage server, preventing unauthorized access during transmission and storage.

[0262] Adhering to the General Data Protection Regulation, which includes obtaining consent for data processing, providing users with access to their data, and ensuring data security.

[0263] In the final output step(80),the culmination of all the processing steps, resulting in a size recommendation and possibly visualization of the user in the selected garment.

[0264] The system in user receives recommendations, presents the suggested clothing size and possibly visual feedback (virtual try-on) to the user.

[0265] In securely stored data part, user data (if stored) is held in a secure environment with appropriate access controls and encryption.

[0266] In the system output step, this signifies the completion of the process and delivery of the results to the user.

[0267] For continuous improvement, the method includes the step of storing anonymized body measurement data, which is securely encrypted and used solely for training improved body measurement models. This data is stripped of any personal identifiers, ensuring user privacy while enabling the system to learn and adapt over time.

[0268] The system also integrates a scoring mechanism to assess the accuracy level of body dimension calculations, providing users with confidence in the reliability of the results. Furthermore, users have the option to compare their body dimensions over time, allowing them to track changes and monitor progress.

[0269] To support seamless integration across platforms, the method includes SDK and API components that enable real-time body measurement synchronization across multiple devices. This ensures consistency and accessibility of data, regardless of the device being used.

[0270] The system dynamically adapts clothing size recommendations based on detected postural variations, ensuring that the recommendations remain accurate even if the user’s posture changes. Additionally, the system automatically detects and corrects minor measurement discrepancies caused by factors such as loose clothing, hair obstructions, or accessories, further improving the accuracy of detected body dimensions.Industrial Applicability

[0271] Online fashion and clothing stores use this platform to provide accurate sizing and reduce returns. Fashion platforms leverage it to simulate clothing on users' bodies. Clothing manufacturers employ it for automatic product tagging. Additionally, it has general applications for estimating the size of objects in images without a reference scale, i

Claims

Claims

1. A computer-implemented method for estimating body dimensions and providing clothing recommendations, the method comprising: a)receiving, from a user device, a frontal image and a side image of a user(10), b) validating the received images using a deep neural network model, wherein the validation process comprises: assessing user posture, ensuring full-body visibility, verifying frontal orientation, confirming an optimal distance from the camera, and checking for sufficient illumination to enhance processing accuracy, c) processing the validated frontal and side images using a convolutional neural network (CNN) to extract geometric features of the user’s body, d)removing background elements and reducing noise from the images using image segmentation techniques and noise reduction algorithms(24), exonerating a simplified three-dimensional model of the user's body based on the extracted geometric features, f)calculating body dimensions(30) by detecting and measuring specific body parts, including shoulders, chest, waist, hips, thighs, and legs, using an object detection model and converting pixel-based measurements from the images into real-world dimensions(32) using a reference scale derived from the user's standardized height, wherein height standardization includes automatically detecting the unit of measurement and converting it into a uniform unit, g) providing a suggested clothing size based on the calculated body dimensions and matching the user’s measurements with predefined clothing size charts, h) displaying, on the user device, a virtual fitting room interface that allows the user to virtually try on clothing items using the frontal and side images,^automatically tagging clothing items in the images using an auto-tagging system, wherein the auto-tagging system identifies clothing type, color, pattern, and material and links the tagged clothing items to an image search engine for recommending similar clothing items,j) encrypting and securely storing(74) the frontal and side images using end-to- end encryption and providing the user with an option to permanently delete the stored images upon request; and k) integrating the method into a software development kit (SDK) or an application programming interface (API) to enable real-time body measurement and virtual fitting room functionality within third-party applications.

2. The method of claim 1 , wherein the step of validating the frontal and side images further includes detecting and correcting image distortions caused by camera angle misalignment, lens warping, or user movement.

3. The method of claim 1 , wherein the step of removing background elements uses deep learning-based segmentation models to distinguish the user's body from the background.

4. The method of claim 1 , wherein the convolutional neural network (CNN) is pre-trained on a large-scale dataset of human body images and further fine-tuned with custom datasets.

5. The method of claim 1 , further comprising the step of detecting asymmetry in the user’s body, wherein left-right differences in body part dimensions are identified and factored into the clothing size recommendations.

6. The method of claim 1 , wherein the three-dimensional model of the user’s body is generated using multi-view reconstruction techniques.

7. The method of claim 1 , further comprising the step of adapting clothing size recommendations based on regional or brand-specific size charts, wherein the system automatically converts the user’s body dimensions into different sizing standards such as US, EU, UK, and Asian sizing systems.

8. The method of claim 1 , wherein the virtual fitting room interface uses augmented reality (AR) overlays to dynamically simulate fabric draping, stretching, and fitting behavior.

9. The method of claim 1 , wherein the system provides real-time feedback on the user’s posture and positioning during image capture, guiding theuser to stand straight, maintain an optimal distance, and ensure proper lighting for the most accurate measurements.

10. The method of claim 1 , further comprising the step of storing anonymized body measurement data for machine learning model improvements, wherein all user data is securely encrypted and used solely for training improved body measurement models without personal identification.

11. The method of claim 1 , wherein the SDK and API integration supports real-time body measurement synchronization across multiple devices.

12. The method of claim 1 , wherein the system dynamically adapts clothing size recommendations based on detected postural variations.

13. The method of claim 1 , wherein the system automatically detects and corrects minor measurement discrepancies caused by loose clothing, hair obstructions, or accessories, improving the accuracy of detected body dimensions.

Citation Information

Patent Citations

  • Machine learning image processing

    US20180012110A1

Cited By

  • Underwear virtual try-on design service method and system

    CN121258637A

  • Virtual fitting method and device based on dynamic fitting perception and medium

    CN122223220A

  • A virtual fitting method and device based on dynamic fit perception and a medium

    CN122223220B