Image-based anthropometric method, apparatus and device
By employing a multi-agent collaboration mechanism, the problem of insufficient universality and robustness of existing image-based human body measurement schemes is solved, enabling high-precision human body measurement in remote measurement scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- LINGDI (ZHEJIANG) TECHNOLOGY CO LTD
- Filing Date
- 2026-04-15
- Publication Date
- 2026-08-04
AI Technical Summary
Existing image-based anthropometric solutions lack differentiated detection strategies, resulting in insufficient versatility, flexibility, and robustness, making it difficult to meet the needs of remote measurement.
By adopting a multi-agent collaboration mechanism, through the coordinated work of the perception agent, task scheduler agent, and execution agent, automatic image understanding, task decomposition, adaptive scheduling, and quality control are achieved, thus constructing a human body auxiliary measurement system with versatility, flexibility, and high robustness.
It improves the versatility, flexibility and robustness of anthropometrics, enabling high-precision anthropometrics on images of different qualities and types, and adapting to remote measurement scenarios.
Smart Images

Figure CN122024020B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and more specifically, to an image-based human body measurement method, apparatus, and device. Background Technology
[0002] In fields such as virtual fitting, body shape assessment, and health monitoring, measuring human body dimensions is one of the most fundamental tasks.
[0003] Traditional methods of measurement, such as using measuring tapes and professional equipment to scan the human body, offer high accuracy but cannot meet the needs of remote measurement in a networked environment. Measurement methods using mobile phones to scan or capture images, on the other hand, are becoming increasingly common due to their low cost, convenience, and suitability for remote measurement scenarios.
[0004] However, existing image-based anthropometric schemes mainly rely on fixed preset strategies or models for detection, lacking differentiated detection strategies for different types of images, resulting in insufficient universality, flexibility and robustness of image-based anthropometrics. Summary of the Invention
[0005] The purpose of this application is to provide an image-based human body measurement method, apparatus, and device to improve the versatility, flexibility, and robustness of image-based human body measurement.
[0006] In a first aspect, embodiments of this application provide an image-based human body measurement method, including:
[0007] Acquire image perception information of the image to be detected;
[0008] Based on the image perception information, a human body measurement task strategy corresponding to the image to be detected is determined; wherein, the human body measurement task strategy includes multiple measurement sub-tasks;
[0009] Human measurement tasks are performed based on the image to be detected according to the human measurement task strategy, and human measurement results are obtained.
[0010] In this embodiment of the application, by acquiring image perception information and generating anthropometric task strategies based on the image perception information, targeted anthropometric measurements can be performed on different types of images, effectively improving the versatility, flexibility and robustness of anthropometric measurements.
[0011] In some embodiments, the image perception information includes image quality assessment information; the image quality assessment information is obtained in the following ways:
[0012] The quality of the image to be detected is evaluated from multiple preset evaluation index dimensions to obtain image quality evaluation information corresponding to the image to be detected.
[0013] In this embodiment of the application, the accuracy of the generated measurement task strategy is further improved by obtaining image quality assessment information from multiple dimensions.
[0014] In some embodiments, the image perception information further includes image semantic understanding information; the image semantic understanding information is acquired in the following ways:
[0015] Using the image to be detected and the image quality assessment information as input, a preset multimodal visual language model is used to obtain image semantic understanding information corresponding to the image to be detected;
[0016] The image semantic understanding information includes at least one of image semantic description information, human body recognition influencing factor information, and initial task agenda information; the image semantic description information is used to characterize the basic situation of the image to be detected, the human body recognition influencing factor information is used to characterize whether the image to be detected has preset degradation factors that affect human body recognition, and the initial task agenda information is used to characterize the type of human body analysis task applicable to the image to be detected in the preliminary prediction.
[0017] In this embodiment, high-level semantic understanding information of the image is obtained by combining quality assessment information, thereby further improving the accuracy of the generation measurement task strategy.
[0018] In some embodiments, determining the anthropometric task strategy corresponding to the image to be detected based on the image perception information includes:
[0019] The image perception information and the preset task execution experience base are combined for comprehensive analysis to determine the human body measurement task strategy corresponding to the image to be detected; wherein, the task execution experience base is used to store the recorded information of historical human body measurement tasks.
[0020] In this embodiment, a human body measurement task strategy is generated by combining image perception information and task execution experience base for comprehensive analysis, thereby further improving the reliability of the generated measurement task strategy.
[0021] In some embodiments, determining the anthropometric task strategy corresponding to the image to be detected based on the image perception information includes:
[0022] By combining the image perception information and the preset user configuration parameters, a comprehensive analysis is performed to determine the human body measurement task strategy corresponding to the image to be detected; wherein, the user configuration parameters are used to characterize the user-configured measurement function activation status, measurement preference information, and execution time constraints.
[0023] In this embodiment, a human body measurement task strategy is generated by combining image perception information and user configuration parameters for comprehensive analysis, thereby further improving the flexibility of generating the measurement task strategy and better meeting the personalized needs of users.
[0024] In some embodiments, performing anthropometric tasks based on the image to be detected according to the anthropometric task strategy to obtain anthropometric results includes:
[0025] Based on the aforementioned anthropometric task strategy, the corresponding measurement sub-tasks and task execution order are determined;
[0026] According to the task execution order, the task execution module corresponding to each measurement sub-task is called to perform human body measurement work and obtain the measurement task results corresponding to each measurement sub-task.
[0027] The results of each measurement task are integrated to obtain the human body measurement results.
[0028] In this embodiment of the application, the accuracy of human body measurement is further improved by obtaining the measurement task results sequentially according to the task type and execution order, and then integrating them to obtain the final human body measurement results.
[0029] In some embodiments, the step of calling the task execution module corresponding to each of the measurement sub-tasks in the order of task execution to perform human body measurement work and obtain the measurement task results corresponding to each of the measurement sub-tasks includes:
[0030] Each of the task execution modules uses a preset hybrid expert model to perform human body measurement tasks and obtains multiple corresponding candidate task results;
[0031] Each candidate task result is scored using a pre-defined scoring model, and the candidate task result with the highest quality score is taken as the measurement task result corresponding to the task execution module.
[0032] In this embodiment of the application, by using a hybrid expert model to generate candidate results in parallel through multiple routes during the execution of each task, and selecting the one with the highest quality score as the measurement result of the task, the accuracy of human body measurement is further improved.
[0033] In some embodiments, the image-based anthropometric method further includes:
[0034] If the quality score of the measurement task result corresponding to the target task execution module is lower than a preset score threshold, a task abnormality event is determined to have occurred, and corresponding abnormality description information is generated; wherein, the target task execution module is any one of the multiple task execution modules;
[0035] The anomaly description information is used to update the human body measurement task strategy;
[0036] Anthropometric tasks will be performed in accordance with the updated anthropometric task strategy.
[0037] In this embodiment of the application, a task rollback mechanism is triggered when the measurement result score of a certain task falls below a threshold, thereby further improving the reliability of human body measurement.
[0038] In some embodiments, the image-based anthropometric method further includes:
[0039] If the number of consecutive task exception events occurring in the target task execution module exceeds a preset threshold, a preset fault tolerance and degradation strategy is executed; wherein, the fault tolerance and degradation strategy includes adjusting the score threshold by lowering it.
[0040] In this embodiment, a fault tolerance degradation strategy is triggered when multiple consecutive failures occur to avoid infinite rollback interruptions, thereby ensuring the reliability of task execution.
[0041] In some embodiments, the step of acquiring image perception information of the image to be detected is implemented by a perception agent, the step of determining a human measurement task strategy corresponding to the image to be detected based on the image perception information is implemented by a task scheduler agent, and the step of performing human measurement work based on the image to be detected according to the human measurement task strategy to obtain human measurement results is implemented by an execution agent.
[0042] In the embodiments of this application, the execution efficiency and robustness of the human body measurement task are further improved by utilizing a perceptual agent, a task scheduler, and an execution agent to implement each step of the process.
[0043] Secondly, embodiments of this application provide an image-based human body measurement device, comprising:
[0044] The information perception module is used to acquire image perception information of the image to be detected;
[0045] The strategy determination module is used to determine a human measurement task strategy corresponding to the image to be detected based on the image perception information; wherein, the human measurement task strategy includes multiple measurement sub-tasks;
[0046] The measurement execution module is used to perform human measurement tasks based on the image to be detected according to the human measurement task strategy, and obtain human measurement results.
[0047] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, can implement the method described in any embodiment of the first aspect.
[0048] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement the method described in any embodiment of the first aspect.
[0049] Fifthly, embodiments of this application provide a computer program product, the computer program product including a computer program, wherein when the computer program is executed by a processor, it can implement the method described in any embodiment of the first aspect. Attached Figure Description
[0050] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0051] Figure 1 A flowchart illustrating an image-based human body measurement method provided in an embodiment of this application;
[0052] Figure 2 A system architecture diagram of an image-based human body measurement system provided in an embodiment of this application;
[0053] Figure 3 A schematic diagram of the structure of an image-based human body measurement device provided in an embodiment of this application;
[0054] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0055] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.
[0056] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0057] It should be noted that anthropometric measurement is a fundamental task in fields such as clothing customization, and current anthropometric technologies are mainly divided into three categories:
[0058] 1. Traditional measuring tape: This method remains fundamental and widely used, especially in high-end customization and specific fields. It relies on the experience of professional tailors for contact measurement, achieving high accuracy (especially for complex body shapes or specific dimensions). However, this method also has significant drawbacks: the process is cumbersome and time-consuming, requires highly skilled tailors, and the standards and techniques of different tailors are difficult to standardize, leading to poor data consistency. Most importantly, this face-to-face contact method is completely unsuitable for the needs of remote, customized tailoring in a network environment.
[0059] 2. Professional 3D Human Body Scanning: This is currently the mainstream non-contact, high-precision automated measurement technology. It primarily utilizes optical principles (such as structured light, lasers, and multi-view stereo vision) to quickly acquire point clouds of the human body surface, reconstruct a 3D model, and automatically extract a large amount of dimensional data (including girth curves and volume, which are difficult to measure using traditional methods). Its accuracy can reach the millimeter level, and its measurement efficiency is high. However, the application of this method is significantly limited: professional scanning equipment is extremely expensive, bulky, and typically requires a fixed installation environment. Furthermore, complex calibration is necessary after the equipment is moved to ensure measurement accuracy. These factors make this measurement method difficult to adapt to flexible, remote measurement needs.
[0060] 3. Emerging and convenient technologies (mobile phone scanning and AI image measurement): This is currently the most active area of development, aiming to solve the pain points of remote customization. It mainly consists of two types:
[0061] 3.1 Mobile 3D Scanning: This method utilizes the depth camera (such as LiDAR) of a smartphone or a regular camera in conjunction with an app to generate a 3D model and estimate its dimensions by having the user or someone else take multiple photos / videos. Its biggest advantages are low cost, high portability, and complete remote operation, making it extremely easy for users to adopt. However, its drawbacks include insufficient accuracy and stability (often with errors at the centimeter level), significant susceptibility to environmental factors, operation, clothing, and algorithms, and difficulties in standardization.
[0062] AI image measurement: This method directly predicts key dimensions (chest, waist, etc.) using a deep learning model based on just a few standard photos (such as front and side views) uploaded by the user. This method is characterized by low cost, convenience, suitability for remote scenarios, and high efficiency. However, its accuracy and reliability are its biggest bottlenecks. The prediction is essentially a statistical estimation, with limited ability to handle complex body shapes, non-standard postures, and loose clothing, and it relies heavily on a large amount of high-quality training data.
[0063] It should be noted that although the above-mentioned existing technologies have achieved human body analysis and measurement in their respective dimensions, most of the methods have high module coupling and lack generalization ability, making it difficult to adapt to the differences in quality of multi-source images or personalized needs.
[0064] To address at least one of the problems existing in the prior art, this application embodiment introduces a multi-agent design to achieve task-level adaptive combination, result quality judgment, and recovery process control, thereby constructing a human body auxiliary measurement system with versatility, flexibility, and high robustness.
[0065] like Figure 1 As shown in the figure, this application provides an image-based human body measurement method, which may include the following steps:
[0066] S1. Obtain image perception information of the image to be detected;
[0067] S2. Determine the human body measurement task strategy corresponding to the image to be detected based on image perception information; wherein, the human body measurement task strategy includes multiple measurement sub-tasks;
[0068] S3. Perform human measurement tasks based on the image to be detected according to the human measurement task strategy to obtain human measurement results.
[0069] In some embodiments, step S1 is an image perception step, which can be implemented by a perception agent; step S2 is a task scheduling step, which can be implemented by a task scheduler agent; and step S3 is a measurement execution step, which can be implemented by an execution agent.
[0070] It should be noted that by utilizing a multi-agent collaboration mechanism to implement each step of the proposed solution, it possesses functions such as automatic understanding of input images, task decomposition, adaptive scheduling, and quality control, thereby further improving the execution efficiency and robustness of the anthropometric task.
[0071] It should be noted that the embodiments of this application aim to construct an image-based human body auxiliary measurement system with high flexibility, interpretability and strong generalization ability, which can complete the structural modeling, posture restoration and key dimension estimation of the human body in 2D and 3D space without relying on professional hardware.
[0072] like Figure 2 As shown, specifically, the core role of the perceptual agent (image perception step) is to comprehensively evaluate and understand the input image (image to be detected), obtain the corresponding image perception information, and formulate an initial task agenda for subsequent anthropometric tasks. This step can integrate low-level image quality perception with high-level visual semantic understanding capabilities to ensure that the system has an accurate judgment on image content, usability, and task feasibility before performing human detection, pose estimation, segmentation, or 3D reconstruction.
[0073] The task scheduler agent (task scheduling step) primarily formulates a series of targeted human analysis and measurement execution paths (human measurement task strategies) adapted to the current input image (the image to be detected) based on the image perception information provided by the image perception step (such as image semantic information, image quality indicators, degradation type, and initial suggested task agenda). These paths serve as the basis for subsequent measurement execution. For example, different human measurement task strategies may include different numbers and types of measurement sub-tasks (e.g., human detection, image matting, human segmentation, pose estimation, SMPL 3D reconstruction, etc.). This agent is the central component connecting "cognition" and "execution" in the system, possessing the capabilities of task planning, adjustment, and personalized customization.
[0074] The executive agent is the core operational unit in a multi-agent collaborative system. It is responsible for progressively invoking specific execution modules (such as image analysis and human body modeling models) according to the measurement task plan issued by the task scheduler agent, completing each analysis and measurement task. This agent can incorporate functions such as parallel execution of multiple models, result selection, stage output, and interrupt response, acting as the "operator" that actually completes the human body measurement process.
[0075] Finally, based on the task execution results of each stage of the execution agent, they can be integrated into the final anthropometric results and output. These anthropometric results can be output and displayed according to a preset structured format or a visual chart.
[0076] Based on this, by acquiring image perception information and generating corresponding anthropometric task strategies according to the image perception information, different anthropometric task strategies can be adopted for images of different qualities and types to perform anthropometric tasks, effectively improving the versatility, flexibility and robustness of anthropometrics.
[0077] In some embodiments, image perception information includes image quality assessment information; the methods for obtaining image quality assessment information include:
[0078] The quality of the image to be detected is evaluated from multiple preset evaluation index dimensions to obtain image quality evaluation information corresponding to the image to be detected.
[0079] Specifically, after receiving an image to be detected, the perceptual agent can invoke various image quality assessment (IQA) models to evaluate the image's sharpness, compression, structural integrity, and perceptual quality. The evaluation metrics used include, but are not limited to, NIQE (Natural Image Quality Evaluator), MUSIQ (Multi-scale Image Quality Transformer), CLIP-IQA (CLIP-based Image Quality Assessment), and MANIQA (Multi-dimension Attention Network for IQA). These metrics are all no-reference image quality assessment models, reflecting the image's performance in terms of natural perceptuality, visual consistency, structural detail, and subjective preference. The perceptual agent can aggregate these multi-dimensional scores into a set of image quality feature vectors to obtain the corresponding image quality assessment information for subsequent analysis.
[0080] In this way, the quality assessment of the current image to be detected can be determined based on the image perception information. For images with different quality assessment results, different anthropometric task strategies can be adaptively adopted for anthropometric measurement. For example, when the quality assessment result in the image perception information indicates that the image has problems such as insufficient image resolution, insufficient clarity or insufficient structural integrity due to large-area occlusion of the human figure, an anthropometric task strategy that does not include 3D modeling is adopted, thereby avoiding the execution of invalid sub-task processes.
[0081] Based on this, by obtaining image quality assessment information from multiple dimensions, the accuracy of the generation measurement task strategy can be further improved.
[0082] In some embodiments, the image perception information further includes image semantic understanding information; the methods for obtaining the image semantic understanding information include:
[0083] Using the image to be detected and image quality assessment information as input, the system utilizes a pre-defined multimodal visual language model to obtain image semantic understanding information corresponding to the image to be detected.
[0084] The image semantic understanding information includes at least one of image semantic description information, human body recognition influencing factor information, and initial task agenda information. The image semantic description information is used to characterize the basic situation of the image to be detected, the human body recognition influencing factor information is used to characterize whether there are preset degradation factors affecting human body recognition in the image to be detected, and the initial task agenda information is used to characterize the type of human body analysis task applicable to the initially predicted image to be detected.
[0085] For example, the system can utilize multimodal visual language models (VLMs), such as LLaVA and Qwen-VL, to perform high-level semantic understanding of images. By using the image to be detected and its IQA (Image Quality Assessment) index as input to the VLM, the following three key information aspects can be generated: (1) Image semantic description information, which is used to describe the basic situation of the image, that is, the natural language expression of the person, pose, background, occlusion, etc. in the image; (2) Degradation list (human body recognition influencing factor information), which is used to describe whether the image has degradation factors such as blurring, occlusion, over-compression, and clothing that cover the body, which affect human body recognition; (3) Initial task agenda, which is used to describe the executable human measurement and analysis tasks that the model predicts are suitable for the current image quality conditions, such as only performing 2D key point extraction, or performing a complete SMPL (Skinned Multi-PersonLinear Model, a human body parameterized representation model based on skin vertices) fitting process.
[0086] For example, another key function of image perception is determining whether to trigger certain sub-processes in the subsequent task chain. For instance, if the system determines that the image resolution is insufficient or that there is extensive occlusion of the human figure, it will mark 3D modeling as not recommended in the initial task agenda information. The subsequent task scheduler will then avoid assigning 3D modeling tasks based on this, thus preventing the execution of invalid tasks. Simultaneously, the perception agent is also responsible for judging spatial conditions such as the number of human figures in the image, pose integrity, and facial visibility, assisting the system in formulating targeted human measurement and processing paths.
[0087] Ultimately, the perceptual agent outputs structured information, including image quality assessment vectors, semantic descriptions, degradation type labels, and suggested task sequences, which are then processed by the task scheduler as the basis for decision-making in subsequent processes. This module's design enables the system to possess the intelligent front-end capabilities of "understanding images, recognizing limitations, and formulating plans," and is a key starting point for the multi-agent architecture of this application.
[0088] Based on this, by combining quality assessment information to obtain high-level semantic understanding information of the image, the accuracy of the generation measurement task strategy can be further improved.
[0089] In some embodiments, determining a human measurement task strategy corresponding to the image to be detected based on image perception information includes:
[0090] By combining image perception information and a pre-set task execution experience base, a comprehensive analysis is performed to determine the anthropometric task strategy corresponding to the image to be detected; the task execution experience base is used to store the recorded information of historical anthropometric tasks.
[0091] For example, the task scheduler agent can combine visual language models (such as LLaMA-Vision, GPT-4V) or language inference models to comprehensively analyze image perception information output by the perception agent, such as image descriptions, degradation lists, and task experience, thereby outputting an ordered list of anthropometric tasks (anthropometric task strategy). The measurement sub-tasks in the anthropometric task strategy include, but are not limited to: human detection, image matting, segmentation, 2D keypoint estimation, 3D reconstruction (such as SMPL), dimensional measurement, and scene calibration. The task scheduler agent will not only decide "which tasks to perform," but also "the order of tasks" and "whether to skip certain tasks," thus ensuring optimal overall workflow.
[0092] For example, to further improve the rationality of task combinations, the task scheduler agent can access the "task execution experience base" recorded in the system, which contains records of historical anthropometric measurements and their success / failure information. For instance, if historical tasks indicate that a certain image feature combination failed during SMPL reconstruction, the task scheduler agent will automatically avoid similar paths and arrange other task flows when planning tasks for similar image feature combinations, such as skipping 3D modeling and only outputting estimation results based on 2D key points. This experience-based feedback reasoning mechanism gives task scheduling memory and self-optimization capabilities.
[0093] Based on this, a comprehensive analysis combining image perception information and task execution experience base is performed to generate anthropometric task strategies, thereby further improving the reliability of the generated measurement task strategies.
[0094] In some embodiments, determining a human measurement task strategy corresponding to the image to be detected based on image perception information includes:
[0095] By combining image perception information and preset user configuration parameters for comprehensive analysis, the human body measurement task strategy corresponding to the image to be detected is determined; among which, the user configuration parameters are used to characterize the user-configured measurement function activation status, measurement preference information, and execution time constraints.
[0096] For example, the task scheduler agent can also customize the planned tasks based on parameters in the user configuration file (such as whether 3D modeling is enabled, measurement preferences, execution time constraints, etc.). For instance, when deployed on a mobile device, the scheduler can choose to skip modules with high computational costs (such as 3D reconstruction); in medical measurement scenarios, it will prioritize tasks that ensure high accuracy, such as human body segmentation and key point localization.
[0097] It should be noted that the output of the task scheduler agent is a clear and orderly task plan (anthropometric task strategy), in the format of structured JSON data or intermediate state diagrams. This data will be passed to the execution agent for specific anthropometric task operations. Through the task scheduler agent, the system possesses "decision-making ability after understanding," capable of transforming ambiguous image semantics into a clear and adaptable measurement process. It is the core scheduling unit in the multi-agent collaborative execution chain.
[0098] Based on this, by combining image perception information and user configuration parameters for comprehensive analysis, a human body measurement task strategy is generated, thereby further improving the flexibility of the generated measurement task strategy and better meeting the personalized needs of users.
[0099] In some embodiments, an anthropometric task is performed based on the image to be detected according to the anthropometric task strategy to obtain anthropometric results, including:
[0100] The corresponding measurement sub-tasks and task execution order are determined based on the anthropometric task strategy;
[0101] The task execution modules corresponding to each measurement subtask are called in the order of task execution to perform human body measurement work and obtain the measurement task results corresponding to each measurement subtask.
[0102] The results of each measurement task are integrated to obtain the anthropometric results.
[0103] For example, the anthropometric task strategy indicates the measurement subtasks to be performed for the current image to be detected, as well as the corresponding task execution order. Multiple measurement subtasks can be executed sequentially according to this order, and the output of one measurement subtask will serve as the input to the next. Furthermore, one or more of the multiple measurement subtasks can be tasks that do not affect each other; therefore, these measurement subtasks can be executed in parallel, and the execution results of each task will be used as the basis for the final result fusion.
[0104] For example, each measurement subtask corresponds to a task execution module. Based on the task execution order, these task execution modules can be used to perform measurement operations on each measurement subtask in a serial or parallel manner to obtain the measurement task results corresponding to each measurement subtask.
[0105] For example, when the anthropometry task strategy requires the "human detection" measurement subtask, the corresponding task execution module will call detection models such as YOLOv8 and DETR to extract human bounding boxes from the image. If the "image matting" measurement subtask is required, image segmentation models such as MODNet and RVM will be scheduled to generate a refined human foreground. If the "human segmentation" measurement subtask is required, parsing models such as CIHP and SCHP will be loaded to divide the body region into semantic parts such as head, torso, and limbs. When entering the "pose estimation" measurement subtask stage, the executing agent will select an appropriate keypoint detector (such as ViTPose, DETR-Pose, or OpenPose) based on image quality and task configuration to extract human joint position information and output a two-dimensional skeleton structure. Based on this, if the anthropometric task strategy includes the measurement sub-task of "SMPL 3D reconstruction", then it further calls SMPL parameter regression models such as PARE, PIXIE, and VIBE to fit the 3D human body mesh structure, outputting body parameters (β), posture parameters (θ) and camera pose, etc., to generate a complete 3D model that can calculate spatial measurement indicators such as height, shoulder width, chest circumference, and leg length.
[0106] Finally, these measurement task results can be fed back to the task scheduler agent (or another result fusion agent) for result integration to obtain the final human body measurement results. Understandably, the executing agent can generate structured measurement results (such as a JSON-formatted table of human body dimensions) based on preset output types, as well as optional human mesh files, pose diagrams, and image visualization overlays, providing standardized results for the backend system or user interface.
[0107] Based on this, by obtaining the measurement task results sequentially according to task type and execution order, and then integrating these results to obtain the final anthropometric results, the accuracy of anthropometric measurements is further improved.
[0108] In some embodiments, the task execution module corresponding to each measurement subtask is called in the order of task execution to perform human body measurement work, and the measurement task results corresponding to each measurement subtask are obtained, including:
[0109] Each task execution module uses a pre-set hybrid expert model to perform anthropometric tasks and obtains multiple candidate task results.
[0110] The results of each candidate task are scored using a pre-defined scoring model, and the candidate task with the highest score is taken as the measurement task result corresponding to the task execution module.
[0111] It should be noted that in each task step, the executing agent is not always based on a single model, but can be based on a Mixture-of-Experts (MoE) mechanism for task execution: for each task, the system can call multiple predefined "expert models" to execute in parallel. For example, two different segmentation models or pose models can be used in the same stage to generate corresponding candidate task results.
[0112] Then, a scoring model consisting of an image quality evaluator (such as HPSv2, NIQE) and a semantic consistency scoring module is used to score the results of all candidate tasks. The result with the highest score is selected as the measurement result of the task step and is used as the input to the next task execution module, or as the basis for the final integrated result.
[0113] Based on this, by using a hybrid expert model to generate candidate results in parallel through multiple routes during each task execution, and selecting the highest quality score as the measurement result for that task, the accuracy of human body measurement is further improved.
[0114] In some embodiments, the image-based anthropometric method further includes:
[0115] If the quality score of the measurement task result corresponding to the target task execution module is lower than the preset score threshold, a task abnormality event is determined to have occurred, and corresponding abnormality description information is generated; wherein, the target task execution module can be any one of multiple task execution modules;
[0116] Update the anthropometric task strategy based on anomaly description information;
[0117] Anthropometric tasks will be performed in accordance with the updated anthropometric task strategy.
[0118] Please continue to refer to this. Figure 2 It should be noted that a task rollback mechanism can also be built into the execution agent. For example, when the quality score of the execution result of a certain measurement subtask (the measurement task result with the highest quality score selected by the task execution module) is lower than a preset threshold, it is determined that a task abnormality event has occurred. The system can locate the cause of this low score based on the scoring logic or the preset abnormality detection strategy (such as severe missing key points, SMPL mesh distortion, segmentation mask damage, etc.) and report the abnormality description information to the task scheduler agent. The task scheduler agent then updates the human body measurement task strategy. For example, it can regenerate the human body measurement task strategy and execute it again in sequence, or rearrange the alternative path only for the task step where the abnormal event occurred (such as replacing 3D modeling with 2D measurement), thereby avoiding the interruption of the overall process or the propagation of errors.
[0119] Understandably, the rollback and reflection mechanism is a core mechanism to ensure the overall stability, robustness, and task continuity of the system. It is specifically designed to address potential execution failures, abnormal results, or insufficient image adaptation during human body analysis. This mechanism is nested within the execution agent and forms a closed-loop feedback relationship with the perception agent and the task scheduler agent, enabling automatic detection, fault reporting, task adjustment, and process repair capabilities.
[0120] Specifically, after the agent completes a task step, the system immediately evaluates the quality of the output from that task execution module. The evaluation method can combine indicators such as image content, structural rationality, and semantic consistency, including no-reference image quality indicators (such as NIQE), human preference scoring, structural symmetry analysis, keypoint count, and 3D model stability scores (such as mesh topology distortion and body collapse in SMPL). If the score of any indicator (or the weighted average of multiple indicator scores) is lower than the system's set quality score threshold, the task step will be judged as a failure or unacceptable (i.e., a task anomaly event has occurred).
[0121] Once a failure is detected, the rollback mechanism will be activated, and corresponding exception description information will be generated immediately, including the failed task name, input image status, failure score details, and a list of executed tasks. This information can be passed back to the task scheduler agent as new input conditions.
[0122] After receiving an anomaly description, the task scheduler agent can use a multimodal language model to perform task reasoning again. For example, it can regenerate a new anthropometric task strategy or adjust the current path based on the existing task agenda. For instance, if SMPL reconstruction fails, the system may choose to backtrack to 2D keypoint measurement; if segmentation fails, the system can prioritize using pose structure for measurement estimation; if occlusion is severe and affects multiple tasks, the system can even terminate the process and prompt the user to replace the image.
[0123] By introducing rollback and reflection mechanisms, the system achieves dynamic perception and real-time adjustment of abnormal execution paths, significantly enhancing the stability and practical applicability of the entire measurement system. This is a key support for building a highly generalized and automatically adaptable human body auxiliary measurement system.
[0124] Understandably, the execution agent, through functions such as model integration and quality control, achieves highly reliable and high-precision execution of key tasks in the human image processing workflow, which is a key guarantee for the practicality and performance of this system.
[0125] Based on this, when the score of a certain task's measurement result falls below a threshold, a task rollback mechanism is triggered, thereby further improving the reliability of human body measurements.
[0126] In some embodiments, the image-based anthropometric method further includes:
[0127] If the number of consecutive task exception events occurring in the target task execution module exceeds a preset threshold, a preset fault tolerance and degradation strategy is executed; the fault tolerance and degradation strategy includes adjusting the scoring threshold by lowering it.
[0128] It should be noted that, based on the task rollback mechanism, the system can support a finite loop logic of try-evaluate-retry. Specifically, for a failed task, the scheduler can attempt to arrange an alternative model or alternative path for re-execution. However, if multiple alternatives fail (i.e., the number of consecutive task exception events in a task execution module exceeds a preset threshold), the system will execute a "fault tolerance degradation strategy." For example, the system can automatically adjust the target accuracy (e.g., lower the scoring threshold to relax the accuracy restriction standard) or result type to ensure controllable convergence of the process, rather than indefinite interruption.
[0129] Furthermore, for example, based on the task rollback and reflection mechanism, the system also supports feeding back the experience of each task failure into the system's "recovery experience base," thereby updating the reference strategy for subsequent perception and scheduling processes. This feedback can generate long-term statistics on image categories, degradation conditions, and task adaptability, thereby enhancing the system's self-evolution capabilities.
[0130] In this embodiment, a fault tolerance degradation strategy is triggered when multiple consecutive failures occur to avoid infinite rollback interruptions, thereby ensuring the reliability of task execution.
[0131] In some embodiments, the system may further include a configuration module for achieving flexible adaptation, multi-scenario customization, and user-level control. Its function is to provide a unified parameterized interface, allowing users or the system to customize the agent's behavior strategies and task execution logic based on specific task requirements, device capabilities, target accuracy, and other factors. The configuration module supports both static presets and dynamic loading at runtime, serving as the fundamental mechanism for the system to achieve "controllable intelligence."
[0132] For example, the configuration module can maintain a set of standardized configuration items, each corresponding to a key system control variable, mainly including the following categories:
[0133] 1. Task selection control
[0134] The configuration module allows users to explicitly enable or disable specific subtasks. For example, users can set whether to enable SMPL 3D reconstruction, or whether to perform human semantic segmentation or image matting. The system will then adjust the task scheduler agent's logic for generating task paths accordingly, avoiding unnecessary calculations or inappropriate task flows.
[0135] 2. Setting Measurement Targets
[0136] Users can select the measurement dimensions they want to focus on, such as "height + shoulder width", "head-to-shoulder ratio", "torso ratio", and "limb length". The configuration module transforms these targets into a combination of intermediate tasks required for the task chain (for example, leg length measurement requires segmentation + pose + pixel scale inference), which are dynamically loaded by the task scheduler agent.
[0137] 3. Restore Preference Definition
[0138] This option controls whether the system prioritizes perceived quality (visual naturalness) or structural accuracy (geometric symmetry, numerical stability) in the result quality assessment. For example, when set to "Fidelity," the system tends to select model outputs with more accurate keypoint localization and smaller dimensional errors during task execution; when set to "Perception," it can focus more on the naturalness and visual continuity of the resulting image.
[0139] 4. Quality tolerance threshold and degradation tolerance
[0140] The configuration module allows setting image quality thresholds (e.g., NIQE < 5.0) and fault tolerance strategies. When the input image quality is below the threshold, the system can automatically disable some highly dependent tasks (such as 3D modeling) to avoid invalid computations. The "Automatic Degradation" option can also be enabled, triggering a rollback path after a critical task fails, and automatically downgrading accuracy limits or adjusting the result type when there are too many consecutive rollbacks.
[0141] 5. Operating mode control
[0142] Options include whether to enable Fast mode (skipping time-consuming models), whether to process subtasks in parallel, and whether to enable the face enhancement path (for head size analysis). These options allow the system to run efficiently in scenarios with limited computing power, high real-time requirements, or lightweight tasks.
[0143] 6. Output Format and Interface Options
[0144] Users can specify the output format (such as JSON parameter table, CSV size data, 3D mesh file, key point skeleton image, etc.) and export path, which facilitates integration into downstream applications such as virtual try-on, data analysis interfaces, etc.
[0145] It should be noted that the configuration module is essentially a policy constraint on the behavior of each intelligent agent in the system. Through a unified parameter interface, it enables a highly personalized human measurement process. This ensures that the system can not only automatically adapt to image content but also dynamically optimize the task chain based on user intent, achieving a shift from "general measurement" to "task-customized measurement."
[0146] Compared with the prior art, the embodiments of this application have at least the following beneficial effects:
[0147] 1. By introducing a multi-agent collaborative architecture, the entire process of perception-planning-execution-reflection in human body measurement tasks is made autonomous: For the first time, a multi-agent system is introduced into the image-based human body measurement process, enabling the system to have image understanding, task scheduling, quality assessment and dynamic path repair capabilities. This breaks through the bottlenecks of rigid process and slow response of traditional measurement methods, and significantly improves the system's intelligence level and task adaptability.
[0148] 2. Construct a task-level hybrid expert model scheduling mechanism to achieve dynamic selection of multiple models for the same task and output of the best results: For each measurement subtask (such as segmentation, pose estimation, SMPL reconstruction, etc.), the system calls multiple expert models in parallel and combines image quality and semantic consistency to score and filter the results, automatically selecting the best result for subsequent processing, thereby greatly improving the overall accuracy and robustness of the system and avoiding the single-model path dependency problem.
[0149] 3. Establish a task rollback and reflection mechanism to achieve fault recovery and path reconstruction at the task flow level: When any measurement subtask fails due to image defects or model misjudgment, the system automatically triggers the rollback mechanism, records the abnormal state of the task, and readjusts the task path or alternative model execution scheme through intelligent reasoning, ensuring that the task flow runs in a closed loop without interruption, which significantly enhances the stability and engineering reliability of the system.
[0150] 4. Supports adaptive task path planning based on visual semantics and image quality, improving the targeting and resource efficiency of task execution: The perceptual agent integrates image quality assessment and visual language model reasoning capabilities, and completes the determination of image structure, degradation type and executable task range before the measurement process begins. This enables the task scheduler to output the most suitable task sequence based on the actual availability of the image and the user's goals, thereby improving overall computational efficiency and execution value.
[0151] 5. Provides a configurable measurement process control interface to achieve unified modeling and flexible deployment of human body measurement tasks in multiple scenarios: Through configuration module design, task type, target dimension, accuracy preference, output format, etc. are parameterized and abstracted to achieve full-process configuration support from lightweight fast measurement to high-precision structural restoration, meeting the needs of different application scenarios such as virtual try-on.
[0152] Please refer to Figure 3 , Figure 3 The diagram illustrates a block diagram of an image-based human body measurement device provided in some embodiments of this application. It should be understood that this image-based human body measurement device is similar to the one described above. Figure 1Corresponding to the method embodiments, it is able to perform each step involved in the above method embodiments. The specific functions of the image-based human body measurement device can be found in the description above. To avoid repetition, detailed descriptions are appropriately omitted here.
[0153] Figure 3 The image-based anthropometric device includes at least one software functional module that can be stored in memory or embedded in the image-based anthropometric device in the form of software or firmware. The image-based anthropometric device includes:
[0154] Information perception module 310 is used to acquire image perception information of the image to be detected;
[0155] The strategy determination module 320 is used to determine the human measurement task strategy corresponding to the image to be detected based on image perception information; wherein, the human measurement task strategy includes multiple measurement sub-tasks;
[0156] The measurement execution module 330 is used to perform human measurement tasks based on the image to be detected according to the human measurement task strategy, and obtain human measurement results.
[0157] It is understood that the above-described device embodiments correspond to the method embodiments of the present invention. The image-based human body measurement device provided by the embodiments of the present invention can implement the image-based human body measurement method provided by any one of the method embodiments of the present invention.
[0158] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the device described above can be referred to the corresponding process in the aforementioned method, and will not be elaborated further here.
[0159] like Figure 4 As shown, some embodiments of this application provide an electronic device 400, which includes: a memory 410, a processor 420, and a computer program stored in the memory 410 and executable on the processor 420, wherein when the processor 420 reads the program from the memory 410 via a bus 430 and executes the program, it can implement any of the methods included in the above-described image-based human body measurement method.
[0160] Processor 420 can process digital signals and may include various computing architectures. For example, it may be a complex instruction set computer architecture, a reduced instruction set computer architecture, or an architecture that implements multiple instruction set combinations. In some examples, processor 420 may be a microprocessor.
[0161] Memory 410 can be used to store instructions executed by processor 420 or data related to the execution of instructions. These instructions and / or data may include code for implementing some or all of the functions of one or more modules described in the embodiments of this application. The processor 420 of this disclosure embodiment can be used to execute instructions in memory 410 to implement the methods shown above. Memory 410 includes dynamic random access memory, static random access memory, flash memory, optical memory, or other memories well known to those skilled in the art.
[0162] Some embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, describes the method described in the method embodiments.
[0163] Some embodiments of this application also provide a computer program product that, when run on a computer, causes the computer to perform the methods described in the method embodiments.
[0164] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0165] It should be understood, in the several embodiments provided in this application, that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0166] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0167] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0168] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0169] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0170] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
Claims
1. An image-based anthropometric method, characterized in that, include: Acquire image perception information of the image to be detected; Based on the image perception information, a human body measurement task strategy corresponding to the image to be detected is determined; wherein, the human body measurement task strategy includes multiple measurement sub-tasks; Human measurement tasks are performed based on the image to be detected according to the human measurement task strategy to obtain human measurement results; The image perception information includes image quality assessment information; the methods for obtaining the image quality assessment information include: The quality of the image to be detected is evaluated from multiple preset evaluation index dimensions to obtain image quality evaluation information corresponding to the image to be detected; The image perception information also includes image semantic understanding information; the methods for acquiring the image semantic understanding information include: Using the image to be detected and the image quality assessment information as input, a preset multimodal visual language model is used to obtain image semantic understanding information corresponding to the image to be detected; The image semantic understanding information includes at least one of image semantic description information, human body recognition influencing factor information, and initial task agenda information; the image semantic description information is used to characterize the basic situation of the image to be detected, the human body recognition influencing factor information is used to characterize whether the image to be detected has preset degradation factors that affect human body recognition, and the initial task agenda information is used to characterize the type of human body analysis task applicable to the image to be detected in the preliminary prediction.
2. The image-based anthropometric method according to claim 1, characterized in that, The step of determining the human measurement task strategy corresponding to the image to be detected based on the image perception information includes: The image perception information and the preset task execution experience base are combined for comprehensive analysis to determine the human body measurement task strategy corresponding to the image to be detected; wherein, the task execution experience base is used to store the recorded information of historical human body measurement tasks.
3. The image-based anthropometric method according to claim 1, characterized in that, The step of determining the human measurement task strategy corresponding to the image to be detected based on the image perception information includes: By combining the image perception information and the preset user configuration parameters, a comprehensive analysis is performed to determine the human body measurement task strategy corresponding to the image to be detected; wherein, the user configuration parameters are used to characterize the user-configured measurement function activation status, measurement preference information, and execution time constraints.
4. The image-based anthropometric method according to claim 1, characterized in that, The step of performing anthropometric tasks based on the image to be detected according to the anthropometric task strategy to obtain anthropometric results includes: Based on the aforementioned anthropometric task strategy, the corresponding measurement sub-tasks and task execution order are determined; According to the task execution order, the task execution module corresponding to each measurement sub-task is called to perform human body measurement work and obtain the measurement task results corresponding to each measurement sub-task. The results of each measurement task are integrated to obtain the human body measurement results.
5. The image-based anthropometric method according to claim 4, characterized in that, The step of calling the task execution module corresponding to each measurement subtask according to the task execution order to perform the human body measurement operation and obtaining the measurement task results corresponding to each measurement subtask includes: Each of the task execution modules uses a preset hybrid expert model to perform human body measurement tasks and obtains multiple corresponding candidate task results; Each candidate task result is scored using a pre-defined scoring model, and the candidate task result with the highest quality score is taken as the measurement task result corresponding to the task execution module.
6. The image-based anthropometric method according to claim 5, characterized in that, Also includes: If the quality score of the measurement task result corresponding to the target task execution module is lower than a preset score threshold, a task abnormality event is determined to have occurred, and corresponding abnormality description information is generated; wherein, the target task execution module is any one of the multiple task execution modules; The anomaly description information is used to update the human body measurement task strategy; Anthropometric tasks will be performed in accordance with the updated anthropometric task strategy.
7. The image-based anthropometric method according to claim 6, characterized in that, Also includes: If the number of consecutive task exception events occurring in the target task execution module exceeds a preset threshold, a preset fault tolerance and degradation strategy is executed; wherein, the fault tolerance and degradation strategy includes adjusting the score threshold by lowering it.
8. The image-based anthropometric method according to any one of claims 1 to 7, characterized in that, The step of acquiring image perception information of the image to be detected is implemented by a perception agent; the step of determining the human body measurement task strategy corresponding to the image to be detected based on the image perception information is implemented by a task scheduler agent; and the step of performing human body measurement work based on the image to be detected according to the human body measurement task strategy to obtain human body measurement results is implemented by an execution agent.
9. An image-based human body measurement device, characterized in that, include: The information perception module is used to acquire image perception information of the image to be detected; The strategy determination module is used to determine a human measurement task strategy corresponding to the image to be detected based on the image perception information; wherein, the human measurement task strategy includes multiple measurement sub-tasks; The measurement execution module is used to perform human measurement tasks based on the image to be detected according to the human measurement task strategy, and obtain human measurement results; The image perception information includes image quality assessment information; the methods for obtaining the image quality assessment information include: The quality of the image to be detected is evaluated from multiple preset evaluation index dimensions to obtain image quality evaluation information corresponding to the image to be detected; The image perception information also includes image semantic understanding information; the methods for acquiring the image semantic understanding information include: Using the image to be detected and the image quality assessment information as input, a preset multimodal visual language model is used to obtain image semantic understanding information corresponding to the image to be detected; The image semantic understanding information includes at least one of image semantic description information, human body recognition influencing factor information, and initial task agenda information; the image semantic description information is used to characterize the basic situation of the image to be detected, the human body recognition influencing factor information is used to characterize whether the image to be detected has preset degradation factors that affect human body recognition, and the initial task agenda information is used to characterize the type of human body analysis task applicable to the image to be detected in the preliminary prediction.
10. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, can implement the image-based human measurement method according to any one of claims 1-8.