Body composition prediction method based on digital anthropometric phenotype
Through digital anthropometric phenotype and machine learning model, the existing body composition measurement methods have limited accuracy and high cost are solved, and fast, economical and accurate body composition measurement is achieved, which is suitable for a variety of people.
Patent Information
- Application Number
- CN202510006567.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-06-03
AI Technical Summary
The existing body composition measurement methods have problems such as limited accuracy, high cost and complex operation, and cannot accurately reflect the obesity level of different groups of people. Especially for Asian people, the globally common BMI critical point is not accurate enough.
Using digital anthropometric phenotype combined with machine learning models, the input digital anthropometric data predicts body components related to fat content/muscle content, and establishes accurate body component prediction methods without complex measurement procedures.
It realizes fast, economical and accurate body composition measurement, with an accuracy of close to traditional high-end measurement technologies such as DXA, and is suitable for a variety of people, reducing time and economic costs.
Smart Images

Figure CN120089353A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of physical anthropology, and particularly relates to a method for predicting body composition based on digital anthropometric phenotypes. Background Art
[0002] Research has shown that obesity is a major risk factor for various chronic diseases, such as type 2 diabetes [1], metabolic syndrome [2], cardiovascular diseases [3], etc. There is an obvious positive correlation between body fat content and death risk. Accurately evaluating body composition such as fat content and fat-free content is of great significance for guiding people to maintain a healthy weight, control overweight and obesity, and the risk of related chronic diseases.
[0003] Body Mass Index (BMI) is a simple measurement index, calculated by dividing body weight (in kilograms) by the square of height (in meters). For decades, BMI has been used as an international standard for indirectly assessing body fat percentage and judging healthy weight. Research has shown that a high BMI is often associated with an increased risk of metabolic diseases and premature death [4]. Due to its simple measurement method and low cost, BMI is often used as an indicator to evaluate total fat mass and judge overweight or obesity. However, BMI has certain limitations. It cannot distinguish between fat content and fat-free content, nor can it distinguish between abdominal fat and gluteofemoral fat; in addition, BMI may not accurately reflect the obesity levels of people of different races, ethnic groups, genders, and living backgrounds.
[0004] An expert consultation meeting of the World Health Organization concluded that for a certain proportion of Asians, when BMI is higher than the overweight critical point (> = 25 kg / m^2) specified by the World Health Organization, the high risks of type 2 diabetes and cardiovascular diseases increase significantly. In the 2002 Deurenberg study, different Asian populations, including Indonesians, Chinese Singaporeans, Malays, and Indians, were compared. Compared with Caucasians, Asian populations had a lower BMI but a higher BF% (Body Fat Percent). The high BF% at a low BMI may be related to differences in body shape, muscle mass, and height. Therefore, the globally applicable BMI critical point is not an accurate method for assessing the risks associated with high BF% in Asian populations.
[0005] Current body composition measurement methods can be roughly divided into two categories: low-cost technologies with limited accuracy and high-precision but expensive technologies, and each of these two types of technologies has significant limitations. Among them, low-precision technologies include Body Calipers and Bioelectrical Impedance Analysis (BIA). Specifically, the Body Calipers method estimates body fat percentage by measuring subcutaneous fat thickness, but it is limited by the experience of the technology user and has limited accuracy. BIA estimates body fat percentage by measuring the resistance value of the human body to a weak current. However, its results are easily affected by water intake, exercise status, and environmental temperature, resulting in unstable data and large errors. High-precision technologies include Air Displacement Plethysmograph (ADP) and Dual-Energy X-ray Absorptiometry (DXA or DEXA). Specifically, ADP uses whole-body density measurement to calculate body composition (the ratio of fat to lean body mass). Although the results are very accurate, it requires expensive equipment and the user needs to measure in a specific enclosed environment, with complex operations. DXA uses low-dose X-rays to measure body composition (such as body fat and bone density). Although DXA technology is considered very accurate, its equipment is expensive and requires trained technicians to operate.
[0006] In summary, although the BMI index is widely used in body composition assessment due to its simplicity, it has significant limitations in reflecting fat content and related indicators. Currently, there are already many high-precision body composition measurement methods that rely on expensive equipment and complex technologies. However, how to develop alternative solutions that are economical, accurate, and convenient remains an important direction for future research. Summary of the Invention
[0007] Aiming at the deficiencies of the above-mentioned existing technologies, the present invention aims to propose a simple, convenient, and low-cost body composition measurement method.
[0008] The body composition measurement method proposed by the present invention uses digital human measurement phenotypes to establish a machine learning model for predicting body composition related to fat content / muscle content, and can obtain accurate body composition data related to fat content / muscle content without complex measurement procedures; the specific steps are as follows:
[0009] (1) Digital human measurement phenotype measurement and body composition measurement.
[0010] The data used in the present invention includes body composition data related to fat content / muscle content and digital anthropometric phenotype data. Among them, the body composition data is collected by DXA, including but not limited to total body fat mass (FM, Fat Mass), total body muscle mass (LM, Lean Mass), fat mass index (FMI, Fat Mass Index), Android to Gynoid ratio A / G, visceral adipose tissue mass VAT (Visceral adipose tissue), etc.; the digital anthropometric phenotype data is collected by a 3D scanner or three-dimensionally reconstructed after 2D image collection, including but not limited to gender, height, weight, age, chest circumference, abdominal circumference, neck circumference, left and right arm circumferences, lower limb circumferences, head height, back width, waist-to-hip ratio, etc.
[0011] Correspond each digital anthropometric phenotype data with the simultaneously measured body composition data, using the digital anthropometric phenotype data as input data and the body composition indicators related to fat content / muscle content as label data.
[0012] (II) Data preprocessing.
[0013] Preprocess the data collected in the first step, mainly including: handling missing values, removing abnormal data, data standardization, normalization, etc.
[0014] (1) Handling missing values: Delete samples with missing data or fill in the missing data.
[0015] (2) Removing abnormal data: Detect data outliers and delete or replace them to smooth the data.
[0016] (3) Data standardization processing: To eliminate the influence of dimension differences on model parameters, it is necessary to standardize the data in advance.
[0017] (4) Data transformation: Convert the data into a form suitable for data modeling through methods such as smoothing aggregation, data generalization, and normalization. This includes normalizing the data, feature selection, dimensionality reduction, etc.
[0018] (III) Constructing multiple machine learning models; including:
[0019] Use the body composition indicators related to fat content / muscle content as the dependent variable and the digital anthropometric phenotype data as the independent variable to construct different machine learning models;
[0020] (IV) Body composition calculation; the specific process is as follows:
[0021] (1) Construct a mathematical model. Through the input digital anthropometric data (such as gender, height, weight, age, chest circumference, waist circumference, hip circumference, etc.) and the established statistical model, predict the body composition related to fat content / muscle content.
[0022] (2) Conduct data mapping. The core of the body composition calculation formula related to fat content / muscle content depends on the training model, and the formula is:
[0023] Body Composition (BC) =
[0024] f (circumferences of various body parts, widths of various body parts, lengths of various body parts, weight, gender, age)
[0025] The function f is a machine learning model trained based on a large number of real data samples;
[0026] (3) Result output: Output the body composition result and provide suggestions according to the standard health guidelines;
[0027] (5) Optimize the machine learning model;
[0028] By changing the parameters of different machine learning models or the feature selection methods, optimize the model. The optimization criteria are the number of independent variables, the goodness of fit of the model, the root mean square error, etc. By adjusting the model parameters (such as the regularization coefficient, the depth of the tree, etc.) or applying the feature selection methods (such as recursive feature elimination, importance ranking screening), optimize the model. The optimization criteria include the goodness of fit of the model (such as R 2 , R Squared), the root mean square error (RMSE), the model complexity (the number of independent variables), and the generalization performance, etc., to ensure that the model has good interpretability and robustness while maintaining high prediction accuracy.
[0029] In the present invention, using the digital anthropometric phenotypes as independent variables and the body composition related to fat content / muscle content as the dependent variable, construct multiple body composition assessment models, and through integrated analysis of various models, a more accurate obesity degree assessment can be obtained.
[0030] In summary, compared with the prior art, the features and beneficial effects of the present invention are as follows:
[0031] (1) The present invention has high efficiency. Only by providing digital anthropometric data can complex body composition measurements be completed. The whole process takes an extremely short time (about 10 seconds to 30 seconds), realizing fast calculation and feedback.
[0032] (2) The present invention has universality: Without complex operation procedures, it supports multiple scenarios, such as home health management, daily monitoring by fitness coaches.
[0033] The present invention has high precision. Through machine learning technology, it predicts body composition related to fat content / muscle content from digital anthropometric phenotype data. Compared with traditional anthropometric indexes such as BMI, the accuracy of this method in obesity assessment is significantly improved, with high clinical application value and can provide more precise guidance. The precision of its measurement results is close to that of traditional high-end measurement technologies such as dual-energy X-ray absorptiometry (DXA). BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a flowchart of the body composition measurement method based on digital anthropometric phenotype of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0035] The present invention will be introduced in detail below with reference to the drawings and specific embodiments.
[0036] The present invention proposes a body composition measurement method based on digital anthropometric phenotype. The process is as Figure 1 shown, and the specific steps are as follows:
[0037] (1) Digital anthropometric phenotype measurement and body composition measurement.
[0038] The data used in the present invention includes body composition data related to fat content / muscle content and digital anthropometric phenotype data. Among them, the body composition data is collected by DXA, including but not limited to total body fat mass FM, fat mass index FMI, total body muscle mass LM, visceral adipose tissue VAT, Android / Gynoid (hereinafter abbreviated as A / G), etc.; the digital anthropometric phenotype data is collected by a 3D scanner, including but not limited to anthropometric phenotypes such as gender, height, weight, age, chest circumference, abdominal circumference, neck circumference, left and right arm circumferences, lower limb circumferences, head height, back width, waist-to-hip ratio, etc.
[0039] Correspond each digital anthropometric phenotype data with the simultaneously measured body composition data, use the digital anthropometric phenotype data as input data, and the body composition indexes related to fat content / muscle content as label data.
[0040] (2) Data preprocessing.
[0041] Preprocess the data collected in the first step, mainly including: handling missing values, removing abnormal data, data standardization, normalization, etc.
[0042] (1) Handling missing values: Delete samples with missing data or fill in the missing data.
[0043] (2) Removing abnormal data: Detect data outliers and delete or replace them to smooth the data.
[0044] (3) Data standardization: To eliminate the influence of dimensional differences on model parameters, data standardization needs to be carried out in advance.
[0045] (4) Data transformation: Transform the data into a form suitable for data modeling through methods such as smoothing aggregation, data generalization, and normalization. This includes normalizing the data, feature selection, dimensionality reduction, etc.
[0046] (III) Constructing multiple machine learning models; including:
[0047] Taking body composition indicators related to fat content / muscle content as the dependent variable and digital anthropometric phenotype data as the independent variable, including gender, height, weight, age, head height, neck height, neck circumference, hip height, hip circumference, abdominal height, abdominal circumference, etc., construct machine learning models, including but not limited to multiple linear regression models, Lasso regression models, ridge regression models, decision tree regression models, and random forest regression models.
[0048] Among them, body composition related to fat content / muscle content includes but not limited to total body fat mass (FM, Fat Mass), total body muscle mass (LM, Lean Mass), fat mass index (FMI, Fat Mass Index), Android to Gynoid ratio A / G, visceral fat mass (VAT, Visceral adipose tissue), etc., and digital anthropometric phenotype data includes basic information such as gender, height, weight, age, and other body measurement values, such as distances, circumferences, widths of various body parts;
[0049] (IV) Body composition calculation; the specific process is as follows:
[0050] (1) Construct a mathematical model to predict body composition related to fat content / muscle content through the input digital anthropometric data (such as gender, height, weight, age, chest circumference, waist circumference, hip circumference, etc.) and the established statistical model.
[0051] (2) Conduct data mapping. The core of the body composition calculation formula related to fat content / muscle content depends on the training model, and the formula is:
[0052] Body composition (BC, Body Composition) =
[0053] f (circumferences of various body parts, widths of various body parts, lengths of various body parts, weight, gender, age)
[0054] The function f is a machine learning model trained based on a large number of real data samples (e.g., multiple linear regression model, Lasso regression model, ridge regression model, decision tree regression model, random forest regression model, etc.); among them, body composition includes but is not limited to total body fat mass FM, total body muscle mass LM, fat mass index FMI, Android to Gynoid ratio A / G, visceral fat mass VAT, etc.; digital anthropometric data includes but is not limited to gender, height, weight, age, head height, neck height, head circumference, transverse shoulder width, chest circumference, hip circumference, thigh circumference, calf circumference, abdominal circumference, arm length, arm circumference, leg length, etc.
[0055] Example formula:
[0056] Body composition = a·waist circumference + b·hip circumference + c·height + d·age + e
[0057] Among them, a, b, c, d, and e are regression coefficients in the model, obtained from the training data.
[0058] (3) Result output: Output the body composition result and provide suggestions according to the standard health guidelines, for example: "The output result is the body fat percentage, and the normal body fat range is 15%-25% (for men) or 20%-30% (for women)."
[0059] (5) Optimize the machine learning model;
[0060] By changing the parameters of different machine learning models or the feature selection method, optimize the model. The optimization criteria are the number of independent variables, the goodness of fit of the model, the root mean square error, etc. Optimize the model by adjusting the model parameters (such as the regularization coefficient, the depth of the tree, etc.) or applying the feature selection method (such as recursive feature elimination, importance ranking screening). The optimization criteria include the goodness of fit of the model (such as R 2 )), root mean square error (RMSE), model complexity (number of independent variables), and generalization performance, etc., to ensure that the model has good interpretability and robustness while maintaining high prediction accuracy.
[0061] In a specific embodiment, use the data of the healthy population cohort to train the machine learning model. Using the digital anthropometric phenotype data as input, 157 kinds of digital anthropometric data are obtained, and the models are established with FM, FMI, LM, VAT, and A / G as labels respectively. Some test results of the models are shown as follows:
[0062] Table 1 Construct a machine learning model for all cohort data
[0063]
[0064]
[0065] Note: FM, total body fat content; FMI, fat mass index; LM, total body muscle content; VAT, visceral fat mass; A / G, Android fat mass / Gynoid fat mass; R 2 , goodness of fit; RMSE, root mean square error.
[0066] As can be seen from Table 1, the R 2 values of FM, FMI, LM, VAT, and A / G predicted by the multiple linear regression model, Lasso regression model, Ridge regression model, decision tree regression model, and random forest regression model are mostly close to 1 (for example, in the regression model for predicting FM, the R 2 of the multiple linear regression model is 0.89, the R 2 of the Lasso regression model is 0.90, the R 2 of the Ridge regression model is 0.91, and the R 2 of the random forest regression model is 0.87), that is, the goodness of fit is relatively high, indicating that the model constructed in this patent has a good fitting effect; in addition, the root mean square errors (RMSE) generated by all regression models are within the range of 0.29 - 0.43 (for example, in the regression model for predicting FM, the RMSE of the multiple linear regression model is 0.34, the RMSE of the Lasso regression model is 0.31, the RMSE of the Ridge regression model is 0.31, and the RMSE of the random forest regression model is 0.36), that is, the error is small, indicating that the model constructed in this patent has a high prediction accuracy.
[0067] Compared with traditional body composition measurement devices, the body composition measurement method proposed by the present invention can greatly improve the utilization rate of digital human measurement data, without complex operation procedures and personnel training, greatly reducing the time cost and economic cost. At the same time, it provides a more efficient, more economical, and more universal non-invasive body composition measurement method, and also shows another possibility for the simple application of human measurement data in daily life or clinical practice.
[0068] References
[0069] [1] Gómez-Ambrosi J, Silva C, Galofré J C, et al. Body adiposity and type 2 diabetes: increased risk with a high body fat percentage even having a normal BMI[J]. Obesity, 2011, 19(7): 1439 - 1444.
[0070] [2]Freisling H, Arnold M, Soerjomataram I, et al. Comparison of general obesity and measures of body fat distribution in older adults in relation to cancer risk: meta-analysis of individual participant data of seven prospective cohorts in Europe[J]. British journal of cancer, 2017, 116(11): 1486-1497.
[0071] [3]Srikanthan P, Horwich T B, Tseng C H. Relation of muscle mass and fat mass to cardiovascular disease mortality[J]. The American journal of cardiology, 2016, 117(8): 1355-1360.
[0072] [4]INA F, SSBE A. BEYOND BMI: HOW TO REDEFINE OBESITY[J]. Nature, 2023, 622: 12.
[0073] [5]Pray R, Riskin S. The history and faults of the body mass index and where to look next: A literature review[J]. Cureus, 2023, 15(11).
Claims
1. A method for predicting body composition based on digital anthropometric phenotypes, characterized in that: Using digital anthropometric phenotypes, a machine learning model for predicting body composition related to fat content / muscle content is established, which can obtain accurate body composition data related to fat content / muscle content; the specific steps are as follows: (i) Collection of digital anthropometric phenotypic data and body composition data; The data used include body composition data related to fat content / muscle content and digital anthropometric phenotypic data; among them, body composition data are collected by DXA, including total body fat mass (FM), total body muscle mass (LM), fat mass index (FMI), Android to Gynoid ratio (A / G), visceral fat mass (VAT); digital anthropometric phenotypic data are collected by 3D scanner or 3D reconstruction after 2D image collection, including gender, height, weight, age, chest circumference, abdominal circumference, neck circumference, left and right arm circumference, lower limb circumference, head height, back width, waist-to-hip ratio; Each digital anthropometric phenotype data is matched with the body composition data measured at the same time, and the digital anthropometric phenotype data is used as input data and the body composition index is used as label data; (ii) data preprocessing; The data collected in step (i) are preprocessed, including: processing missing values, removing abnormal data, data standardization, and normalization; wherein: (1) Handling missing values: deleting samples with missing data or filling in missing data; (2) Clear abnormal data: detect data outliers and delete or replace them to smooth the data; (3) Data standardization: To eliminate the impact of dimensional differences on model parameters, the data is standardized; (4) Data transformation: transforming data into a form suitable for data modeling through smooth aggregation, data generalization, and normalization; including data normalization, feature selection, and dimensionality reduction; (III) Build multiple machine learning models, including: A machine learning model was constructed using body composition indices as dependent variables and digital anthropometric phenotype data as independent variables; (IV) Body composition calculation; the specific process is as follows: (1) Constructing a mathematical model to predict body composition related to fat content / muscle content using input digital anthropometric data and established statistical models; (2) Perform data mapping. The core of the body composition calculation formula related to fat content / muscle content depends on the training model. The formula is: Body composition (BC) f(body circumference, body width, body length, weight, gender, age); Function f is a machine learning model trained based on a large number of real data samples; (3) Result output: Output body composition results and provide recommendations based on standard health guidelines; (V) Optimizing machine learning models; The model is optimized by adjusting model parameters or applying feature selection methods; the optimization criteria include the number of independent variables, goodness of fit, root mean square error (RMSE), model complexity, and generalization performance of the model to ensure that the model has good interpretability and robustness while maintaining high prediction accuracy.
2. A method for predicting body composition based on digital anthropometric phenotype according to claim 1, characterized in that: The machine learning model is a multiple linear regression model, a Lasso regression model, a ridge regression model, a decision tree regression model or a random forest regression model.