Children obesity risk prediction and grading intervention management system based on machine learning
By using a machine learning-based system and combining multi-dimensional health data, a composite risk level matrix of childhood obesity risk is generated, and personalized intervention paths are generated based on the levels. This solves the problems of insufficient predictive performance and lack of precision in intervention in existing technologies, and improves the effectiveness of childhood obesity prevention and control.
Patent Information
- Application Number
- CN202511404616.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2025-11-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies are inadequate in predicting the risk of childhood obesity, lack precise and personalized interventions, and fail to fully utilize longitudinal patterns of children's health data.
The system employs a machine learning-based approach, which combines multi-dimensional health data with a data acquisition module, a feature engineering module, a risk prediction module, and a tiered intervention module. It utilizes gradient boosting trees and attention networks to generate a composite risk level matrix and then generates personalized tree-like intervention paths based on the risk levels.
It enables accurate prediction and personalized intervention of childhood obesity risk, improves prevention and control effectiveness, and provides a scientific basis for early intervention.
Smart Images

Figure CN120932897A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of children's health management technology, and in particular to a machine learning-based system for predicting and grading the risk of childhood obesity. Background Technology
[0002] Early identification of children at high risk of obesity and timely targeted intervention are crucial for improving children's health and preventing chronic diseases in adulthood. Currently, while some potential modifiable risk factors have been identified, including high maternal pre-pregnancy body mass index, excessive weight gain during pregnancy, and high birth weight, their predictive performance when considering all these factors comprehensively lacks in-depth research. Furthermore, most existing predictive tools rely primarily on traditional regression methods, selecting only a few features and failing to fully utilize the longitudinal patterns of children's data.
[0003] Machine learning, as a powerful data-driven technology, can handle high-dimensional and complex data, uncovering potential patterns and correlations to build more accurate predictive models. However, research and practice on fully applying machine learning technology to predict childhood obesity risk and combining it with comprehensive children's health data for in-depth analysis are still relatively limited. In terms of intervention and management of childhood obesity, existing intervention measures often lack precision and personalization, failing to tier interventions based on the specific risk levels of different children, resulting in inconsistent intervention outcomes.
[0004] Therefore, developing a machine learning-based system for predicting and tiered interventions for childhood obesity risk, making full use of health data from multiple dimensions to achieve accurate prediction of childhood obesity risk, and implementing personalized tiered interventions based on risk levels, is of great practical significance and urgent need for the effective prevention and control of childhood obesity. Summary of the Invention
[0005] This invention provides a machine learning-based system for predicting and tiered intervention in childhood obesity risk, comprising: a data acquisition module, a feature engineering module, a risk prediction module, a tiered intervention module, and an intervention guidance module. The data acquisition module is used to collect children's medical and health data, behavioral time-series data, maternal-fetal tracing data, and dynamically supplemented data; The feature engineering module is used to extract children's static features, growth features, and derived features from the four types of data collected. The risk prediction module has a built-in hierarchical obesity risk prediction model, which is used to output a composite risk level matrix based on the extracted features. The tiered intervention module has a built-in standardized intervention atom library, which is used to generate tree-like intervention paths based on the composite risk level matrix; The intervention guidance module is used to guide parents in a visual way to implement the intervention plan at each node along the path.
[0006] The aforementioned machine learning-based childhood obesity risk prediction and tiered intervention management system uses static features extracted from medical and health data and maternal-fetal tracing data to provide a benchmark for risk prediction; its extraction logic is as follows: Calculate the deviation of various indicators in medical and health data and maternal-fetal tracing data from the baseline of the same age / sex group; The coefficient of synergy between the deviation of fetal period indicators and the deviation of childhood period indicators; The current BIM value of the child, the deviation calculation result, and the coordination coefficient are used as three components to generate static features.
[0007] The aforementioned machine learning-based childhood obesity risk prediction and tiered intervention management system uses growth characteristics extracted from children's longitudinal growth data (height and weight) to capture risk signals during the growth process. The extraction logic is as follows: Calculate the growth rate of height and weight at the beginning and end of different growth stages; Identify the key inflection points in the height and weight growth curves within each growth stage; Growth characteristics are calculated based on the rate of increase in height and weight and key inflection points.
[0008] The machine learning-based childhood obesity risk prediction and hierarchical intervention management system described above includes a hierarchical obesity risk prediction model that specifically comprises: a basic risk prediction layer, a time-series dynamic adjustment layer, and a composite risk aggregation layer. The basic risk prediction layer is used to output the child's current basic obesity risk based on the extracted static features; The time-series dynamic adjustment layer is used to dynamically correct the basic risks by combining the extracted growth features; The composite risk aggregation layer is used to combine derived features to output a composite risk level matrix that includes short-term, medium-term, and long-term dimensions.
[0009] The machine learning-based childhood obesity risk prediction and tiered intervention management system described above includes, in particular, a tiered intervention module comprising: a risk level analysis submodule, an intervention atomic matching submodule, and an intervention path generation submodule. The risk level analysis submodule is used to combine the extracted static features, growth features, and derived features to locate the core risk factors and output structured risk analysis results. The intervention atom matching submodule is used to select suitable intervention atoms from the standardized intervention atom library based on the structured risk analysis results. The intervention path generation submodule is used to organize the selected intervention atoms into a tree-like intervention path.
[0010] The beneficial effects achieved by this invention are as follows: Through multi-dimensional data collection and feature engineering, combined with a hierarchical machine learning model, accurate prediction of childhood obesity risk is achieved, outputting a short-term, medium-term, and long-term composite risk matrix. Personalized tree-structured intervention paths are generated based on a standardized intervention atomic library and visualized through an intervention guidance module, addressing the problems of insufficient predictive performance and lack of precision and personalization in existing technologies, thereby improving the effectiveness of childhood obesity prevention and control and providing a scientific basis for early intervention. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0012] Figure 1 This is a schematic diagram of a machine learning-based childhood obesity risk prediction and graded intervention management system provided in Embodiment 1 of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] Example 1
[0015] like Figure 1 As shown, Embodiment 1 of this application provides a machine learning-based childhood obesity risk prediction and graded intervention management system, including: a data acquisition module 11, a feature engineering module 12, a risk prediction module 13, a graded intervention module 14, and an intervention guidance module 15; The data acquisition module 11 is used to collect children's medical and health data, behavioral time-series data, maternal-fetal tracing data, and dynamic supplementary data; Medical and health data includes static basic data (routine physical examination indicators such as height, weight, BMI, and blood pressure) and dynamic clinical data (blood routine, endocrine indicators, and gut microbiota sequencing data); behavioral time-series data includes wearable device data (exercise calorie consumption and sleep structure) and dietary data (identifying the dietary structure at meals through AI vision terminals and automatically converting it into nutritional composition data through image recognition algorithms); maternal-fetal tracing data includes maternal pregnancy data (pre-pregnancy BMI and gestational weight gain curve) and fetal data (fetal growth indicators such as abdominal circumference and femur length from ultrasound examinations); and dynamic supplementary data includes short-term health events such as colds and fevers. These four types of data are standardized and stored in a local database.
[0016] Feature engineering module 12 is used to extract children's static features, growth features and derived features from the four types of data collected; Static features, extracted from medical and health data and maternal-fetal tracing data, are used to provide a benchmark for risk prediction; their extraction logic is as follows: ① Calculate the deviation of various indicators in medical and health data and maternal-fetal tracing data from the baseline of the same age / sex group. The formula for calculating the deviation is as follows: Where D is the calculated deviation, i takes values from 1 to n, and n is the total number of indicators included in the medical and health data and maternal-fetal tracing data. Let i be the value of the current child's i-th indicator. The median of the i-th indicator for the same age / gender group. represents the standard deviation of the i-th indicator for the same age / gender group; ② Quantify the correlation coefficient between the deviation of fetal period indicators and the deviation of childhood period indicators; Substituting the current fetal data and static baseline data into the above deviation calculation formula yields the following results. and Then and Substituting into the quantization formula: From this, the coefficient of coordination C between the deviation of fetal period indicators and the deviation of childhood period indicators can be obtained, where It is a smoothing factor (to avoid a denominator of 0); ③ Generate static features by taking the current child's BIM value, deviation calculation result D, and coordination coefficient C as three components; Note that the three components need to be uniformly mapped to the [0,1] interval (using min-max standardization) to eliminate dimensional differences.
[0017] Growth characteristics, extracted from children's longitudinal growth data of height and weight, are used to capture risk signals during the growth process. The extraction logic is as follows: ① Calculate the growth rate of height and weight at the beginning and end of different growth stages; The growth stages include two phases: infancy (0-12 months) and early childhood (12-24 months). For each phase, the average monthly growth rate of height and weight in the first three months and the last three months is calculated, with units of kg / month and cm / month, respectively, forming a height growth rate sequence and a weight growth rate sequence.
[0018] ② Identify the key inflection points in the height and weight growth curves within each growth stage; There are many existing algorithms for identifying inflection points, such as derivative-based algorithms, Pelt algorithms, and CSS algorithms. We will not limit ourselves to any of these algorithms. We will arrange the identified key inflection points in chronological order to obtain the key inflection point sequences for height and weight.
[0019] ③ Calculate growth characteristics based on the rate of increase in height and weight and key inflection points; Substitute the height growth rate sequence, the height key inflection point sequence, the weight growth rate sequence, and the weight key inflection point sequence into the growth characteristic component calculation formula respectively: In the process, two components of growth characteristics are obtained: height and weight. For the calculated components, For adjustable weights, , Let represent the growth rates of the indicators at the beginning and end of the s-th growth stage, respectively. This represents the total duration of the s-th growth stage, in months, where s ranges from 1 to k, and k is the number of growth stages set. Indicates the first The key inflection point and the first Rate of change of acceleration between key inflection points The value ranges from 1 to L, where L is the total number of critical inflection points.
[0020] Derived features, extracted from behavioral time-series data and dynamically supplemented data, are used to reflect multi-dimensional behavioral interaction risks and quantify the impact of lifestyle on obesity. Their extraction formula is expressed as: ,in For the extracted derived features, The coupling strength coefficient is adjustable. This is a behavioral pattern index. , These are the weekly exercise calorie expenditure sequence, dietary structure sequence, and sleep structure sequence, respectively. Return to the exponential moving average of exercise calorie expenditure. The entropy value is the sequence of dietary structures (the higher the proportion of high sugar / high fat, the closer the entropy value is to 1). The entropy value is the sequence of sleep structures (the more times you wake up at night, the closer the entropy value is to 1). As a health event early warning index, , This represents the number of health events experienced by children over the past month. This represents the average number of health events per month for the same age and gender group.
[0021] Risk prediction module 13 has a built-in hierarchical obesity risk prediction model, which is used to output a composite risk level matrix based on the extracted features. The hierarchical obesity risk prediction model specifically includes: a basic risk prediction layer, a time-series dynamic adjustment layer, and a composite risk aggregation layer; 1. Basic risk prediction layer, used to output the child's current basic obesity risk based on the extracted static features; The static features are typical structured data (numerical components) and there is a "non-linear relationship between the synergy coefficient C and BMI" (e.g., the risk of children with high BMI and high synergy coefficient increases exponentially). Gradient boosting trees are good at handling non-linear relationships in structured data and are more robust to small sample data than deep learning models. They also support feature importance output. Therefore, gradient boosting trees are used as the base model for this layer.
[0022] On the other hand, to address the issue of "uneven contribution" among static feature components—for example, the "maternal-fetal co-occurrence coefficient C" has a higher predictive value for obesity risk than BMI alone—it is necessary to dynamically amplify the influence of key components through an attention network. This involves dynamically assigning weights to each component in the static features, weighting the sum, and then feeding it to the base model. The base model learns the mapping relationship between features and risk through multiple rounds of gradient boosting, outputting a basic risk score S ranging from [0,1]. base .
[0023] 2. Temporal dynamic adjustment layer, used to dynamically correct the basic risk by combining the extracted growth characteristics; This layer uses the formula. The basic risks are dynamically adjusted, among which The revised risk score, It is the Sigmoid function. These are the extracted growth characteristics. yes The scaling factor, It is the offset of the sigmoid function. It is the limiting factor for the correction range.
[0024] 3. Composite risk aggregation layer, used to combine derived features to output a composite risk level matrix containing short-term, medium-term, and long-term dimensions; This layer learns different time dimensions (short-term=1, medium-term=2, long-term=3) through an attention network. and derived features The association weights are used to obtain the weight matrix W, which is represented as follows: Then, the fusion score S is calculated for the short-term, medium-term, and long-term dimensions respectively. t ,Right now Where t is the time dimension index, , They are respectively under the time dimension t , The association weight.
[0025] The combined scores S from the short-term, medium-term, and long-term dimensions are respectively... t Mapped to risk levels (mapping rules are set by the user as needed), and then processed to obtain a composite risk level matrix. .
[0026] The graded intervention module 14 has a built-in standardized intervention atom library, which is used to generate tree-like intervention paths based on the composite risk level matrix; specifically, it includes: risk level parsing submodule, intervention atom matching submodule, and intervention path generation submodule. The standardized atom library stores multi-dimensional intervention atoms, which are classified according to intervention type, applicable risk level, and core factors. For example, dietary atom 2: dietary fiber supplementation ≥5 times per week (applicable to medium-term low risk, dietary imbalance); exercise atom 1: moderate-intensity exercise ≥30 minutes per day (applicable to short-term high risk, insufficient exercise).
[0027] 1. Risk level analysis submodule, which is used to locate core risk factors by combining extracted static features, growth features, and derived features, and output structured risk analysis results; The identification of core risk factors is achieved by setting identification rules. For example, if the dietary entropy value H(E2) in the derived features is greater than a threshold, then "dietary imbalance" is returned; the exponential moving average of exercise calorie expenditure is also used. If the threshold is exceeded, "insufficient exercise" is returned.
[0028] The output of structured risk analysis results includes: time dimension (short-term / medium-term / long-term), risk level (low / medium / high), and core risk factors (dietary imbalance, lack of exercise).
[0029] 2. The intervention atom matching submodule is used to select suitable intervention atoms from the standardized intervention atom library based on the structured risk analysis results; A three-dimensional index is constructed based on "time dimension + risk level + core factors" to retrieve the intervention atom library and output the candidate intervention atom set.
[0030] 3. The intervention path generation submodule is used to organize the selected intervention atoms into a tree-like intervention path; Using the timeline (short-term → medium-term → long-term) as the trunk, candidate intervention atoms are assigned to corresponding nodes according to time logic and progressive relationship to generate a tree-like intervention path.
[0031] Intervention guidance module 15 is used to guide parents in a visual way to implement the intervention plan at each node along the path; The tree-like intervention path is transformed into a visual interface that is easy for parents to understand, such as a timeline + node card display. The nodes are displayed according to the timeline of "short-term → medium-term → long-term". Each node card is marked with the intervention plan (such as "30 minutes of exercise per day") and the execution time. Parents can click to view details (such as example videos of exercise types). The execution time can be modified by parents as needed.
[0032] Example 2
[0033] Embodiment 2 of this application provides a method for predicting and tiered intervention management of childhood obesity risk based on machine learning, including: Step S210: Collect children's medical and health data, behavioral time-series data, maternal-fetal tracing data, and dynamic supplementary data; Step S220: Extract the child's static characteristics, growth characteristics, and derived characteristics from the four types of data collected; Step S230: Using a hierarchical obesity risk prediction model, output a composite risk level matrix based on the extracted features; Step S240: Based on the standardized intervention atomic library, generate a tree-like intervention path according to the composite risk level matrix; Step S50: Guide parents visually to implement the intervention plan at each node along the path.
[0034] Corresponding to the above embodiments, the present invention provides a computer storage medium, including: at least one memory and at least one processor; The memory is used to store one or more program instructions; A processor for running one or more program instructions to execute a machine learning-based method for predicting and managing childhood obesity risk.
[0035] Corresponding to the above embodiments, this embodiment of the invention provides a computer-readable storage medium containing one or more program instructions, which are executed by a processor to provide a machine learning-based method for predicting and managing the risk of childhood obesity.
[0036] The embodiments disclosed in this invention provide a computer-readable storage medium storing computer program instructions. When the computer program instructions are executed on a computer, the computer performs the aforementioned method for predicting and tiered intervention management of childhood obesity risk based on machine learning.
[0037] In this embodiment of the invention, the processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0038] The various methods, steps, and logic diagrams disclosed in the embodiments of this invention can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The processor reads information from the storage medium and, in conjunction with its hardware, completes the steps of the above methods.
[0039] The storage medium can be memory, such as volatile memory or non-volatile memory, or may include both volatile and non-volatile memory.
[0040] Among them, non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory.
[0041] Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (Synchlink DRAM, SLDRAM), and direct memory bus RAM (DRRAM).
[0042] The storage media described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memory.
[0043] Those skilled in the art will recognize that, in one or more of the examples above, the functions described in this invention can be implemented using a combination of hardware and software. When applied as software, the corresponding functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transmission of computer programs from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0044] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solution of the present invention should be included within the scope of protection of the present invention.
Claims
1. A machine learning-based system for predicting and tiered intervention in childhood obesity risk, characterized in that, include: The module includes a data acquisition module, a feature engineering module, a risk prediction module, a tiered intervention module, and an intervention guidance module. The data acquisition module is used to collect children's medical and health data, behavioral time-series data, maternal-fetal tracing data, and dynamically supplemented data; The feature engineering module is used to extract children's static features, growth features, and derived features from the four types of data collected. The risk prediction module has a built-in hierarchical obesity risk prediction model, which is used to output a composite risk level matrix based on the extracted features. The tiered intervention module has a built-in standardized intervention atom library, which is used to generate tree-like intervention paths based on the composite risk level matrix; The intervention guidance module is used to guide parents in a visual way to implement the intervention plan at each node along the path.
2. The machine learning-based childhood obesity risk prediction and graded intervention management system according to claim 1, characterized in that, Static features, extracted from medical and health data and maternal-fetal tracing data, are used to provide a benchmark for risk prediction; their extraction logic is as follows: Calculate the deviation of various indicators in medical and health data and maternal-fetal tracing data from the benchmark of the same age and gender group; The coefficient of synergy between the deviation of fetal period indicators and the deviation of childhood period indicators; The current BIM value of the child, the deviation calculation result, and the coordination coefficient are used as three components to generate static features.
3. The machine learning-based childhood obesity risk prediction and tiered intervention management system according to claim 1, characterized in that, Growth characteristics, extracted from children's longitudinal growth data of height and weight, are used to capture risk signals during the growth process. The extraction logic is as follows: Calculate the growth rate of height and weight at the beginning and end of different growth stages; Identify the key inflection points in the height and weight growth curves within each growth stage; Growth characteristics are calculated based on the rate of increase in height and weight and key inflection points; Substitute the height growth rate sequence, the height key inflection point sequence, the weight growth rate sequence, and the weight key inflection point sequence into the growth characteristic component calculation formula respectively: In the process, two components of growth characteristics are obtained: height and weight. For the calculated components, For adjustable weights, , Let represent the growth rates of the indicators at the beginning and end of the s-th growth stage, respectively. This represents the total duration of the s-th growth stage, in months, where s ranges from 1 to k, and k is the number of growth stages set. Indicates the first The key inflection point and the first Rate of change of acceleration between key inflection points The value ranges from 1 to L, where L is the total number of critical inflection points.
4. The machine learning-based childhood obesity risk prediction and graded intervention management system according to claim 1, characterized in that, The hierarchical obesity risk prediction model specifically includes: a basic risk prediction layer, a time-series dynamic adjustment layer, and a composite risk aggregation layer; The basic risk prediction layer is used to output the child's current basic obesity risk based on the extracted static features; The time-series dynamic adjustment layer is used to dynamically correct the basic risks by combining the extracted growth features; The composite risk aggregation layer is used to combine derived features to output a composite risk level matrix that includes short-term, medium-term, and long-term dimensions.
5. The machine learning-based childhood obesity risk prediction and graded intervention management system according to claim 4, characterized in that, The standardized atom library stores multi-dimensional intervention atoms, categorized by intervention type, applicable risk level, and core factors.
6. The machine learning-based childhood obesity risk prediction and graded intervention management system according to claim 5, characterized in that, The tiered intervention module specifically includes: a risk level analysis submodule, an intervention atomic matching submodule, and an intervention path generation submodule; The risk level analysis submodule is used to combine the extracted static features, growth features, and derived features to locate the core risk factors and output structured risk analysis results. The intervention atom matching submodule is used to select suitable intervention atoms from the standardized intervention atom library based on the structured risk analysis results. The intervention path generation submodule is used to organize the selected intervention atoms into a tree-like intervention path.