A method and device system for constructing a childhood obesity risk assessment model

By constructing a childhood obesity risk assessment model based on multi-gene expression, screening key genes and determining weight coefficients, the problem of early identification of the risk of metabolic complications from childhood obesity was solved, achieving high-precision risk assessment and scientific decision support.

CN121922380BActive Publication Date: 2026-07-17SICHUAN ACADEMY OF MEDICAL SCI SICHUAN PROVINCIAL PEOPLES HOSPITAL

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN ACADEMY OF MEDICAL SCI SICHUAN PROVINCIAL PEOPLES HOSPITAL
Filing Date
2026-03-24
Publication Date
2026-07-17

Smart Images

  • Figure CN121922380B_ABST
    Figure CN121922380B_ABST
Patent Text Reader

Abstract

This invention relates to the field of bioinformatics technology, specifically to a method and apparatus system for constructing a childhood obesity risk assessment model. The method includes: S1, acquiring a childhood obesity dataset, collecting sets of genes related to lipid metabolism and glucose metabolism, and taking their intersection; S2, performing differential analysis on the childhood obesity dataset using the R package limma; setting thresholds for differentially expressed genes and identifying them; S3, taking the intersection of the data obtained in S2 and S1 to obtain sets of differentially expressed genes related to lipid metabolism and glucose metabolism; S4, screening key genes using machine learning algorithms and determining their weight coefficients; and constructing a risk scoring model. The key genes include APOE, ADM, and MAPK14 genes. This invention, by screening relevant genes and fusing multiple data sets, achieves scientific processing of core related gene expression data, helping doctors extract useful information from large amounts of data and assisting in clinical work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of bioinformatics and medical data analysis, and in particular to a method and device system for constructing a childhood obesity risk assessment model, which can construct a childhood obesity assessment model based on the expression levels of multiple genes to assist clinicians in their work. Background Technology

[0002] Childhood obesity has become a major global public health problem, with its prevalence rising dramatically over the past few decades. Currently, hundreds of millions of children worldwide (including those under 5 years old) are threatened by overweight and obesity. Approximately 80% of childhood obesity persists into adulthood, significantly increasing long-term health risks. Obesity in childhood is not only closely related to impaired fasting glucose, impaired glucose tolerance, and insulin resistance, but may also trigger early abnormalities in glucose and lipid metabolism, increasing the risk of diabetes, hyperlipidemia, and cardiovascular disease, leading to metabolic syndrome, and further promoting the development of type 2 diabetes, cardiovascular disease, and even cancer. Notably, these microvascular and macrovascular-related metabolic abnormalities often begin in childhood, making individuals more susceptible to the early onset and accelerated progression of various complications in adulthood.

[0003] In this context, early identification and effective intervention of high-risk individuals are particularly crucial. While body mass index (BMI) and other metrics are widely used for screening childhood obesity, they cannot accurately predict the risk of long-term complications. Not all obesity is accompanied by metabolic abnormalities; based on whether or not metabolic abnormalities are present, obesity is classified into metabolically normal obesity and metabolically abnormal obesity. Diagnosis of metabolically abnormal obesity often only occurs after metabolic complications have developed, resulting in a significant delay in the detection window and missing the golden window for early prediction and intervention. Therefore, obesity diagnosis based on BMI cannot reveal the intrinsic molecular subtype of obese children, nor can it accurately predict the risk of long-term complications, limiting the precision of prevention and intervention strategies. Secondly, without early molecular subtyping of obesity, it is difficult to guide targeted therapy, leading to poor treatment outcomes for childhood obesity, and the inability to achieve ideal results despite significant time and effort.

[0004] Traditional Chinese medicine emphasizes a proactive approach to health management, generally believing that prevention is better than cure. Modern preventive medicine views disease as a result of a disruption in the body's internal and external balance, highlighting the crucial role of health management. Patients often exhibit various signs before the onset of illness, such as a sub-healthy state. However, the lack of reliable methods for diagnosing childhood obesity leads to a deficiency in early proactive intervention, exacerbating the health problems of obese children.

[0005] There remains a lack of consensus on how to identify the risk of metabolic complications, particularly insulin resistance and related metabolic abnormalities, in early childhood. This highlights the urgent need to develop early and reliable obesity risk assessment tools. While modern medicine can obtain a large amount of patient diagnostic data through various methods, methods for fusing multiple data sets and processing them scientifically are still lacking. This makes it difficult for doctors to extract useful information from large amounts of data, especially when faced with significant data noise.

[0006] Therefore, there is an urgent need in this field for a high-precision, quantifiable risk assessment tool to enable early warning and precise prevention and control of childhood metabolic obesity. Summary of the Invention

[0007] The purpose of this invention is to overcome the problem of the lack of information processing methods for scientifically and accurately judging childhood obesity in the existing technology, and to provide a method and device system for constructing a childhood obesity risk assessment model.

[0008] In the first aspect, the present invention provides a novel method for constructing a risk assessment model for childhood obesity, thereby achieving the construction of a scientific and standardized scoring model and obtaining a model with accurate and reliable scores. This helps doctors better aggregate effective information from multiple aspects, eliminate noise interference, scientifically process the risk scores, and obtain more instructive intermediate score data to better assist doctors in their clinical work.

[0009] A method for constructing a childhood obesity risk assessment model includes the following steps:

[0010] S1. Obtain a dataset of children with obesity, including obese children and control samples.

[0011] Collect the sets of genes related to lipid metabolism and the sets of genes related to glucose metabolism, and take their intersection to obtain the sets of genes related to lipid metabolism and glucose metabolism.

[0012] S2. For the childhood obesity dataset, a statistical test method based on a linear model was used to test the significance of the difference in expression levels of each gene between the childhood obesity group and the control group, and to identify differentially expressed genes that meet the threshold.

[0013] The specific method for significance testing is as follows: set a threshold of differential expression fold > 0.5 and a corrected p-value of < 0.05 as the threshold for differentially expressed genes, and identify differentially expressed genes that meet the threshold.

[0014] S3. Compare and analyze the differentially expressed genes obtained in S2 with the lipid metabolism and glucose metabolism-related gene sets obtained in S1, and take the intersection to obtain the differentially expressed gene sets related to lipid metabolism and glucose metabolism.

[0015] S4. Based on the differentially expressed gene set related to lipid metabolism and glucose metabolism, key genes are screened using machine learning algorithms, and their weight coefficients are determined.

[0016] Based on the key genes and weight coefficients obtained through screening, a risk scoring model is constructed based on the key genes and weight coefficients.

[0017] This invention, based on a child obesity dataset, innovatively screens a set of genes related to lipid and glucose metabolism. Utilizing the differential analysis results of the child obesity dataset, it identifies differentially expressed gene sets. Through machine learning algorithms, key genes and weights are selected to construct a risk assessment model including APOE, ADM, and MAPK14 genes, enabling a scientific scoring technique for childhood obesity risk. The model's scoring provides objective, quantitative, and highly accurate guidance for risk assessment. The output scores offer scientific reference for clinical decision-making, helping doctors make better informed decisions regarding the need for further investigation or lifestyle interventions.

[0018] Among them, the collection of lipid metabolism-related gene sets and glucose metabolism-related gene sets, and taking the intersection means screening gene sets that simultaneously satisfy the requirements of lipid metabolism-related and glucose metabolism-related.

[0019] Furthermore, it also includes S5, which uses an independent validation dataset to validate the effectiveness of the constructed risk assessment model.

[0020] Furthermore, S4 also includes reference suggestions for the scoring results output based on the risk scoring model.

[0021] If the risk scoring model score is less than the first threshold, the first reference suggestion will be output.

[0022] If the risk scoring model score is greater than the first threshold, a second reference suggestion will be output.

[0023] If the score exceeds the second threshold, a third reference suggestion is output. This score-based suggestion provides guidance to users or healthcare professionals, offering more easily understood reference information than the score itself, thus helping them make informed and informed decisions.

[0024] Furthermore, in S1, the acquisition of the childhood obesity dataset includes at least one of the datasets GSE9624, GSE104815, and GSE139400.

[0025] Furthermore, in S1, the sets of lipid metabolism-related genes and glucose metabolism-related genes are retrieved and extracted from specialized gene function and disease association databases and / or peer-reviewed biomedical literature databases. In particular, large, widely recognized comprehensive or specialized biomedical databases, such as GeneCards, are preferred.

[0026] Preferably, the lipid metabolism-related gene set and the glucose metabolism-related gene set are obtained from the GeneCards database and / or the PubMed literature database.

[0027] Furthermore, in S1, after obtaining the children's obesity dataset, data preprocessing is performed, and batch effect correction algorithm based on the empirical Bayesian framework (specifically using the R package sva) is used to remove batch effects, resulting in the integrated GEO dataset, i.e., the integrated dataset.

[0028] Furthermore, the integrated dataset is standardized.

[0029] Preferably, the GEO integrated dataset is standardized.

[0030] Preferably, the difference in expression values ​​before and after standardization is compared using box plots to ensure that the dataset is a reliable data foundation for gene expression research related to childhood obesity, and to ensure the accuracy and reproducibility of the analysis.

[0031] Furthermore, in step S2, specifically, a linear model for comparison between the two groups is constructed, and the variance of the model estimate is robustly adjusted using the empirical Bayesian method to calculate the statistical significance (P-value) and fold change of the differences in gene expression.

[0032] Preferably, in S2, differential analysis is performed on the childhood obesity dataset using the R package limma (Linear Models for Microarray Data, which implements statistical testing methods for linear models); a threshold of |logFC|>0.5 and p<0.05 is set as the threshold for differentially expressed genes, and differentially expressed genes that meet the threshold are identified.

[0033] Furthermore, in S2, differential expression analysis is performed on the integrated dataset (such as the GEO integrated dataset). The threshold for differentially expressed genes is set as |logFC|>0.5 and p value<0.05, and differentially expressed genes that meet the threshold are identified.

[0034] Furthermore, in S4, a preliminary screening is first performed, and logistic regression analysis is conducted on the differentially expressed gene set related to lipid metabolism and glucose metabolism to screen out genes that are significantly associated with childhood obesity. Then, based on the genes obtained from the preliminary screening, key genes are further screened using the support vector machine algorithm.

[0035] Furthermore, the key genes were subjected to logistic regression analysis as binary variables; the binary variables were the grouping information of the obese children's samples and the control samples.

[0036] Furthermore, in S4, an SVM model is constructed using the support vector machine algorithm, and key genes are selected based on the principle of lowest error rate and highest accuracy.

[0037] Preferably, 3-6 key genes are selected. For example, 3, 4, or 5 key genes.

[0038] Furthermore, in S4, based on the key genes selected, minimum absolute shrinkage and selection LASSO regression analysis are performed to determine the weight coefficients of each key gene and form a risk scoring model.

[0039] Furthermore, in S4, the key genes include the APOE, ADM, and MAPK14 genes.

[0040] Furthermore, in S4, the risk score calculation formula of the risk scoring model is based on the selected gene GenX. i Scoring calculation:

[0041]

[0042] in,

[0043] RiskScore is a risk rating.

[0044] n is the selected GenX i The total number of items is 3 ≤ n ≤ 6.

[0045] GenX i This represents the expression level of the selected i-th gene.

[0046] w i The weight coefficient represents the i-th gene.

[0047] GenX i Including APOE, ADM, and MAPK14; APOE represents the expression level of the apolipoprotein E gene, ADM represents the expression level of the adrenal medullarin gene, and MAPK14 represents the expression level of the mitogen-activated protein kinase 14 gene.

[0048] use a , b , c The weight coefficients corresponding to genes APOE, ADM, and MAPK14 satisfy: -1.6 ≤ a ≤ -0.9, 0.6≤ b ≤ 1.3, -1.5 ≤ c ≤ -0.8.

[0049] Furthermore, in S4, the risk score calculation formula of the risk scoring model includes the score calculation of APOE, ADM, and MAPK14:

[0050] RiskScore = a × APOE + b × ADM + c × MAPK14

[0051] in,

[0052] RiskScore is a risk rating;

[0053] APOE represents the expression level of the apolipoprotein E gene;

[0054] ADM represents the expression level of the adrenomedullin gene;

[0055] MAPK14 represents the expression level of the mitogen-activated protein kinase 14 gene;

[0056] a , b , c Let be the weight coefficients corresponding to each gene, and satisfy: -1.6 ≤ a ≤ -0.9, 0.6 ≤ b ≤ 1.3, -1.5≤ c ≤ -0.8.

[0057] Preferably, if other genes are selected, the weighting coefficients are determined using the same method as above, and the results are summed to obtain the risk score calculation formula.

[0058] Furthermore, it also includes S5, the model validation step: in the integrated dataset and / or independent datasets, the diagnostic accuracy, calibration or clinical applicability of the risk scoring model is validated by at least one of the following methods: receiver operating characteristic (ROC) curve analysis, calibration curve analysis or decision curve analysis.

[0059] Furthermore, in S5, the effectiveness of the constructed risk assessment model is validated using an independent validation dataset.

[0060] In a second aspect, the present invention provides a novel device system for assessing the risk of childhood obesity.

[0061] A childhood obesity risk assessment device system, comprising:

[0062] The data acquisition module is used to acquire datasets on childhood obesity, sets of genes related to lipid metabolism, and sets of genes related to glucose metabolism.

[0063] The data preprocessing module is used to preprocess the data obtained by the data acquisition module;

[0064] The differentially expressed gene identification module is used to identify preprocessed data based on a threshold and output differentially expressed genes that meet the threshold.

[0065] The model building module is used to use machine learning algorithms to screen key genes from differentially expressed genes and determine weight coefficients to build a risk assessment model.

[0066] Furthermore, the aforementioned childhood obesity risk assessment device system is used to implement the aforementioned method for constructing a childhood obesity risk assessment model.

[0067] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0068] 1. This invention constructs a model based on an independently developed assessment method. The corresponding model can perform scientific scoring, and can objectively, quantitatively, and with high precision assess the risk of childhood obesity. It outputs scoring information that can provide scientific reference for clinical decision-making, helping doctors to make better scientific decisions and determine whether further examination or lifestyle intervention is needed. Therefore, the method for constructing a childhood obesity risk assessment model of this invention is of great significance for providing services to clinical medicine.

[0069] 2. This invention achieves scientific processing of core related gene expression data by screening relevant genes and fusing multiple data sources. It screens key genes such as APOE, ADM, and MAPK14 genes, and the resulting evaluation model can help doctors extract useful information from large amounts of data to assist in clinical work. Attached Figure Description

[0070] Figure 1 A flowchart for a comprehensive analysis of differentially expressed genes related to lipid and glucose metabolism in childhood obesity.

[0071] Figure 2 Remove correlation plots for batch effects in the GSE9624, GSE104815, and GSE139400 datasets.

[0072] Figure 3 This is a graph showing the correlation between differential gene expression analysis and other data.

[0073] Figure 4 Charts related to the diagnostic model for childhood obesity.

[0074] Figure 5 Charts and graphs related to the diagnosis and verification of childhood obesity.

[0075] Figure 6 Charts and graphs related to the diagnosis and verification of childhood obesity.

[0076] Figure 7 Figures related to differential expression validation and correlation analysis. Detailed Implementation

[0077] In addition, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described method for constructing a childhood obesity risk assessment model.

[0078] This invention provides a novel childhood obesity risk assessment model that achieves a scientific and standardized scoring model. This model helps doctors to scientifically process a large amount of complex information to obtain risk scores, aggregate effective information from multiple aspects, eliminate noise interference, and obtain more instructive intermediate scoring data to better assist doctors in their clinical work.

[0079] A childhood obesity risk assessment model obtains expression level data of three core genes from the same sample, wherein the three core genes are APOE gene, ADM gene and MAPK14 gene;

[0080] Risk scores were calculated based on expression level data of three core genes:

[0081] RiskScore = a × APOE + b × ADM + c × MAPK14

[0082] in,

[0083] RiskScore is a risk rating;

[0084] APOE represents the expression level of the apolipoprotein E gene;

[0085] ADM represents the expression level of the adrenomedullin gene;

[0086] MAPK14 represents the expression level of the mitogen-activated protein kinase 14 gene;

[0087] a, b, and c are the weight coefficients corresponding to each gene, and satisfy the following conditions: -1.6 ≤ a ≤ -0.9, 0.6 ≤ b ≤ 1.3, -1.5 ≤ c ≤ -0.8.

[0088] This invention presents a novel childhood obesity risk assessment model that integrates multi-center genomics data for the first time. It utilizes systematic bioinformatics for model analysis and employs a selected and validated combination of three core genes: APOE, ADM, and MAPK14, to construct a molecular diagnostic model for childhood obesity. The model calculates a risk score by quantifying the expression levels of these three genes. It enables scientific and accurate processing of various data, allowing physicians to better utilize large amounts of chemical and analytical data. The resulting score provides significant reference value for high-precision diagnosis and risk stratification of childhood obesity.

[0089] Furthermore, the weighting coefficients take the following values: -1.2 ≤ a ≤ -0.8, 0.6 ≤ b ≤ 1.1, -1.2 ≤ c ≤ -0.8.

[0090] Furthermore, the weighting coefficients are as follows: a = -0.994, b = 0.843, c = -0.98.

[0091] Furthermore, it also includes: comparing the risk score with a preset threshold; and based on the comparison result, outputting risk level information for the target individual for doctors' reference.

[0092] Furthermore, when the risk score exceeds the first threshold, a first warning message is output; when the risk score is less than the first threshold, a second prompt message is output.

[0093] When the weighting coefficients a=-0.994, b=0.843, and c=-0.980, the first threshold is -7.29.

[0094] When the RiskScore ≥ -7.29, the first warning message is output, indicating to the doctor that the subject is at high risk of childhood obesity or abnormal glucose and lipid metabolism; when the RiskScore < -7.29, it indicates to the doctor that the subject is not obese or at low risk of abnormal glucose and lipid metabolism.

[0095] In addition, the present invention provides a kit for the diagnosis or risk stratification of childhood obesity. The kit includes reagents for detecting the expression levels of the APOE gene, ADM gene, and MAPK14 gene, as well as instructions or a data processing unit for calculating a risk score based on the expression levels of the genes and outputting diagnostic or risk assessment results. The calculation of the risk score is based on the aforementioned risk assessment model.

[0096] This invention presents a risk scoring model for assessing childhood obesity, derived through bioinformatics analysis, screening, and validation of multicenter genomics data. By assessing multiple key factors contributing to childhood obesity, such as the risk of abnormal glucose and lipid metabolism, the model enables scientific and accurate processing of various data, allowing physicians to utilize large amounts of chemical and analytical data.

[0097] The present invention will now be described in further detail with reference to specific embodiments. However, this should not be construed as limiting the scope of the present invention to the following embodiments; all technologies implemented based on the content of the present invention fall within the scope of the present invention.

[0098] The following are explanations of some abbreviations used in this invention: CO (Childhood Obesity), GSEA (Gene Set Enrichment Analysis), GSVA (Gene Set Variation Analysis), DEGs (Differentially Expressed Genes), LMRGs (Lipid Metabolism-Related Genes), GMRGs (Glucose Metabolism-Related Genes), LGMRGs (Lipid Metabolism and Glucose Metabolism-Related Genes), and LGMRDEGs (Differentially Expressed Genes Related to Lipid and Glucose Metabolism). PCA (Principal Component Analysis). Combined Datasets: Integrated GEO datasets. ROC: Receiver Operating Characteristic Curve (ROC) constructed based on risk scores. ADM: Adrenal medullaris; APOE: Apolipoprotein; MAPK14: Mitogen-activated protein kinase 14; ROC: Receiver Operating Characteristic; AUC: Area Under the Curve; CI: Confidence Interval; SVM: Support Vector Machine; LASSO regression: Least Absolute Shrinkage and Selection Operator.

[0099] The definitions of the English abbreviations for "gene" used in this invention are summarized below:

[0100] ADM: Adrenomedullin; ANGPTL4: Angiopoietin-like protein 4; APEX1: Depurin / depyrimidine endonuclease 1; APOC1: Apolipoprotein C1; APOE: Apolipoprotein E; ATF6B: Activating transcription factor 6β; CDK4: Cyclin-dependent kinase 4; CDKN1A: Cyclin-dependent kinase inhibitor 1A (p21); CDKN1B: Cyclin-dependent kinase inhibitor 1B (p27); CDKN2A: Cyclin-dependent kinase inhibitor 2A (p16); CDKN2B: Cyclin-dependent kinase inhibitor 2B (p15); CEBPA: CCAAT enhancer-binding protein α; CETP: cholesterol ester transporter; CNR1: cannabinoid receptor 1; EDN1: endothelin 1; FASN: fatty acid synthase; FOXM1: forkhead box protein M1; G6PD: glucose-6-phosphate dehydrogenase; GAL: glycopropyl peptide; GDRD5: dopamine receptor D5; GPBAR1: G protein-coupled bile acid receptor 1; GTF2I: universal transcription factor II-I; GNPDA: glucosamine-6-phosphate deaminase; HMGCR: 3-hydroxy-3-methylglutaryl-CoA reductase; IDH1: isocitrate dehydrogenase 1; KIF20A: kinin family member 20A; Lipin 1: Lipoprotein 1; LTB: Lymphotoxin β; LPIN1: Lipoprotein 1; MAPK14: Mitogen-activated protein kinase 14; MPV17: Mitochondrial inner membrane protein MPV17; MYOM1: Myosin 1; NAMPT: Nicotinamide phosphoribosyltransferase; NCBP1: Nuclear cap-binding protein subunit 1; NFATC2: Activated T cell nuclear factor 2; PCK1: Phosphoenolpyruvate carboxykinase 1; PFKM: Phosphofructokinase, muscle type; P2RX7: Purinergic receptor P2X ligand-gated ion channel 7; PGAM2: phosphoglycerate mutase 2; PIK3R1: phosphatidylinositol 3-kinase regulatory subunit 1; PKM: pyruvate kinase M1 / 2; PLA2G7: lipoprotein-associated phospholipase A2; PDIA6: protein disulfide isomerase A6; PEA15: phosphoprotein enriched in astrocytes 15; RORC: RAR-associated orphan receptor C; STX1A: synaptic fusion protein 1A; SERPINH1: serine protease inhibitor H1 (Heat shock protein 47); SLC16A1: Solute carrier family 16 member 1 (monocarboxylic acid transporter 1); SLC2A4: Glucose transporter 4; SOD3: Superoxide dismutase 3; SREBF1: Sterol regulatory element-binding protein 1; THSD7A: Platelet-reactive protein type 1 domain protein 7A; TNFRSF11B: Tumor necrosis factor receptor superfamily member 11B; TPO: Thyroid peroxidase; TUBB4B: Tubulin β-4B chain; TUBB6: Tubulin β-6 chain;TYMP: Thymidine phosphorylase.

[0101] Example 1

[0102] This embodiment details the construction and validation of a childhood obesity risk scoring model based on multicenter genomics data, resulting in a gene expression scoring model for childhood obesity diagnosis and risk stratification of abnormal glucose and lipid metabolism. The flowchart is shown below. Figure 1 As shown.

[0103] 1. Data Download

[0104] Using the standard data retrieval and acquisition interface of the R package, and employing the GEO database (GEOquery Version 2.70.0), download the childhood obesity datasets GSE9624, GSE104815, and GSE139400 from the international public functional genomics data storage and retrieval repository (i.e., the GEO database).

[0105] Table 1: Microarray chip information for gene expression comprehensive database

[0106] Dataset GSE9624 GSE104815 GSE139400 platform GPL570 GPL21827 GPL570 Species Humans Humans Humans tissue samples greater omentum adipose tissue abdominal subcutaneous fat rectus femoris Childhood obesity group samples 5 4 5 control group samples 6 4 5 refer to PMID: 25856673 PMID: 29884798

[0107] The samples in datasets GSE9624, GSE104815, and GSE139400 were all from *Homo sapiens*, with tissue origins of omental adipose, abdominal subcutaneous adipose, and rectus femoris muscle, respectively. Specific information is shown in Table 1. Dataset GSE9624 was run on the GPL570 microarray platform and included 5 obese children and 6 control samples. Dataset GSE104815 was run on the GPL21827 microarray platform and included 4 obese children and 4 control samples. Dataset GSE139400 was run on the GPL570 microarray platform and included 5 obese children and 5 control samples. All obese children and control samples were included in this study.

[0108] Lipid metabolism-related genes were collected using comprehensive human gene function annotation and knowledge databases, such as GeneCards. These databases provide comprehensive information on human genes. Using "Lipid metabolism" as the search keyword, and retaining only lipid metabolism-related genes with "Protein Coding" and "Relevance Score > 2," a total of 3326 lipid metabolism-related genes were obtained.

[0109] Similarly, using "Lipid metabolism" as a keyword, authoritative, peer-reviewed biomedical science databases, such as PubMed, yielded a set of 15,604 lipid metabolism-related genes (LMRGs). After merging and removing duplicates, a total of 15,705 lipid metabolism-related genes were obtained. After integrating the dataset and removing batch effects, some gene data are shown in the table below.

[0110] Table 2: Data on some lipid metabolism-related genes

[0111]

[0112] In the table, A1BG: α-1-B glycoprotein; A1BG-AS1: α-1-B glycoprotein-antisense chain 1; A1CF: APOBEC1 complement factor; ZZEF1: Zinc finger ZZ type and EF hand domain containing protein 1; ZYX: X-linked zinc finger protein.

[0113] Genes related to glucose metabolism were collected using the comprehensive human gene function annotation and knowledge database (GeneCards). Using "Glucose Metabolism" as the search keyword, only genes related to glucose metabolism with "ProteinCoding" and "Relevance Score > 2" were retained, resulting in 991 such genes. Similarly, using "Glucose Metabolism" as the keyword, authoritative, peer-reviewed biomedical science databases, such as PubMed, yielded 7 glucose metabolism-related genes. After merging and removing duplicates, a total of 998 glucose metabolism-related genes were obtained. Taking the intersection of lipid metabolism-related genes and glucose metabolism-related genes yielded 930 genes related to both lipid metabolism and glucose metabolism.

[0114] 2. Merging of Childhood Obesity Datasets. First, batch effect correction algorithms based on empirical Bayesian statistical models, such as the R package sva (version v3.50.0), were used to remove batch effects from datasets GSE9624 and GSE104815, resulting in a merged GEO dataset containing 9 obese children and 10 control samples.

[0115] To verify the effectiveness of batch effect removal, principal component analysis was used. The principal component analysis plots of the expression matrices before and after batch effect removal were analyzed, as shown below. Figure 2 As shown in CD, the comparison demonstrates the low-dimensional feature distribution of the samples, and the results show that the batch effect in the children's obesity dataset is basically eliminated after batch effect processing.

[0116] In addition, box plots of the distribution of the integrated GEO dataset before and after batch processing are shown, such as... Figure 2 As shown in Figure AB, the difference in expression values ​​before and after batch effect removal was compared, further supporting our analytical results.

[0117] Specifically, Figure 2 In Figure A, the box plot shows the distribution of the GEO dataset before batch processing. The horizontal axis represents the different sample sources, namely two independent gene expression datasets, GSE9624 (green) and GSE104815 (orange). The vertical axis represents the quantitative value of gene expression (the raw signal intensity before standardization), ranging from 0 to 15.

[0118] Figure 2 Figure B shows the box plot of the integrated GEO dataset distribution after batch processing. The horizontal axis represents the different sample sources, namely two independent gene expression datasets GSE9624 (green) and GSE104815 (orange). The vertical axis represents the quantitative value of gene expression (after log transformation), ranging from 0 to 15.

[0119] Figure 2 The graph in C is the PCA plot of the integrated GEO dataset before batch processing. The horizontal axis represents the first principal component (Dim1), the direction of the largest variation in the data. Dim1 (83.4%) means that up to 83.4% of the total variation before batch processing can be explained by the first principal component, mainly driven by batch differences. The vertical axis represents the second principal component (Dim2), the direction of the second largest variation in the data. Dim2 is 7.8%. Samples from different datasets (distinguished by shape and color) are completely separated on the Dim1 axis, a typical characteristic of batch effects.

[0120] Figure 2 The diagram in Figure D shows the PCA plot of the integrated GEO dataset after batch processing removal. The horizontal axis represents the first principal component (Dim1), which is the direction of greatest variance in the data. Dim1 (21.5%) indicates that the variance explained by the first principal component drops sharply to 21.5% after batch processing removal. The vertical axis represents the second principal component (Dim2), which is the direction of second-greatest variance in the data, at 16.3%. After batch processing removal, the samples from the two datasets are mixed together in a low-dimensional space, no longer clustered by batch, and the variance explained by each principal component becomes more balanced, proving that batch effects are largely eliminated.

[0121] Subsequently, we used the same R package `sva` to standardize the GSE139400 dataset, obtaining the standardized dataset. To evaluate the effectiveness of the standardization process, such as... Figure 2As shown in Figure EF, the difference in expression values ​​before and after standardization was compared using a distribution box plot.

[0122] Figure 2 The box plot in E represents the distribution of the GSE139400 dataset before standardization. The horizontal axis represents all samples within a single dataset, i.e., GSE139400. The vertical axis represents the quantitative value of gene expression (signal value before standardization), ranging from 0 to 10. This shows that the distribution of expression values ​​among samples may exhibit some skewness or scale differences.

[0123] Figure 2 The box plot in Figure F shows the distribution of the standardized dataset GSE139400. The horizontal axis represents all samples within a single dataset, i.e., GSE139400. The vertical axis represents the quantitative value of gene expression (standardized signal value), ranging from 0 to 10. This shows that the median expression values ​​of all samples are aligned, and the distribution range becomes consistent, ensuring the data are on a comparable standard scale for subsequent analysis.

[0124] Figure 2 In the diagram, green represents the children's obesity dataset GSE9624, orange represents the children's obesity dataset GSE104815, and brown represents the children's obesity dataset GSE139400.

[0125] These steps and results provide a reliable data foundation for research on gene expression related to childhood obesity, ensuring the accuracy and reproducibility of the analysis.

[0126] 3. Differentially expressed genes related to lipid and glucose metabolism in childhood obesity

[0127] In investigating differentially expressed genes related to lipid and glucose metabolism in children with obesity, the GEO integrated dataset was used to divide the samples into an obese group and a control group. Differential expression analysis was performed using a statistical inference method based on a linear model (e.g., R package limma), with |logFC| > 0.5 and p-value < -0.05 set as the threshold for differentially expressed genes. In the analysis, a total of 783 differentially expressed genes meeting this threshold were identified, of which 424 were upregulated genes (logFC > 0.5) and 359 were downregulated genes (logFC < -0.5). The differential expression analysis results were visualized using a volcano plot.

[0128] Subsequently, to identify differentially expressed genes related to lipid and glucose metabolism associated with childhood obesity, an intersection analysis was performed on all differentially expressed genes and those related to lipid and glucose metabolism. The results showed a total of 60 differentially expressed genes related to lipid and glucose metabolism, which were visualized using a Venn diagram. Furthermore, a heatmap was generated using the R package pheatmap to illustrate the expression differences of these differentially expressed genes, showing the expression of the top 20 genes. (See...) Figure 3 AC (Chinese)

[0129] exist Figure 3 Figure A shows a volcano plot of differentially expressed genes between the obese children and the control group in the integrated GEO dataset. The x-axis represents log2FC (logarithmic fold change), the fold change in gene expression in the obese children relative to the control group (logarithmically transformed to base 2), with the dashed line marking the threshold of |log2FC| = 0.5. The y-axis represents -log... 10 (p-value), the p-value of differential expression analysis is decremented by -log 10 The p-value is used to measure the statistical significance of differences. A larger value (higher position of the point) indicates more significant differential expression of the gene (i.e., a smaller original p-value). A horizontal dashed line typically corresponds to a significance threshold of 0.05.

[0130] Figure 3The abbreviations of the English names involved in the Chinese A are explained below.ADM: Adrenomedullin, ANGPTL4: Angiopoietin-like protein 4, APEX1: Depurin / depyrimidine endonuclease 1, APOC1: Apolipoprotein C1, APOE: Apolipoprotein E, ATF6B: Activating transcription factor 6β, CDK4: Cyclin-dependent kinase 4, CDKN1A: Cyclin-dependent kinase inhibitor 1A (p21), CDKN1B: Cyclin-dependent kinase inhibitor 1B (p27), CDKN2A: Cyclin-dependent kinase inhibitor 2A (p16), CDKN2B: Cyclin-dependent kinase inhibitor 2B (p15), CEBPA: CC AAT enhancer-binding protein α, CETP: cholesterol ester transporter, CNR1: cannabinoid receptor 1, EDN1: endothelin 1, FASN: fatty acid synthase, FOXM1: forkhead box protein M1, G6PD: glucose-6-phosphate dehydrogenase, GAL: glycopropyl peptide, GDRD5: dopamine receptor D5, GPBAR1: G protein-coupled bile acid receptor 1, GTF2I: universal transcription factor II-I, GNPDA: glucosamine-6-phosphate deaminase, HMGCR: 3-hydroxy-3-methylglutaryl-CoA reductase, IDH1: isocitrate dehydrogenase 1, KIF20A: kinin family member 20A, Lipin 1: Lipoprotein 1, LTB: Lymphotoxin β, LPIN1: Lipoprotein 1, MAPK14: Mitogen-activated protein kinase 14, MPV17: Mitochondrial inner membrane protein MPV17, MYOM1: Myosin 1, NAMPT: Nicotinamide phosphoribosyltransferase, NCBP1: Nuclear cap-binding protein subunit 1, NFATC2: Activated T cell nuclear factor 2, PCK1: Phosphoenolpyruvate carboxykinase 1, PFKM: Phosphofructokinase, muscle type, P2RX7: Purinergic receptor P2X ligand-gated ion channel 7, PGAM2: Phosphoglycerate mutase 2, PIK3R1: Phosphatidylinositol 3-kinase regulatory subunit 1, PKM: Pyruvate kinase M1 / 2 type, PLA2G7: Lipoprotein-associated phospholipase A2, PDIA6: Protein disulfide isomerase A6, PEA15: Phosphoproteins enriched in astrocytes 15, RORC: RAR-associated orphan receptor C, STX1A: Synaptic fusion protein 1A, SERPINH1: Serine protease inhibitor H1 (heat shock protein 47), SLC16A1: Solute carrier family 16 member 1 (monocarboxylic acid transporter 1), SLC2A4: Glucose transporter 4, SOD3: Superoxide dismutase 3, SREBF1: Sterol regulatory element binding protein 1, THSD7A: Platelet-reactive protein type 1 domain protein 7A, TNFRSF11B: Tumor necrosis factor receptor superfamily member 11B, TPO: Thyroid peroxidase, TUBB4B: Tubulin β-4B chain, TUBB6: Tubulin β-6 chain, TYMP: Thymidine phosphorylase.

[0131] Figure 3 B is a Venn diagram of differentially expressed genes and lipid metabolism and glucose metabolism-related genes in the integrated GEO dataset.

[0132] Figure 3 The Venn diagram shown in Figure B indicates that the pink circles on the left represent all differentially expressed genes, while the blue circles on the right represent differentially expressed genes related to lipid and glucose metabolism. Of these, 723 are unique to all differentially expressed genes and are unrelated to lipid and glucose metabolism. The blue circles on the right represent 870 unique to differentially expressed genes related to lipid and glucose metabolism; these are differentially expressed genes involved only in lipid / glucose metabolism. The overlapping area in the middle represents genes belonging to both groups, totaling 60 genes—those that are both differentially expressed overall and involved in the regulation of lipid and glucose metabolism.

[0133] Figure 3 The image in section C is a heatmap of differentially expressed genes related to lipid and glucose metabolism from the integrated GEO dataset. The pink bars at the top represent the control group, and the orange bars represent the obese children group. Red indicates high expression, and blue indicates low expression. The horizontal axis represents individual samples. Each bar represents an independent sample. Samples are typically clustered and labeled (divided into obese and control groups). The vertical axis represents differentially expressed genes related to lipid and glucose metabolism. Each row represents a specific lipid and glucose metabolism-related gene. The heatmap shows the top 20 genes with the most significant or core differential expression, and the legend color intensity represents the Z-score normalized value of gene expression levels.

[0134] Red: indicates that the expression level of the gene in this sample is higher than the average expression level of the gene in all samples.

[0135] Blue: indicates that the expression level of the gene in this sample is lower than the average expression level of the gene in all samples.

[0136] The shade of color indicates the degree of deviation from the average.

[0137] Figure 3The definitions of all the English abbreviations involved in the Chinese character C are as follows. SPP1: secretory phosphoprotein 1, HSD11B1: 11-β-hydroxysteroid dehydrogenase 1, ADM: adrenomedullin, P2RX7: P2X purine receptor 7, PLA2G7: phospholipase A2 group VII, MYOM1: myosin 1, SOD3: superoxide dismutase 3, LEP: leptin, GPBAR1: G protein-coupled bile acid receptor 1, CDKN2B: cyclin-dependent kinase inhibitor 2B, PCK1: phosphoenolpyruvate carboxykinase 1, KIF20A: kinin family member 20A, APOE: apolipoprotein E, APOC1: apolipoprotein C1, FASN: fatty acid synthase, RORC: RAR-associated orphan receptor C, STX1A: synaptic fusion protein 1A, NAMPT: nicotinamide phosphoribosyltransferase, PCBD1: pterin-4α-carbamoyl dehydratase 1, PFKM: phosphofructokinase.

[0138] 4. Construction of a diagnostic model for childhood obesity

[0139] To obtain a diagnostic model for childhood obesity based on the integrated GEO dataset, logistic regression analysis was performed on differentially expressed genes related to lipid and glucose metabolism, with the dependent variable being a binary variable (childhood obesity group and control group). First, using a p-value <0.05 as the criterion, genes significantly associated with CO were screened from the differentially expressed genes related to lipid and glucose metabolism, and an initial logistic regression model was constructed. The grouping expression of these genes was then visualized using a forest plot. Next, logistic regression analysis was performed based on 60 LGMRDEGs, and a model was constructed as follows: Figure 4 As shown in Figure A, the forest plot of 29 differentially expressed genes related to lipid and glucose metabolism in the logistic regression model for childhood obesity diagnosis is presented. The results show that 29 LGMRDEGs were statistically significant (p < 0.05). Next, using these 29 genes, an SVM model was constructed using the support vector machine algorithm, as shown below. Figure 4 The B-side visualization identifies the part with the lowest error rate, such as... Figure 4 The C-type visualization displays the most accurate data, and key genes are selected based on this principle.

[0140] The SVM model achieved the highest accuracy when the number of genes was 3. Subsequently, based on the 3 LGMRDEGs selected by the SVM model, a logistic regression model was constructed using the minimum absolute shrinkage and selection operator regression method. Specifically, this embodiment uses the R package glmnet (Version 4.1.8) with set.seed(1000) and family="binomial" as parameters for LASSO regression analysis. LASSO regression reduces overfitting and improves the model's generalization ability by adding a penalty term (lambda × absolute value of the slope), thus constructing a diagnostic model for childhood obesity. The LASSO regression model is plotted as follows: Figure 4 As shown in D; draw the variable trajectory graph, as shown. Figure 4 As shown in Figure E, the results are visualized, revealing that the model contains three differentially expressed genes (i.e., model genes) related to lipid metabolism and glucose metabolism: APOE, ADM, and MAPK14.

[0141] Figure 4 Figure A shows a forest plot of 29 differentially expressed genes related to lipid and glucose metabolism included in a logistic regression model for childhood obesity diagnosis. The horizontal axis represents the regression coefficient (Beta), indicating the magnitude and direction of the impact of changes in gene expression levels on the risk of childhood obesity when each gene is used as a predictor. The vertical line (X = 0) is the null line. If the 95% confidence interval (95% CI) of a gene lies entirely to the right (Beta > 1), its upregulation is positively correlated with obesity risk. If it lies entirely to the left (Beta < 1), it is negatively correlated. If the horizontal line intersects the null line, there is no statistical significance. The vertical axis represents the differentially expressed genes related to lipid and glucose metabolism.

[0142] Figure 4The abbreviations related to "A" in the text are explained below: ADM: Adrenal medullaris, ANGPTL4: Angiopoietin-like protein 4, APEX1: Depurin / depyrimidine endonuclease 1, APOE: Apolipoprotein E, CETP: Cholesterol ester transport protein, EDN1: Endothelin 1, ENO1: Enolase 1, FASN: Fatty acid synthase, G6PD: Glucose-6-phosphate dehydrogenase, GAL: Glycopropyl peptide, GTF2I: Universal transcription factor II-I, HMGCR: 3-hydroxy-3-methylglutaryl-CoA reductase, HSD11B1: 11-β-hydroxysteroid dehydrogenase 1, LPIN1: Lipokinase 1, MAPK14: Mitogen-activated protein kinase 14, MPV17: MPV1 7. Mitochondrial inner membrane proteins: NCBP1: nuclear cap-binding protein subunit 1, NFATC2: activated T cell nuclear factor 2, PCBD1: pterin-4α-carbamoyl dehydratase 1, PDIA6: protein disulfide isomerase A6, PDK1: pyruvate dehydrogenase kinase 1, PIK3R1: phosphatidylinositol 3-kinase regulatory subunit 1, SLC2A4: glucose transporter 4, SOD3: superoxide dismutase 3, SPP1: osteopontin, THSD7A: platelet-reactive protein type 1 domain protein 7A, TNFRSF11B: tumor necrosis factor receptor superfamily member 11B, TPO: thyroid peroxidase, TUBB4B: tubulin β-4B chain.

[0143] Figure 4 In this context, B represents the number of genes with the lowest error rate obtained by the support vector algorithm. Figure 4 The middle column (C) represents a visualization of the number of genes with the highest accuracy. Figure 4 Plot B shows the cross-validation error rate curve of the SVM model. This subplot is used to determine how many genes are needed to minimize the prediction error of the SVM model. The horizontal axis represents the number of genes included in order of importance during modeling. The vertical axis represents the cross-validation error rate, the proportion of prediction errors made by the model when evaluated using cross-validation. The curve shows that the error rate drops to its lowest point (approximately 0.0556) when the number of genes is 3, indicating that 3 genes are sufficient to build a robust model with low error. Figure 4 Figure C represents the cross-validation accuracy curve of the SVM model. C is complementary to B, determining the optimal number of genes from an accuracy perspective. The horizontal axis represents the number of genes (consistent with B). The vertical axis represents the cross-validation accuracy. This curve shows the proportion of correct predictions made by the model when evaluated using cross-validation. The curve shows that the accuracy peaks (approximately 0.944) when the number of genes is 3, further confirming that 3 genes are the optimal choice.

[0144] Figure 4Figure D shows the diagnostic model of the minimum absolute contraction and selection operator regression model, illustrating the analysis of cross-validation error rate and regularization parameter selection. It illustrates the relationship between model prediction error determined through cross-validation and the strength of regularization, used to select the optimal penalty parameter λ. The horizontal axis represents the logarithm of the regularization parameter λ (log(λ)). From left to right, a decreasing λ value indicates a weaker constraint on the model coefficients and an increased model complexity. The vertical axis represents the model's misclassification error rate; a lower value indicates better predictive performance. The red curve and dots represent the average cross-validation error rate of the model at different λ values. Its typical "U"-shaped trajectory reveals the bias-variance tradeoff: Left side (large λ value): excessive penalty, the model is too simple, underfitting, leading to a high error rate. Right side (small λ value): insufficient penalty, the model may be too complex and overfitting the training data, leading to increased generalization error. The trough at the bottom corresponds to the λ range with the best generalization performance (approximately log(λ) between -6 and -4), where the optimal balance between model complexity and predictive ability is found. Error bars (gray): Represent the range of variation (e.g., standard deviation) of multiple cross-validation results. Shorter error bars near the trough indicate stable and reliable model performance in this region. Numbers above the X-axis: Indicate the number of non-zero coefficients in the model at the corresponding λ value (i.e., the number of features selected and retained). In the optimal interval, the model achieves the best predictive performance with 3 features.

[0145] Figure 4 E in the diagram represents the variable trajectory plot, which is a LASSO coefficient path and feature selection analysis. This plot visually shows how the regression coefficients of all predictor variables (genes) change with model complexity, fully demonstrating the dynamic process of feature selection. The vertical axis represents the standardized regression coefficients of each variable. The sign of the coefficient indicates the direction of its effect on the outcome variable, and the absolute value represents the strength of the effect. The horizontal axis represents the proportion of bias explained by the model. From left to right, the regularization strength λ gradually decreases, and the amount of information incorporated by the model increases. The colored trajectory lines represent the coefficient change path of a specific gene. When the constraints are extremely strong (leftmost), all coefficients are compressed to 0. As the constraints are relaxed (moving to the right), the coefficients of important variables first separate from 0 and their absolute values ​​increase. The coefficients of redundant or unimportant variables fluctuate around 0. Final model interpretation: Based on... Figure 4 The optimal λ determined by D (corresponding to an explanatory bias of approximately 0.8) has most coefficients of 0. Figure 4 In the final screening, only three key genes were identified to form the model: APOE (apolipoprotein E, green line): coefficient approximately -0.99, showing a significant negative correlation; ADM (adrenergic medullaris, red line): coefficient approximately 0.84, showing a significant positive correlation; and MAPK14 (mitogen-activated protein kinase 14, black line): coefficient approximately -0.98, showing a significant negative correlation.

[0146] Finally, the risk score is calculated based on the risk coefficients obtained from LASSO regression analysis. The calculation formula is as follows:

[0147] RiskScore = APOE * (-0.994) + ADM * (0.843) + MAPK14 * (-0.98).

[0148] To comprehensively evaluate the efficacy of the childhood obesity diagnostic model, a systematic validation analysis was conducted on the GEO integrated dataset and the independent dataset GSE139400. First, on the GEO integrated dataset, the receiver operating characteristic (ROC) curves constructed based on risk scores showed that it had high diagnostic accuracy (AUC > 0.9). Figure 5 As shown in Figure A.

[0149] Figure 5 Figure A shows the receiver operating characteristic (ROC) curve of the risk score in the integrated GEO dataset. The x-axis represents 1 - Specificity (False Positive Rate), indicating the proportion of samples that were incorrectly classified as "childhood obesity" by the model in the actual "control" sample. The y-axis represents sensitivity (True Positive Rate), indicating the proportion of samples that were correctly classified as "childhood obesity" by the model. The area under the curve (AUC) > 0.9, indicating that the risk score calculated based on the three model genes has extremely high accuracy in distinguishing between the childhood obesity group and the control group.

[0150] Figure 5 In the middle, B is a nomogram of the integrated model genes in the GEO dataset used in the childhood obesity diagnostic model. (Example:) Figure 5 As shown in Figure B, a nomogram illustrating the relationships between the three model genes was plotted. The results indicate that APOE expression contributes the most to the model, while MAPK14 contributes relatively less. The horizontal axis represents the score: a scale assigned to each predictor variable (here, the expression values ​​of the three genes). The vertical axis (left-hand variable list) represents the predictor variables, i.e., the three model genes: ADM, APOE, and MAPK14. Each gene corresponds to a "score axis." The expression values ​​of the three genes are plotted on their respective score axes and projected upwards onto the "score" axis to obtain individual scores. The sum of all individual scores is then plotted on the "total score" axis at the bottom and projected vertically downwards onto the "childhood obesity risk" axis at the bottom to obtain the probability that the sample is predicted to be obese. APOE has the longest score axis, indicating that changes in its expression level have the greatest impact on the total score (i.e., the final predicted risk), contributing the most; MAPK14 has the shortest score axis, contributing the least.

[0151] To further evaluate the model's calibration and clinical applicability, calibration curves were plotted and analyzed. The results show that the predicted probabilities closely match the actual probabilities, indicating that the model has good calibration. The results are as follows: Figure 5 As shown in C. Figure 5 The graph in Figure C represents the calibration curve of the childhood obesity diagnostic model based on the risk score of the integrated GEO dataset. The horizontal axis represents the model's predicted probability of childhood obesity. The vertical axis represents the actual observed probability of childhood obesity. The thin gray solid line (ideal line) represents the ideal situation where the predicted probability is exactly equal to the actual probability, serving as a reference baseline. The closer other curves are to this line, the better the model is calibrated. The black solid line is the logical calibration line, indicating the calibration status of the diagnostic model. This curve is slightly lower than the ideal line in most areas, and the upper indicator shows Slope = 0.756 (i.e., slope = 0.756). This indicates a slight systematic underestimation in the original model's predictions. The black dashed line is the nonparametric calibration line, supplementing the logical calibration. This curve almost completely coincides with the ideal line. This indicates that, from a flexible nonparametric perspective, the model's predicted probability is extremely accurate with minimal bias. C: Area under the ROC curve, equal to (Dxy + 1) / 2. Its value ranges from [0.5, 1]. C = 1.000 also indicates that the model has perfect discriminative power.

[0152] Dxy: Rank correlation index, ranging from -1 to 1. A value closer to 1 indicates better discrimination. Dxy = 1.000 means the model can perfectly distinguish all samples. Slope: Slope, Intercept: Intercept, R0 2 : Pseudo-R-squared, Emax: The maximum absolute deviation between the calibration curve and the ideal line. E90 and Eavg: The deviation and mean absolute deviation at the 90th quantile, respectively. E90=0.000 and Eavg=0.000 confirm the extreme accuracy of the model's predictions. S:z and S:p (Hosmer-Lemeshow test): Tests the goodness of fit between the model's predicted probabilities and the actual probabilities. S:z is the test statistic, and S:p is the p-value. S:p=1.000 indicates that the null hypothesis is fully accepted, meaning that the model's predicted probabilities are not significantly different from the actual probabilities, and the fit is excellent.

[0153] Decision curve analysis also confirmed that the model has good clinical net benefit within a certain threshold range, such as... Figure 5 As shown in D. Figure 5Figure D shows the decision curve analysis of the childhood obesity diagnostic model based on risk scores from the integrated GEO dataset. The horizontal axis represents the high-risk threshold, i.e., the lowest acceptable probability of disease in clinical decision-making (such as diagnosing and intervening in obesity). The vertical axis represents the standardized net benefit. This curve is used to evaluate the clinical applicability of the model at different decision thresholds. Specifically, the gray solid line (all positive) represents the net benefit assuming all samples are childhood obesity. The black solid line (all negative) represents the net benefit assuming all samples are not childhood obesity. The red solid line (risk assessment model) represents the net benefit of using the diagnostic model for decision-making. When the model's curve is significantly higher than the "all positive" and "all negative" lines within a large threshold range, it indicates that using the model to guide clinical decision-making yields a higher net benefit and has clinical application value.

[0154] Furthermore, functional similarity analysis was used to assess the potential importance of these genes in biological processes, such as... Figure 5 As shown in E. Figure 5 The box plot in Figure E shows the results of the Friends analysis of the model genes' functional similarity. The horizontal axis represents the functional similarity score, which quantifies the strength of the functional association between the target gene and a set of genes known to be involved in a specific biological process. The vertical axis represents the gene names, i.e., the three model genes (ADM, APOE, and MAPK14). A higher score indicates a stronger functional association between the gene and the target biological process (e.g., childhood obesity), suggesting a potentially more important role in the disease mechanism.

[0155] Independent analysis of individual model genes showed that ADM expression level had the highest diagnostic value (AUC > 0.9), while APOE and MAPK14 also showed some diagnostic accuracy (0.7 < AUC ≤ 0.9). Figure 5 As shown in FH. Figure 5 In the middle, F represents the model gene ADM. Figure 5 G stands for APOE. Figure 5 In the figure, H represents MAPK14, and the ROC curve is shown in the integrated GEO dataset. The horizontal axis represents 1 - specificity (false positive rate), and the vertical axis represents sensitivity (true positive rate). These three subplots evaluate the efficacy of individual model genes as diagnostic biomarkers.

[0156] When AUC > 0.5, it indicates that the expression of the molecule promotes the occurrence of the event. The closer the AUC is to 1, the better the diagnostic effect. AUC between 0.7 and 0.9 has a certain accuracy, and AUC above 0.9 has high accuracy. Figure 5 The results from G showed that the AUC of APOE was >0.9, indicating that the APOE gene alone has extremely high discrimination accuracy. Figure 5 China F and Figure 5In H, the AUC of ADM and MAPK14 was 0.7 < AUC ≤ 0.9, indicating that these two genes also had certain discriminative accuracy alone, but lower than that of APOE.

[0157] Finally, in the independent validation set GSE139400, we also evaluated the diagnostic efficacy of the risk score and the expression levels of each model gene for childhood obesity, and plotted the corresponding nomogram and calibration curve to further verify the stability and generalization ability of the model. For the diagnosis and validation analysis of childhood obesity, we obtained Figure 6 the relevant charts shown below.

[0158] Figure 6 In Figure A is the receiver operating characteristic curve (ROC curve) of the model risk score (GSE139400 validation set). Abscissa: 1 - specificity, i.e., false positive; Ordinate: sensitivity, i.e., true positive rate. Blue curve: represents the ROC curve based on the risk score of three genes. Grey line: represents the random guess line (diagonal line, AUC = 0.5). The curve is above the diagonal line but closer to it. The AUC = 0.640 is between 0.5 and 0.7, indicating that this risk score has certain discriminative ability in the GSE139400 dataset, but the accuracy is moderately low.

[0159] Figure 6 In Figure B is the nomogram in GSE139400 of the model genes in the childhood obesity diagnosis model. Abscissa (at the bottom of each variable axis): the expression levels (normalized) of the three model genes APOE, ADM, and MAPK14. Ordinate (at the far left) and the upper axis: "Score" axis: the single score corresponding to each gene expression value. "Total score" axis: the total score obtained by adding the single scores of the three genes. "Childhood obesity risk" axis: the probability of childhood obesity predicted according to the total score.

[0160] Figure 6 In Figure C is the calibration curve of the risk score of the childhood obesity diagnosis model based on GSE139400 (Calibration Curve). Abscissa: the probability of childhood obesity predicted by the model. Ordinate: the actually observed probability of childhood obesity. Grey solid line: ideal perfect calibration line. Black solid line: the calibration curve of the model on GSE139400, black dashed line: non-parametric calibration line. The calibration curve is close to but slightly deviated from the ideal diagonal line, indicating that the probability values predicted by the model are basically consistent with the actual incidence rate, but there are certain systematic biases (possibly slightly overestimated or underestimated). Even if the AUC discriminative ability is not high, good calibration means that the risk values predicted by the model still have reference value. The abbreviations and their explanations are as follows: C: area under the ROC curve. Dxy: rank correlation index. Slope: slope, Intercept: intercept, R 2Pseudo-R-squared, Emax: The maximum absolute deviation between the calibration curve and the ideal line. E90 and Eavg: The deviation and mean absolute deviation at the 90th quantile, respectively. S:z and S:p (Hosmer-Lemeshow test): Goodness-of-fit tests to determine whether the predicted probabilities of the model match the actual probabilities. S:z is the test statistic, and S:p is the p-value.

[0161] Figure 6 D in the model gene is ADM. Figure 6 In the middle, E stands for APOE. Figure 6 F in the figure represents the ROC curve of MAPK14 in the integrated GEO dataset. The x-axis represents 1 - specificity, and the y-axis represents sensitivity. When AUC > 0.5, it indicates that the expression of the molecule promotes the occurrence of the event; the closer the AUC is to 1, the better the diagnostic effect. AUC between 0.5 and 0.7 indicates lower accuracy, while AUC between 0.7 and 0.9 indicates some accuracy.

[0162] 6. To further explore the expression characteristics and interactions of model genes in childhood obesity, we conducted differential expression validation and correlation analysis on the GEO integrated dataset. First, by plotting group comparison diagrams, the results are as follows: Figure 7 As shown, the expression differences of three model genes between the obese children group and the control group were evaluated.

[0163] Figure 7 Plot A shows the comparison of model genes in the integrated GEO dataset between the obese children's group and the control group. The horizontal axis (X-axis) represents the names of the three model genes, from left to right: ADM, APOE, and MAPK14. The vertical axis (Y-axis) represents the standardized gene expression level, indicating the relative expression level of the gene. A higher value indicates a higher expression level of the gene in the sample. The pink box plot represents the expression distribution of the gene in the control group. The orange box plot represents the expression distribution of the gene in the obese children's group. * indicates p-value < 0.05, statistically significant; ** indicates p-value < 0.01, highly statistically significant; *** indicates p-value < 0.001, extremely statistically significant. The results showed that the expression levels of APOE differed significantly between the two groups (p < 0.001), the expression levels of ADM differed significantly (p < 0.01), and the expression levels of MAPK14 differed significantly (p < 0.05), which confirmed the association between these genes and childhood obesity.

[0164] Subsequently, we analyzed the expression correlations among these genes using the Spearman algorithm and visualized the results using heatmaps. Figure 7As shown in B.

[0165] Figure 7 Figure B shows a heatmap of the correlation between model genes in the integrated GEO dataset between the obese and control groups of children. The x and y axes represent the names of the model genes; red indicates a positive correlation, blue indicates a negative correlation, and white indicates no correlation. Each off-diagonal cell displays the specific Pearson correlation coefficient and significance marker. Group comparisons are shown, with pink representing the control sample and orange representing the obese children. Red indicates a positive correlation, and blue indicates a negative correlation. The intensity of the color represents the strength of the correlation. A p-value < 0.05 indicates statistical significance. Correlation analysis shows that in the integrated dataset, there is the strongest significant negative correlation between ADM and APOE (r = -0.51, p < 0.05), revealing a potential biological relationship of antagonism or regulation between these two key genes in childhood obesity.

[0166] 7. The optimal diagnostic threshold (Cut-off Value) is -7.29. If the RiskScore ≥ -7.29, the sample is classified as a child at high risk of abnormal glucose and lipid metabolism / obese. If the RiskScore < -7.29, the sample is classified as a person at low risk of abnormal glucose and lipid metabolism / healthy. This threshold also showed good classification performance on the independent validation set.

[0167] Based on the above research results, we selected key genes and weight coefficients, and constructed a risk scoring model based on the key genes and weight coefficients.

[0168] Risk scoring models use formulas to calculate risk scores for APOE, ADM, and MAPK14.

[0169] RiskScore = a × APOE + b × ADM + c × MAPK14

[0170] in,

[0171] RiskScore is a risk rating;

[0172] APOE represents the expression level of the apolipoprotein E gene;

[0173] ADM represents the expression level of the adrenomedullin gene;

[0174] MAPK14 represents the expression level of the mitogen-activated protein kinase 14 gene;

[0175] a , b ,c The first threshold is -7.29 when the weight coefficients a=-0.994, b=0.843, and c=-0.980 are the weight coefficients corresponding to each gene.

[0176] Example 2

[0177] This embodiment provides a device system for constructing a childhood obesity risk assessment model as described in Embodiment 1, comprising: a data acquisition module for acquiring a childhood obesity dataset, a lipid metabolism-related gene set, and a glucose metabolism-related gene set; a data preprocessing module for preprocessing the data acquired by the data acquisition module; a differentially expressed gene identification module for identifying the preprocessed data according to a threshold and outputting differentially expressed genes that meet the threshold; and a model construction module for using machine learning algorithms to screen key genes from the differentially expressed genes and determine weight coefficients to construct a risk assessment model.

[0178] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0179] Furthermore, the use of terms such as "first," "second," and "third" in terminology is merely for distinguishing descriptions of identical or similar components and should not be interpreted as emphasizing or implying the relative importance of a particular component.

[0180] Furthermore, in the description of the embodiments of the present invention, "several", "more than", and "a number of" represent at least two. The number can be any number, such as two, three, four, five, six, seven, eight, or nine, and can even exceed nine.

Claims

1. A method for constructing a childhood obesity risk assessment model, characterized in that, Includes the following steps: S1. Obtain a childhood obesity dataset, including obese children and control samples; Collect the sets of genes related to lipid metabolism and the sets of genes related to glucose metabolism, and take their intersection to obtain the sets of genes related to lipid metabolism and glucose metabolism. S2. For the childhood obesity dataset, a statistical test method based on a linear model was used to test the significance of the difference in expression levels of each gene between the childhood obesity group and the control group, and to identify differentially expressed genes that meet the threshold. The specific method for significance testing is as follows: set the fold change in differential expression |logFC| > 0.5 and the corrected p-value of the differential test < 0.05 as the threshold for differentially expressed genes, and identify differentially expressed genes that meet the threshold; S3. Compare and analyze the differentially expressed genes obtained in S2 with the lipid metabolism and glucose metabolism-related gene set obtained in S1, and take the intersection to obtain the differentially expressed gene set related to lipid metabolism and glucose metabolism. S4. Based on the differentially expressed gene set related to lipid metabolism and glucose metabolism, key genes are screened using machine learning algorithms. 3-6 key genes are selected, and their weight coefficients are determined. The key genes include APOE, ADM, and MAPK14 genes. Based on the key genes and weight coefficients obtained through screening, a risk scoring model is constructed based on the key genes and weight coefficients. The risk scoring model's risk score calculation formula is based on the selected gene GenX. i Scoring calculation: in, RiskScore is a risk rating; n is the selected GenX i Total number of items, 3 ≤ n ≤ 6; GenX i This represents the expression level of the selected i-th gene; w i The weight coefficient representing the i-th gene; GenX i Including APOE, ADM, and MAPK14; APOE represents the expression level of the apolipoprotein E gene, ADM represents the expression level of the adrenal medullarin gene, and MAPK14 represents the expression level of the mitogen-activated protein kinase 14 gene. use a , b , c The weight coefficients corresponding to genes APOE, ADM, and MAPK14 satisfy: -1.6 ≤ a ≤ -0.9, 0.6 ≤ b ≤ 1.3, -1.5 ≤ c ≤ -0.

8.

2. The method for constructing a childhood obesity risk assessment model according to claim 1, characterized in that, It also includes S5, which uses an independent validation dataset to validate the effectiveness of the constructed risk assessment model.

3. The method for constructing a childhood obesity risk assessment model according to claim 1, characterized in that, In S1, the acquisition of the childhood obesity dataset includes at least one of the datasets GSE9624, GSE104815, and GSE139400; The sets of lipid metabolism-related genes and glucose metabolism-related genes were retrieved and extracted from professional gene function and disease association databases and / or peer-reviewed biomedical literature databases. After obtaining the childhood obesity dataset, data preprocessing was performed, and batch effect correction algorithm based on empirical Bayesian framework was used to remove batch effects, resulting in an integrated dataset.

4. The method for constructing a childhood obesity risk assessment model according to claim 3, characterized in that, Standardize the integrated dataset; The differences in expression values ​​before and after standardization were compared using box plots.

5. The method for constructing a childhood obesity risk assessment model according to claim 3, characterized in that, In S2, differential expression analysis was performed on the integrated dataset. |logFC|>0.5 and p<0.05 were set as the threshold for differentially expressed genes, and differentially expressed genes that met the threshold were identified.

6. The method for constructing a childhood obesity risk assessment model according to claim 1, characterized in that, In S4, a preliminary screening is first performed, and logistic regression analysis is conducted on the differentially expressed gene sets related to lipid metabolism and glucose metabolism to screen out genes that are significantly associated with childhood obesity. Then, based on the genes obtained from the initial screening, key genes are further screened using the support vector machine algorithm.

7. The method for constructing a childhood obesity risk assessment model according to claim 6, characterized in that, In S4, an SVM model is constructed using the support vector machine algorithm, and key genes are selected based on the principle of lowest error rate and highest accuracy.

8. The method for constructing a childhood obesity risk assessment model according to claim 1, characterized in that, In S4, based on the key genes selected through screening, minimum absolute shrinkage and selection operator regression analysis are performed to determine the weight coefficients of each key gene and form a risk scoring model. The risk scoring calculation formula of the risk scoring model includes the scoring calculation of APOE, ADM, and MAPK14: RiskScore = a × APOE + b × ADM + c × MAPK14 in, RiskScore is a risk rating; APOE represents the expression level of the apolipoprotein E gene; ADM represents the expression level of the adrenomedullin gene; MAPK14 represents the expression level of the mitogen-activated protein kinase 14 gene; a , b , c Let be the weight coefficients corresponding to each gene, and satisfy: -1.6 ≤ a ≤ -0.9, 0.6 ≤ b ≤ 1.3, -1.5 ≤ c ≤ -0.

8.

9. A device for assessing the risk of childhood obesity, characterized in that, The method for constructing a childhood obesity risk assessment model according to any one of claims 1-8 includes: The data acquisition module is used to acquire datasets on childhood obesity, sets of genes related to lipid metabolism, and sets of genes related to glucose metabolism. The data preprocessing module is used to preprocess the data obtained by the data acquisition module; The differentially expressed gene identification module is used to identify preprocessed data based on a threshold and output differentially expressed genes that meet the threshold. The model building module is used to use machine learning algorithms to screen key genes from differentially expressed genes and determine weight coefficients to build a risk assessment model.