A child spine shape evaluation and detection method based on multi-source data fusion

By using a multi-source data fusion method, the problem of single-indicator dependence and difficulty in distinguishing the effects of mixed pollutants in the assessment of children's spinal morphology was solved, and more accurate assessment and risk identification of children's spinal morphology were achieved.

CN120954732BActive Publication Date: 2025-12-12CHILDRENS HOSPITAL OF FUDAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511467683.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2025-12-12
Estimated Expiration
2045-10-15

AI Technical Summary

Technical Problem

Existing technologies lack multivariate joint models when assessing spinal morphology in children, making it difficult to reflect heterogeneity within a population, and the effects of mixed pollutants are difficult to distinguish, leading to difficulties in the correlation analysis between phenotype and molecular/metabolic indicators.

Method used

A multi-source data fusion approach was adopted, including acquiring and preprocessing observation indicators, constructing a potential profile analysis model, performing analysis of variance and generalized linear regression, assessing the PFAS exposure effect, analyzing the mixed exposure effect through weighted quantiles and regression models, correcting for the influence of life trajectories, and removing noise to obtain an accurate spine dataset.

Benefits of technology

It improves the repeatability and readability of phenotypic definitions, enables targeted risk identification and intervention selection, reduces systematic errors, and enhances the sensitivity and specificity of results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954732B_ABST
    Figure CN120954732B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of spinal evaluation treatment, and particularly relates to a children's spinal form evaluation and detection method based on multi-source data fusion. The method comprises the following steps: S1: obtaining observation indexes of a child to be measured and performing pretreatment; S2: constructing a latent profile analysis model based on standardized values, obtaining a fitting index of each model, and performing weighted discrimination on the fitting index to determine an optimal latent category number; S3: performing variance analysis test difference between standardized values of each observation index of each latent category, and inducing a phenotype characteristic of each latent category and performing naming according to a difference direction and effect size of each index between different latent categories. The present application can be more targeted in risk identification and intervention target selection by using an adjusted generalized linear model on single pollutant analysis and introducing weighted quantile regression and its multiple imputation expansion on mixed exposure analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spinal assessment and treatment technology, and in particular to a method for assessing and detecting spinal morphology in children based on multi-source data fusion. Background Technology

[0002] As research into the impact of environmental pollutants on child development deepens, long-chain and short-chain polyfluoroalkyl and polyalkyl substances (collectively referred to as PFAS), as a class of chemicals with persistent and bioaccumulative characteristics, are increasingly attracting attention in studies on their potential effects on child growth and development, metabolic pathways, and musculoskeletal system during perinatal and prenatal exposure. Several existing epidemiological studies have reported correlations between PFAS and children's metabolic markers, gut-related metabolites, and individual physiological development indicators; however, these studies generally face the following limitations:

[0003] First, there is the limitation of phenotypic characterization and reliance on subjective thresholds. Traditional spinal morphology assessments often rely on single indicators or preset thresholds to determine abnormalities, lacking a classification strategy based on multivariate joint models. This makes it difficult to reflect the heterogeneity within a population and also hinders the one-to-one correspondence between morphological phenotypes and molecular / metabolic indicators for interpretation and subsequent correlation analysis.

[0004] Second, it is difficult to distinguish between mixed exposures and the effects of single pollutants. Environmental exposures are usually multi-component mixtures, and regression analysis using a single pollutant may miss the synergistic or antagonistic effects between components. Furthermore, existing mixture analysis methods are not yet closely coupled with phenotyping, behavioral trajectories, or spatial environmental factors in terms of handling measurement gaps, multiple corrections, and weight estimation.

[0005] In view of the above problems, there is an urgent need for an evaluation method based on multi-source data. Summary of the Invention

[0006] To overcome the problems mentioned in the background art, the present invention provides a method for assessing and detecting spinal morphology in children based on multi-source data fusion.

[0007] Technical solution: A method for assessing and detecting spinal morphology in children based on multi-source data fusion, comprising the following steps:

[0008] S1: Obtain the observation indicators of the children to be tested and perform preprocessing;

[0009] S2: Construct a potential profile analysis model based on standardized values, obtain the fitting index of each model, and perform weighted discrimination on the fitting index to determine the optimal number of potential categories;

[0010] S3: performing variance analysis on the standardized values of each observation index between each potential category to test the difference, and summarizing the phenotypic characteristics of each potential category and naming according to the difference direction and effect size of each index between different potential categories;

[0011] S4: using generalized linear regression model to evaluate the exposure effect of PFAS single pollutant, using weighted quantile regression to evaluate the exposure effect of PFAS complex, using GLM model to evaluate the influence of different types of cord blood PFAS exposure on children's fecal bile acid metabolism and the evaluation results of children's fecal bile acid metabolism on spine shape;

[0012] S5: correcting the evaluation results of children's spine to obtain the corrected children's spine data set.

[0013] Preferably, the acquisition of the observation indexes of the children to be tested and the preprocessing thereof include: collecting multiple spine and pelvis shape observation indexes of the children to be tested, preprocessing and standardizing each observation index to obtain the standardized values of each observation index.

[0014] Preferably, the construction of the latent profile analysis model based on the standardized values, the acquisition of the fitting index of each model, and the weighted discrimination of the fitting index to determine the optimal number of potential categories include: weighting the fitting index based on hierarchical clustering analysis method, and determining the optimal number of potential categories according to the weighted comprehensive evaluation value, wherein the fitting index is AIC, BIC, AWE, CLC and KIC.

[0015] Preferably, the variance analysis on the standardized values of each observation index between each potential category to test the difference, and the summarizing of the phenotypic characteristics of each potential category and the naming according to the difference direction and effect size of each index between different potential categories further include: using the variance analysis method to detect the difference between the children's spine shape and each of the optimal potential categories, and after summarizing the phenotypic characteristics of each potential category and naming according to the difference direction and effect size of each index between different potential categories, acquiring the prior probability and the potential category probability of the person to be tested; the prior probability reflects the relative proportion of each potential category in the total population, and the sum of the prior probabilities of all potential categories is 1; the potential category probability reflects the possibility of an individual belonging to each potential category under the condition of observation data, and the sum of all potential category probabilities of each individual is 1.

[0016] Preferably, the PFAS single pollutant exposure effect evaluation using the generalized linear regression model comprises: obtaining the covariates of the subjects, wherein the covariates include maternal education level, parity, gestational age at delivery, child gender, child birth weight, weaning age, age at spine follow-up, and length of weekly vigorous physical activity at spine follow-up; performing multiple imputation on the covariates based on a chain equation multiple imputation algorithm to generate a preset number of imputed datasets, combining the statistical results of all the imputed datasets using Rubin's rule, and calculating overall parameter estimates, standard errors, and preset confidence intervals; and analyzing the correlation between cord blood PFAS exposure and child spine morphology using the generalized linear regression model according to the covariates.

[0017] Preferably, the PFAS composite exposure effect evaluation using the weighted quantile regression comprises: preprocessing and multiple imputation on the preset number of types of PFAS concentrations in the test sample, quantifying each PFAS level by quantile, and constructing a weighted quantile regression model, wherein the weighted quantile regression model realizes estimation of mixed effects of two types of positive and negative effects, and the weight of each PFAS is obtained through repeated sampling or cross-validation, the mixed exposure index is calculated using the weight, and the mixed exposure index is taken as an independent variable to be included in a linear regression or generalized linear regression model for evaluating the overall effect of mixed exposure on the spine morphology outcome, and the overall effect estimation of mixed exposure, the relative weight ranking of each component, and the outcome indicators that are significantly under the positive model or the negative model are output.

[0018] Preferably, the use of the GLM model to evaluate the influence of different types of cord blood PFAS exposure on child fecal bile acid metabolism and the evaluation results of child fecal bile acid metabolism on spine morphology comprise: wherein in the model of the influence of cord blood PFAS exposure on child fecal bile acid metabolism, the corrected covariates include maternal education level, parity, gestational age at delivery, child gender, child birth weight, and weaning age; in the model of the influence of child fecal bile acid metabolism on spine morphology, the corrected covariates are the covariates; the mediating effect of fecal bile acid metabolism between cord blood PFAS exposure and child spine morphology is tested and evaluated based on a structural equation model; and the mediating effect is classified according to the test results.

[0019] Preferably, the use of the GLM model to evaluate the influence of different types of cord blood PFAS exposure on child fecal bile acid metabolism and the evaluation results of child fecal bile acid metabolism on spine morphology comprise: wherein in the model of the influence of cord blood PFAS exposure on child fecal bile acid metabolism, the corrected covariates include maternal education level, parity, gestational age at delivery, child gender, child birth weight, and weaning age; in the model of the influence of child fecal bile acid metabolism on spine morphology, the corrected covariates are the covariates; the mediating effect of fecal bile acid metabolism between cord blood PFAS exposure and child spine morphology is tested and evaluated based on a structural equation model; and the mediating effect is classified according to the test results.

[0020] ;

[0021] ;

[0022] ;

[0023] wherein, is the independent variable; is the dependent variable; is the independent variable on the dependent variable ; is the effect of the independent variable on the mediator variable ; is the effect of the mediator variable on the dependent variable , controlling for the independent variable ; is the direct effect of the independent variable on the dependent variable , controlling for the mediator variable ; the residual is denoted by ;

[0024] In the mediation effect model, the mediation effect is equivalent to the indirect effect, which is the product of and ; the relationship between the coefficients is described as:

[0025] .

[0026] Preferably, the mediation effect is classified according to the test results, comprising: performing mediation effect test by the following steps:

[0027] testing the direct effect , i.e. the effect of cord PFAS exposure on spine shape;

[0028] testing the coefficients and , i.e. the effect of cord PFAS exposure on fecal bile acid metabolism and the effect of fecal bile acid exposure on spine shape, by the GLM model;

[0029] if at least one of the coefficients and is not significantly associated, testing using the Bootstrap method, and then testing the coefficient , and confirming that and are of the same sign, by the structural equation model;

[0030] the mediation effect is classified as non-mediation effect, complete mediation effect, partial mediation effect and masking effect according to the test results, wherein the complete mediation effect is on​ The effect is entirely determined by the mediating variable Transmission, control at this time back right The direct effect disappears; the partial mediating effect indicates It directly affects And through Indirect impact The masking effect is The introduction of enhancement or reversal right The direct effects.

[0031] Preferably, the step of correcting the child's spine based on the assessment results to obtain a corrected child spine dataset includes: acquiring the child's daily life trajectory data based on the assessment results, the daily life trajectory data including the child's motion altitude data sequence; obtaining an altitude correction term based on the motion altitude data sequence using an altitude correction term formula, and removing it from the child's spine assessment results to obtain an altitude-corrected child spine dataset, wherein the altitude correction term formula is:

[0032] ;

[0033] In the formula, For the first The target variable for each child after altitude correction; For the first The original target measurements for each child; For adjustment coefficients; For the first The average altitude of the children in the exercise altitude data sequence; This is the standard altitude.

[0034] The beneficial effects of this invention are:

[0035] 1. This invention combines latent profile analysis with variance analysis of observed indicators. This method not only determines the optimal number of latent categories using the fitting index of the information criterion, but also examines the differences between categories on the original standardized index and summarizes and names them accordingly. This makes each latent category not only reasonable in terms of model fitting, but also has characteristic fingerprints that can be interpreted clinically and physiologically, which significantly improves the repeatability and readability of phenotypic definitions.

[0036] 2. This method employs an adjusted generalized linear model for single pollutant analysis and introduces weighted quantiles, regression, and multiple imputation extensions for mixed exposure analysis. It can simultaneously provide effect estimates for single components and ranking of the relative contributions of each component in mixed exposure, thus making it more targeted in risk identification and intervention target selection.

[0037] 3、The altitude correction formula provided by the present application eliminates the mixed effects of trajectory average altitude and climbing characteristics on target variables, the constructed altitude reliability weight is used for weighted regression or sample screening, can effectively reduce the system error caused by trajectory measurement noise or regional terrain difference without reducing the overall sample utilization rate, and enhances the sensitivity and specificity of the results to the true biological effect. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1 A flowchart of a child spine morphology evaluation and detection method based on multi-source data fusion is provided in the present application.

[0039] Figure 2 A mediation effect test flowchart of the child spine morphology evaluation and detection method is provided in the present application. DETAILED DESCRIPTION

[0040] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the present application.

[0041] A child spine morphology evaluation and detection method based on multi-source data fusion, as shown in Figure 1 includes the following steps:

[0042] S1: obtaining observation indexes of a child to be measured and performing pretreatment;

[0043] S2: constructing a potential profile analysis model based on standardized values, obtaining fitting indexes of each model, and performing weighted discrimination on the fitting indexes to determine the optimal number of potential categories;

[0044] S3: performing variance analysis test difference between each potential category and the standardized values of each observation index, and inducing the phenotype characteristics of each potential category and naming according to the difference direction and effect size of each index between different potential categories;

[0045] S4: performing PFAS single pollutant exposure effect evaluation using a generalized linear regression model, performing PFAS complex exposure effect evaluation using weighted quantile regression, and using a GLM model to evaluate the influence of different types of cord blood PFAS exposure on child fecal bile acid metabolism and the evaluation results of child fecal bile acid metabolism on spine morphology;

[0046] S5: correcting according to the evaluation results of the child spine to obtain a corrected child spine data set.

[0047] Collecting multiple spine and pelvis morphological observation indexes of the children to be tested, pre-processing and standardizing each observation index to obtain the standardized value of each observation index.

[0048] The multiple spine and pelvis morphological observation indexes are defined in advance, including but not limited to thoracic kyphosis angle, lumbar lordosis angle, spinal inclination angle, sacral inclination angle, pelvic height difference, pelvic inclination angle and trunk inclination; the collection method is to use standardized measurement equipment and record the measurement unit and measurement time point of each index. Each observation value should be associated with the subject identification, sampling time and sampling environment (indoor / outdoor, sampling personnel) for quality traceability; check the multiple measurement states of each child, i.e. the same follow-up or comparable time point, and if necessary, mark the data at different time points and correct or include time as a covariate in subsequent analysis; eliminate obvious data entry errors and mark outliers; identify missing as completely random missing, random missing or non-random missing, and select an interpolation strategy accordingly, using chain equation multiple interpolation method (MICE) to preserve the relationship between covariates. 20 interpolated data sets are preset; the combined estimation uses Rubin's rule to compare the variable distribution before and after interpolation to confirm that the interpolation does not introduce abnormal bias; record the number of iterations and convergence criteria of interpolation; the purpose of standardization is to eliminate the dimensional differences of different indicators and facilitate LPA and subsequent comparison, so each observation index must be converted to a standardized value.

[0049] Based on the analytic hierarchy process, the fitting indexes are weighted, and the best number of latent classes is determined according to the weighted comprehensive evaluation value, wherein the fitting indexes are AIC, BIC, AWE, CLC and KIC.

[0050] Latent profile analysis (LPA) is used to evaluate the comprehensive characteristics of the spine morphology of children of different genders, and the main steps include determining the best model, describing the parameters and naming the classes. The first step is to determine the best model, calculate the model fitting index when the number of latent classes is 1-5, and determine the best number of latent classes based on the analytic hierarchy process (AHP) of the five model fitting indexes. The five fitting indexes used for weighting and the weights are as follows:

[0051] 1) Akaike information criterion (AIC), weight 0.2323;

[0052] 2) Bayesian information criterion (BIC), weight 0.2525;

[0053] 3) Approximate Weight of Evidence (AWE), with a weight of 0.1129;

[0054] 4) Classification-Likelihood-Criterion (CLC), with a weight of 0.0922;

[0055] 5) Kullback-Information-Criterion (KIC), with a weight of 0.3101.

[0056] The hierarchical clustering method is used to perform data-driven weighting of the relative importance of fitting indices. The steps are as follows: Constructing indicator behavior vectors: The normalized index column is regarded as a behavior vector, that is, the index is the object and the performance of model m is the vector component of the object in each model; Calculating similarity / distance between indices: The similarity (e.g., Pearson correlation coefficient) or distance is calculated for each pair of indices to quantify the consistency of the performance patterns of the two indices in different models; Hierarchical clustering implementation: Hierarchical clustering is implemented on the distance matrix of the indices, for example, using the Ward method or the average connectivity method, to generate a dendrogram of the indices; Determining the clustering scheme: Based on the number of clusters obtained by the elbow method or the pruning dendrogram, the 5 fitting indices are divided into several clusters. Each cluster index is considered a subset with similar behavior; intra-cluster variability is calculated for the indices, for example, the sum of the variances of the indices across all models. The smaller the intra-cluster variability, the more stable the cluster index is in model evaluation and the more reliable it is as a source of evaluation. The index weights are assigned as shown above, and a weighted summation method is used to obtain the comprehensive score. The candidate models are ranked according to the comprehensive score, and the model with the highest comprehensive score is selected as the optimal number of potential categories.

[0057] Analysis of variance was used to detect the differences between children's spinal morphology and each of the optimal potential categories. Based on the direction and effect size of the differences of each indicator among different potential categories, the phenotypic characteristics of each potential category were summarized and named. The prior probability and potential category probability of the test subjects were obtained. The prior probability reflects the relative proportion of each potential category in the total population, and the sum of the prior probabilities of all potential categories is 1. The potential category probability reflects the probability that an individual belongs to each potential category under the conditions of the observed data, and the sum of the probabilities of all potential categories of each individual is 1.

[0058] Overall framework of the steps: (1) Input: the standardized value matrix of each observation index and the LPA optimal model partitioning result, the number of categories given by the model and the posterior probability of each category corresponding to each observation record.

[0059] (2) Test: Analysis of variance (ANOVA) is performed for each observed indicator to determine whether there is a significant difference between groups. After testing, the phenotypic characteristics of each latent class are summarized and named according to the direction and effect size of the difference.

[0060] (3) Probability calculation: The prior probability of each latent class (the proportion of the population) and the latent class probability of each person being tested (the posterior membership probability output by LPA) are calculated, and the above results are output and archived in tabular form;

[0061] The second step of LPA is to describe parameters and name classes. First, ANOVA is used to evaluate the differences in each spinal shape score between different latent classes, and the characteristics of each class are summarized and named. Second, the prior probability and latent class probability are calculated. The prior probability reflects the relative proportion of each latent class in the total population, and the sum of the prior probabilities of all latent classes is 1. The greater the prior probability, the more common the latent class in the sample. The latent class with the highest prior probability is usually considered a typical group, but the typical state also needs to be further judged in combination with the feature distribution. The latent class probability reflects the likelihood of an individual belonging to each latent class given the observed data, and the sum of all latent class probabilities for each individual is 1. The higher the latent class probability, the more consistent the individual's characteristics with the distribution characteristics of the latent class;

[0062] Class characteristic induction and naming rules: For each latent class, the characteristics are summarized in the following order:

[0063] List the indicators that are significantly higher than the overall or other classes (after post-hoc comparison) in this class (and give the effect size); list the indicators that are significantly lower than the overall or other classes; evaluate the biological / clinical meaning of these indicators and summarize them with a short descriptive name, such as "excessive curvature" or "pelvic tilt";

[0064] Naming rules: If there are ≥2 key morphological indicators that are significantly different and consistent in direction (both positive or negative) with an effect size ≥ moderate, name them as "something type (main feature)";

[0065] If only a single indicator is significantly different with a small effect size, add "mild / marginal" to the name to avoid medical diagnostic language;

[0066] The name should be reviewed and agreed upon by at least two experts in the field (e.g., pediatric / orthopedic or human engineering experts).

[0067] Obtaining covariates of the to-be-tested personnel, the covariates including maternal education level, parity, gestational age at delivery, child gender, birth weight of the child, weaning age, age at spine follow-up and length of weekly vigorous physical activity at spine follow-up; performing multiple imputation of the covariates based on a chain equation multiple imputation algorithm to generate a preset number of imputed data sets, merging statistical results of all the imputed data sets using Rubin's rule, and calculating overall parameter estimates, standard errors and pre-set confidence intervals; and analyzing the association between cord blood PFAS exposure and child spine morphology according to the covariates using a generalized linear regression model.

[0068] Input data and variable definition:

[0069] Independent variable (X): single PFAS concentration (cord blood measurement value), if the sample contains multiple PFAS, each PFAS is analyzed as a single pollutant;

[0070] Outcome variable (Y): child spine morphology index, which is a standardized continuous index (z-score) or vector (item-by-item regression), and z-score is recommended as the outcome to be compatible with LPA input.

[0071] Covariate set (Z): the covariates explicitly listed and used in the examples include maternal education level, parity, gestational age at delivery, child gender, birth weight of the child, weaning age, age at spine follow-up (accurate to year) and length of weekly vigorous physical activity at spine follow-up.

[0072] Collection and preprocessing of covariates: coding rules: define a coding table for categorical covariates (such as maternal education level); preparation before missing value processing: perform preliminary logical checks and outlier labeling on all variables for imputation to ensure the stability of the imputation model;

[0073] GLM model details and implementation suggestions: to improve the skew distribution of PFAS, it is preferred to take the logarithmic transformation of X; model selection: use linear regression (ordinary least squares) for continuous outcomes, and use robust standard error (Huber-White) when there is heteroscedasticity; use Poisson or negative binomial regression for count outcomes; interaction and stratification: stratified analysis is performed on child gender. Implementation includes two types of implementation: implementation mode A (stratification): split the sample into two groups according to gender and repeat the imputation and regression analysis for each group; implementation mode B (interaction term): add an interaction term Xxsex in a single model and test the interaction coefficient (if significant, do stratified interpretation). Model diagnosis: check the normality of residuals, heteroscedasticity and high leverage points for linear models; check the goodness of fit and influential points for logistic models.

[0074] Generalized linear regression models were used to analyze the association between PFAS exposure and spinal shape in childhood. Stratified analysis was performed according to the gender of the child. Based on previous literature, the covariates adjusted in the model included maternal education level, parity, gestational age at birth, gender of the child, birth weight of the child, age at weaning, age at spinal follow-up (rounded to the nearest year), and duration of weekly vigorous physical activity at the time of spinal follow-up. Based on the chain equation multiple imputation algorithm, 20 imputed datasets were generated, and the statistical results of all imputed datasets were combined using the Rubin rule to calculate the overall parameter estimates, standard errors, and 95% confidence intervals. To assess the robustness of the study results, a sensitivity analysis was performed to examine the impact of age-adjusted BMI as a potential confounding variable on the association between PFAS exposure and spinal shape in children. Based on the above model, further control of age-adjusted BMI, a multiple linear regression model was constructed to compare the differences in the effects of PFAS exposure on spinal shape.

[0075] The concentrations of a predetermined number of types of PFAS in the test sample are pre-processed and multiple imputed, each PFAS level is quantified according to the quantile, and a weighted quantile sum regression model is constructed, in which the estimation of mixed effects of both positive and negative types is realized, and the weight of each PFAS is obtained through repeated sampling or cross-validation, the mixed exposure index is calculated using the weight, and the mixed exposure index is included as an independent variable in the linear regression or generalized linear regression model to evaluate the overall effect of mixed exposure on the outcome of spinal shape, and the overall effect estimate of mixed exposure, the relative weight order of each component, and the significantly different outcome indicators under the positive model or the negative model are output.

[0076] The combined exposure effect of 8 PFAS mixtures was calculated using weighted quantile sum regression. WQS is a weighted quartile sum method combined with linear regression, which can not only estimate the comprehensive positive and negative effects of mixtures, but also calculate the importance weight of each exposure component based on multiple imputation. All statistical analyses were completed in R version 4.4.2. Multiple imputation used the "mice" package, data cleaning used the "tidyverse" framework, latent profile analysis used the "tidyLPA" and "mclust" packages, linear regression analysis used the "tidymodels" framework, and combined exposure analysis used the "gWQS" and "miWQS" packages. The "ggplot2" package was used for plotting, and the "gt" and "gtsummary" packages were used for table drawing. Statistical tests were two-sided, with a test level of 0.05.

[0077] Among the models of the influence of cord PFAS exposure on fecal bile acid metabolism in children, the covariates corrected included: maternal education level, parity, gestational age at delivery, child gender, child's birth weight, and weaning age; among the models of the influence of fecal bile acid metabolism in children on spinal shape, the covariates corrected were the covariates; the mediation effect of fecal bile acid metabolism between cord PFAS exposure and spinal shape in children was tested and evaluated based on structural equation model; and the mediation effect was classified according to the test results.

[0078] The relationship between variables in the mediation effect model is described as:

[0079] ;

[0080] ;

[0081] ;

[0082] In the formula, is the independent variable; is the dependent variable; is the independent variable on the dependent variable ; is the effect of the independent variable on the mediator variable ; is the influence of the mediator variable on the dependent variable under the control of the independent variable ; is the direct effect of the independent variable on the dependent variable under the control of the mediator variable ; the residual error is represented by ;

[0083] In the mediation effect model, the mediation effect is equivalent to the indirect effect, and its size is the product of and ;the relationship between the coefficients is described as:

[0084] .

[0085] The mediation effect test of this study is divided into the following three steps:

[0086] 1) Test the direct effect (coefficient c), that is, the influence of cord PFAS exposure on spinal shape. This test has been completed in the first part of the study, and the significant correlation is based on the mediation effect, and the remaining correlation is based on the masking effect, all of which are included in the subsequent analysis.

[0087] 2) Test the coefficients a and b, i.e. the effect of cord PFAS exposure on fecal bile acid metabolism and the effect of fecal bile acid exposure on spine shape, in sequence. This test is done in the GLM model of this part of the study.

[0088] 3) For at least one of the coefficients a and b not being significant, test ab using the Bootstrap method, then test the coefficients and confirm whether ab and are of the same sign. This test is done using structural equation modeling.

[0089] As shown in Figure 2 , the three kinds of mediation effects described above describe different mechanisms of how the independent variable X affects the dependent variable Y, where the complete mediation effect means that the effect of X on Y is completely transmitted by the mediator M, which is understood as "X must pass through M to affect Y", at which time the direct effect of X on Y disappears after controlling M; the partial mediation effect indicates that X directly acts on Y and indirectly affects Y through M, which is understood as "X has both direct and indirect influence paths on Y"; and the masking effect means that the introduction of M enhances or reverses the direct effect of X on Y, i.e. "M masks the true effect of X", or even makes the originally insignificant X-Y relationship significant.

[0090] The mediation effect test is performed by the following steps:

[0091] Test the direct effect , i.e. the effect of cord PFAS exposure on spine shape;

[0092] Test the coefficients and , i.e. the effect of cord PFAS exposure on fecal bile acid metabolism and the effect of fecal bile acid exposure on spine shape, through the GLM model;

[0093] For at least one of the coefficients and not being significant, test ab using the Bootstrap method, then test the coefficients , and confirm whether and are of the same sign, through structural equation modeling;

[0094] According to the test results, the mediation effect is classified as no mediation effect, complete mediation effect, partial mediation effect and masking effect, where the complete mediation effect is the effect of is completely transmitted by the mediator Transmission, control at this time back right The direct effect disappears; the partial mediating effect indicates It directly affects And through Indirect impact The masking effect is The introduction of enhancement or reversal right The direct effects.

[0095] Based on the assessment results of the child's spine, the child's daily life trajectory data is obtained, which includes the child's movement altitude data sequence. An altitude correction term is obtained from the movement altitude data sequence using an altitude correction term formula, and then removed from the child's spine assessment results to obtain an altitude-corrected child spine dataset. The altitude correction term formula is as follows:

[0096] ;

[0097] In the formula, For the first The target variable for each child after altitude correction; For the first The original target measurements for each child; For adjustment coefficients; For the first The average altitude of the children in the exercise altitude data sequence; This is the standard altitude.

[0098] To adjust the coefficients, by Obtain, in the formula For the first Other covariate matrices for each child, such as sex and birth weight; The corresponding coefficients were obtained through historical experience; This is the coefficient for the altitude deviation term; To obtain the value that minimizes the sum of squared residuals value; The standard altitude is set by region when there is significant heterogeneity in the region or population.

[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A child spine shape evaluation and detection method based on multi-source data fusion, characterized in that, Includes the following steps: S1: Obtain the observation indicators of the children to be tested and perform preprocessing; S2: Construct a potential profile analysis model based on standardized values, obtain the fitting index of each model, and perform weighted discrimination on the fitting index to determine the optimal number of potential categories; S3: Perform ANOVA on the standardized values ​​of each observation indicator for each potential category to test for differences, and summarize and name the phenotypic characteristics of each potential category based on the direction and magnitude of the differences between each indicator and different potential categories. S4: The generalized linear regression model was used to assess the single pollutant exposure effect of PFAS, the weighted quantile and regression model was used to assess the combined exposure effect of PFAS, and the GLM model was used to assess the effects of different types of umbilical cord blood PFAS exposure on fecal bile acid metabolism in children and the assessment results of fecal bile acid metabolism in children on spinal morphology. S5: Correct the spine based on the assessment results of the child's spine and obtain a dataset of the corrected child's spine.

2. The method of claim 1, wherein the method is based on multi-source data fusion for children's spine shape assessment and detection. The process of acquiring and preprocessing the observation indicators of the children to be tested includes: collecting multiple morphological observation indicators of the spine and pelvis of the children to be tested, preprocessing and standardizing each observation indicator, and obtaining the standardized value of each observation indicator.

3. The method of claim 2, wherein the method is based on multi-source data fusion for children's spine shape assessment and detection. The process of constructing a potential profile analysis model based on standardized values, obtaining a fitting index for each model, and weighting the fitting index to determine the optimal number of potential categories includes: weighting the fitting index based on hierarchical clustering analysis, and determining the optimal number of potential categories based on the weighted comprehensive evaluation value, wherein the fitting index is AIC, BIC, AWE, CLC, and KIC.

4. The method of claim 1, wherein the method is based on multi-source data fusion for children's spine shape assessment and detection. The process of performing ANOVA on the standardized values ​​of each observation indicator for each potential category to test for differences, and summarizing and naming the phenotypic characteristics of each potential category based on the direction and effect size of the differences between each indicator and different potential categories, further includes: using ANOVA to detect the difference between the child's spinal morphology and each of the optimal potential categories; summarizing and naming the phenotypic characteristics of each potential category based on the direction and effect size of the differences between each indicator and different potential categories; and obtaining the prior probability and potential category probability of the test subject. The prior probability reflects the relative proportion of each potential category in the total population, and the sum of the prior probabilities of all potential categories is 1. The potential category probability reflects the probability that an individual belongs to each potential category under the conditions of the observed data, and the sum of the probabilities of all potential categories for each individual is 1.

5. The method of claim 1, wherein the method is based on multi-source data fusion for children's spine shape assessment and detection. The PFAS single pollutant exposure effect evaluation by using the generalized linear regression model comprises the following steps: obtaining the covariates of the to-be-tested personnel, wherein the covariates comprise the mother's education level, the parity, the gestational age at delivery, the child's gender, the child's birth weight, the weaning age, the age at the time of spinal follow-up, and the length of weekly strenuous physical activity at the time of spinal follow-up; performing multiple imputation on the covariates based on a chain equation multiple imputation algorithm to generate a preset number of imputed data sets, combining the statistical results of all the imputed data sets by using the Rubin rule, and calculating the overall parameter estimate, the standard error, and the preset confidence interval; and analyzing the correlation between the cord blood PFAS exposure and the child's spinal morphology by using the generalized linear regression model according to the covariates.

6. The method of claim 1, wherein the method is based on multi-source data fusion for children's spine shape assessment and detection. The PFAS composite exposure effect evaluation by using the weighted quantile regression comprises the following steps: preprocessing and multiple imputation are performed on a preset number of types of PFAS concentrations in the to-be-tested sample, each type of PFAS level is quantified according to a quantile, and a weighted quantile regression model is constructed, in which the estimation of mixed effects of two types of positive and negative effects is realized, the weight of each type of PFAS is obtained through repeated sampling or cross-validation, the mixed exposure index is calculated by using the weight, the mixed exposure index is taken as an independent variable and is introduced into a linear regression or generalized linear regression model, the overall effect of mixed exposure on the spinal morphology outcome is evaluated, and the overall effect estimate of mixed exposure, the relative weight ranking of each component, and the outcome indicators that are respectively significant under a positive model or a negative model are output.

7. The method of claim 1, wherein the method is based on multi-source data fusion for children's spine shape assessment and detection. The evaluation results of the influence of different types of cord blood PFAS exposure on the child's fecal bile acid metabolism and the evaluation results of the child's fecal bile acid metabolism on the spinal morphology comprise the following steps: in the model of the influence of cord blood PFAS exposure on the child's fecal bile acid metabolism, the corrected covariates comprise the mother's education level, the parity, the gestational age at delivery, the child's gender, the child's birth weight, and the weaning age; in the model of the influence of the child's fecal bile acid metabolism on the spinal morphology, the corrected covariates are the covariates; the mediating effect of the fecal bile acid metabolism between the cord blood PFAS exposure and the child's spinal morphology is tested and evaluated based on a structural equation model; and the mediating effect is classified according to the test results.

8. The method of claim 7, wherein the method is based on multi-source data fusion for spinal morphology assessment of children. The mediating effect of the fecal bile acid metabolism between the cord blood PFAS exposure and the child's spinal morphology is tested and evaluated based on a structural equation model, and the relationship between variables in the mediating effect model is described as follows: ; ; ; where is the dependent variable; is the dependent variable; is the dependent variable on the mediator variable ; the total effect of on the dependent variable; is the effect of on the dependent variable , controlling for the mediator variable ; the influence of on the dependent variable , controlling for the mediator variable ; the direct effect of on the dependent variable , controlling for the mediator variable ; the residual variance is represented by ; In the mediation effect model, the mediation effect is equivalent to the indirect effect, and its size is the product of and ; the relationship between the coefficients is described as: 。 9. The method of claim 7, wherein the method is based on multi-source data fusion for spine morphology assessment of children. The mediating effect is classified according to the test results, and the mediating effect test is performed by the following steps: Testing direct effects i.e. the effect of cord blood PFAS exposure on spinal shape; Test coefficients And The effects of cord blood PFAS exposure on fecal bile acid metabolism and the effects of fecal bile acid exposure on spinal shape were completed by the GLM models. For the coefficients and At least one insignificant association, using the Bootstrap method to test Then test the coefficients And confirm and The same number of states, through structural equation modeling; Based on the test results, the mediation effect is classified into four categories: no mediation effect, complete mediation effect, partial mediation effect, and masking effect. The complete mediation effect is... right The effect is entirely determined by the mediating variable Transmission, control at this time back right The direct effect disappears; the partial mediating effect indicates It directly affects And through Indirect impact The masking effect is The introduction of enhancement or reversal right The direct effects.

10. The method of claim 1, wherein the method is based on multi-source data fusion for children's spine shape assessment and detection. The child's spinal data set after correction is obtained according to the evaluation results of the child's spinal column, and the steps comprise the following steps: obtaining the child's life trajectory data according to the evaluation results of the child's spinal column, wherein the life trajectory data comprise a child's movement altitude data sequence; obtaining an altitude correction term from the child's spinal evaluation results by using an altitude correction term formula, and obtaining an altitude-corrected child's spinal data set, wherein the altitude correction term formula is as follows: ; In the formula, For the first The target variable for each child after altitude correction; For the first The original target measurements for each child; For adjustment coefficients; For the first The average altitude of the children in the exercise altitude data sequence; This is the standard altitude.

Citation Information

Patent Citations

  • Spine monitoring method and device for teenagers and children, electronic equipment and storage medium

    CN115153514A

  • Large-scale mixed exposure data analysis method based on machine learning

    CN116738172A