Children spine form evaluation and detection method based on multi-source data fusion

The method for assessing pediatric spinal morphology by fusing multi-source data solves the problems of reliance on single indicators and difficulty in distinguishing the effects of mixed contaminants, achieving more accurate assessment of pediatric spinal morphology and improving the repeatability of phenotype definition and risk identification capabilities.

CN120954732AActive Publication Date: 2025-11-14CHILDRENS HOSPITAL OF FUDAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511467683.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2025-11-14
Estimated Expiration
2045-10-15

AI Technical Summary

Technical Problem

Existing technologies for assessing spinal morphology in children suffer from problems such as reliance on single indicators, difficulty in distinguishing the effects of mixed contaminants, and a lack of multivariate joint models in data analysis methods. This makes it difficult to reflect heterogeneity within the population and to explain the correlation between morphological phenotypes and molecular/metabolic indicators one-to-one.

Method used

A multi-source data fusion method was adopted to construct a potential profile analysis model by acquiring and preprocessing observation indicators, and to conduct analysis of variance and generalized linear regression to evaluate the PFAS exposure effect. The mixed exposure effect was analyzed using weighted quantiles and regression models, and noise was removed by combining the altitude correction formula to obtain the corrected children's spine dataset.

Benefits of technology

It significantly improves the repeatability and readability of phenotypic definitions, enables targeted risk identification and intervention selection, reduces systematic errors, and enhances the sensitivity and specificity of results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954732A_ABST
    Figure CN120954732A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of spine assessment and treatment, in particular to a children spine form assessment and detection method based on multi-source data fusion. Comprising the following steps: S1, obtaining observation indexes of a to-be-detected child and performing preprocessing; s2, constructing a potential section analysis model based on the standardized value, obtaining a fitting index of each model, and carrying out weighted discrimination on the fitting indexes to determine an optimal potential category number; and S3, performing variance analysis on each potential category among the standardized values of the observation indexes to check differences, and summarizing phenotypic features of each potential category according to difference directions and effect sizes of the indexes among different potential categories and naming the phenotypic features. According to the method, the adjusted generalized linear model is adopted on single pollutant analysis, and weighted quantiles, regression and multi-interpolation expansion thereof are introduced on mixed exposure analysis, so that the method is more targeted in risk identification and intervention target selection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spinal assessment and treatment technology, and in particular to a method for assessing and detecting spinal morphology in children based on multi-source data fusion. Background Technology

[0002] As research into the impact of environmental pollutants on child development deepens, long-chain and short-chain polyfluoroalkyl and polyalkyl substances (collectively referred to as PFAS), as a class of chemicals with persistent and bioaccumulative characteristics, are increasingly attracting attention in studies on their potential effects on child growth and development, metabolic pathways, and musculoskeletal system during perinatal and prenatal exposure. Several existing epidemiological studies have reported correlations between PFAS and children's metabolic markers, gut-related metabolites, and individual physiological development indicators; however, these studies generally face the following limitations:

[0003] First, there is the limitation of phenotypic characterization and reliance on subjective thresholds. Traditional spinal morphology assessments often rely on single indicators or preset thresholds to determine abnormalities, lacking a classification strategy based on multivariate joint models. This makes it difficult to reflect the heterogeneity within a population and also hinders the one-to-one correspondence between morphological phenotypes and molecular / metabolic indicators for interpretation and subsequent correlation analysis.

[0004] Second, it is difficult to distinguish between mixed exposures and the effects of single pollutants. Environmental exposures are usually multi-component mixtures, and regression analysis using a single pollutant may miss the synergistic or antagonistic effects between components. Furthermore, existing mixture analysis methods are not yet closely coupled with phenotyping, behavioral trajectories, or spatial environmental factors in terms of handling measurement gaps, multiple corrections, and weight estimation.

[0005] In view of the above problems, there is an urgent need for an evaluation method based on multi-source data. Summary of the Invention

[0006] To overcome the problems mentioned in the background art, the present invention provides a method for assessing and detecting spinal morphology in children based on multi-source data fusion.

[0007] Technical solution: A method for assessing and detecting spinal morphology in children based on multi-source data fusion, comprising the following steps:

[0008] S1: Obtain the observation indicators of the children to be tested and perform preprocessing;

[0009] S2: Construct a potential profile analysis model based on standardized values, obtain the fitting index of each model, and perform weighted discrimination on the fitting index to determine the optimal number of potential categories;

[0010] S3: Perform ANOVA on the standardized values ​​of each observation indicator for each potential category to test for differences, and summarize and name the phenotypic characteristics of each potential category based on the direction and magnitude of the differences between each indicator and different potential categories.

[0011] S4: The generalized linear regression model was used to assess the single pollutant exposure effect of PFAS, the weighted quantile and regression model was used to assess the combined exposure effect of PFAS, and the GLM model was used to assess the effects of different types of umbilical cord blood PFAS exposure on fecal bile acid metabolism in children and the assessment results of fecal bile acid metabolism in children on spinal morphology.

[0012] S5: Correct the spine based on the assessment results of the child's spine and obtain a dataset of the corrected child's spine.

[0013] Preferably, the step of acquiring and preprocessing the observation indicators of the child to be tested includes: collecting multiple spinal and pelvic morphological observation indicators of the child to be tested, preprocessing and standardizing each observation indicator to obtain the standardized value of each observation indicator.

[0014] Preferably, the step of constructing a potential profile analysis model based on standardized values, obtaining a fitting index for each model, and weighting the fitting index to determine the optimal number of potential categories includes: weighting the fitting index based on hierarchical clustering analysis, and determining the optimal number of potential categories based on the weighted comprehensive evaluation value, wherein the fitting index is AIC, BIC, AWE, CLC, and KIC.

[0015] Preferably, the step of performing ANOVA to test the differences between the standardized values ​​of each observation index for each potential category, and summarizing and naming the phenotypic characteristics of each potential category based on the direction and effect size of the differences between each index and different potential categories, further includes: using ANOVA to detect the differences between the child's spinal morphology and each of the optimal potential categories; summarizing and naming the phenotypic characteristics of each potential category based on the direction and effect size of the differences between each index and different potential categories; and then obtaining the prior probability and potential category probability of the test subject. The prior probability reflects the relative proportion of each potential category in the total population, and the sum of the prior probabilities of all potential categories is 1. The potential category probability reflects the probability that an individual belongs to each potential category under the conditions of the observed data, and the sum of the probabilities of all potential categories for each individual is 1.

[0016] Preferably, the assessment of the single pollutant exposure effect of PFAS using a generalized linear regression model includes: obtaining covariates for the subjects, wherein the covariates include maternal education level, parity, gestational age at delivery, child sex, child's birth weight, weaning age, age at spinal follow-up, and weekly duration of vigorous physical activity at spinal follow-up; imputing the covariates based on a chain equation multiple imputation algorithm to generate a preset number of imputed datasets; merging the statistical results of all imputed datasets using Rubin's rule; calculating the overall parameter estimate, standard error, and preset confidence interval; and analyzing the association between umbilical cord blood PFAS exposure and children's spinal morphology using a generalized linear regression model based on the covariates.

[0017] Preferably, the assessment of the combined PFAS exposure effect using weighted quantiles and regression includes: preprocessing and multiple imputation of the concentrations of a predetermined number of PFAS types in the sample to be tested; quantifying the level of each PFAS by quantiles; constructing a weighted quantile and regression model; estimating the positive and negative mixed effects in the weighted quantile and regression model; obtaining the weight of each PFAS through repeated sampling or cross-validation; calculating the mixed exposure index using the weights; and incorporating the mixed exposure index as an independent variable into a linear regression or generalized linear regression model to assess the overall effect of mixed exposure on spinal morphological outcomes; and outputting the overall effect estimate of mixed exposure, the relative weight ranking of each component, and the significant outcome indicators under the positive or negative models, respectively.

[0018] Preferably, the evaluation of the effects of different types of umbilical cord blood PFAS exposure on fecal bile acid metabolism in children and the evaluation results of the effects of fecal bile acid metabolism on spinal morphology using the GLM model includes: in the model of the effect of umbilical cord blood PFAS exposure on fecal bile acid metabolism in children, the corrected covariates include: maternal education level, parity, gestational age at delivery, child sex, child's birth weight, and weaning age; in the model of the effect of fecal bile acid metabolism on spinal morphology in children, the corrected covariates are the aforementioned covariates; the mediating role of fecal bile acid metabolism between umbilical cord blood PFAS exposure and spinal morphology in children is tested and evaluated based on structural equation modeling; and the mediating effect is classified according to the test results.

[0019] Preferably, the structural equation model-based testing and evaluation of the mediating role of fecal bile acid metabolism in the relationship between umbilical cord blood PFAS exposure and spinal morphology in children includes: describing the relationship between variables in the mediation effect model as follows:

[0020] ;

[0021] ;

[0022] ;

[0023] In the formula, As the independent variable; The dependent variable; as independent variable For dependent variable The overall effect; To reflect the independent variable For mediating variables Its function; In order to control the independent variable In the case of mediator variables For dependent variable The impact; To control for mediating variables In the case of independent variable For dependent variable The direct effect of residuals; express;

[0024] In the mediation effect model, the mediation effect is equivalent to the indirect effect, and its magnitude is... and product The relationship between the coefficients can be described as follows:

[0025] .

[0026] Preferably, classifying the mediation effect based on the test results includes: performing a mediation effect test through the following steps:

[0027] Testing direct effects The effect of umbilical cord blood PFAS exposure on spinal morphology;

[0028] Test coefficient and The effects of umbilical cord blood PFAS exposure on fecal bile acid metabolism and fecal bile acid exposure on spinal morphology were investigated using the GLM model.

[0029] For coefficients and There is at least one non-significant association; use the Bootstrap test. Then test the coefficients. and confirm and The same sign state is achieved through structural equation modeling;

[0030] Based on the test results, the mediation effect is classified into four categories: no mediation effect, complete mediation effect, partial mediation effect, and masking effect. The complete mediation effect is... right The effect is entirely determined by the mediating variable Transmission, control at this time back right The direct effect disappears; the partial mediating effect indicates It directly affects And through Indirect impact The masking effect is The introduction of enhancement or reversal right The direct effects.

[0031] Preferably, the step of correcting the child's spine based on the assessment results to obtain a corrected child spine dataset includes: acquiring the child's daily life trajectory data based on the assessment results, the daily life trajectory data including the child's motion altitude data sequence; obtaining an altitude correction term based on the motion altitude data sequence using an altitude correction term formula, and removing it from the child's spine assessment results to obtain an altitude-corrected child spine dataset, wherein the altitude correction term formula is:

[0032] ;

[0033] In the formula, For the first The target variable for each child after altitude correction; For the first The original target measurements for each child; For adjustment coefficients; For the first The average altitude of the children in the exercise altitude data sequence; This is the standard altitude.

[0034] The beneficial effects of this invention are:

[0035] 1. This invention combines latent profile analysis with variance analysis of observed indicators. This method not only determines the optimal number of latent categories using the fitting index of the information criterion, but also examines the differences between categories on the original standardized index and summarizes and names them accordingly. This makes each latent category not only reasonable in terms of model fitting, but also has characteristic fingerprints that can be interpreted clinically and physiologically, which significantly improves the repeatability and readability of phenotypic definitions.

[0036] 2. This method employs an adjusted generalized linear model for single pollutant analysis and introduces weighted quantiles, regression, and multiple imputation extensions for mixed exposure analysis. It can simultaneously provide effect estimates for single components and ranking of the relative contributions of each component in mixed exposure, thus making it more targeted in risk identification and intervention target selection.

[0037] 3. The altitude correction formula proposed in this invention eliminates the confounding effect of trajectory average altitude and climbing characteristics on the target variable. The constructed altitude confidence weight is used for weighted regression or sample screening. It can effectively reduce systematic errors caused by trajectory measurement noise or regional terrain differences without reducing the overall sample utilization rate, and enhance the sensitivity and specificity of the results to real biological effects. Attached Figure Description

[0038] Figure 1 This is a flowchart of a method for assessing and detecting spinal morphology in children based on multi-source data fusion, according to the present invention.

[0039] Figure 2 This is a flowchart of the mediation effect test for the pediatric spinal morphology assessment and detection method of the present invention. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] A method for assessing and detecting pediatric spinal morphology based on multi-source data fusion, such as Figure 1 As shown, it includes the following steps:

[0042] S1: Obtain the observation indicators of the children to be tested and perform preprocessing;

[0043] S2: Construct a potential profile analysis model based on standardized values, obtain the fitting index of each model, and perform weighted discrimination on the fitting index to determine the optimal number of potential categories;

[0044] S3: Perform ANOVA on the standardized values ​​of each observation indicator for each potential category to test for differences, and summarize and name the phenotypic characteristics of each potential category based on the direction and magnitude of the differences between each indicator and different potential categories.

[0045] S4: The generalized linear regression model was used to assess the single pollutant exposure effect of PFAS, the weighted quantile and regression model was used to assess the combined exposure effect of PFAS, and the GLM model was used to assess the effects of different types of umbilical cord blood PFAS exposure on fecal bile acid metabolism in children and the assessment results of fecal bile acid metabolism in children on spinal morphology.

[0046] S5: Correct the spine based on the assessment results of the child's spine and obtain a dataset of the corrected child's spine.

[0047] Multiple spinal and pelvic morphological observation indicators were collected from the children to be tested. Each observation indicator was preprocessed and standardized to obtain the standardized value of each observation indicator.

[0048] The various spinal and pelvic morphological observation indicators were predefined, including but not limited to: thoracic lordosis angle, lumbar lordosis angle, spinal lateral tilt angle, sacral tilt angle, pelvic height difference, pelvic anterior tilt angle, and trunk tilt. Data collection was conducted using standardized measuring equipment, and the measurement unit and time point for each indicator were recorded. Each observation value was associated with the subject identification, sampling time, and sampling environment (indoor / outdoor, sampling personnel) for quality traceability. Multiple measurement statuses for each child were verified, i.e., the same follow-up or comparable time points. Data from different time points were noted if necessary and corrected for time in subsequent analyses or included as covariates. Obvious data entry errors were eliminated, and outliers were marked. Missing data were identified as completely random, random, or non-random, and imputation strategies were selected accordingly. Multiple interpolation using chain equations (MICE) was employed to preserve relationships between covariates. 20 imputation datasets are generated by default; the Rubin rule is used to merge the estimates and compare the distribution of variables before and after imputation to confirm that imputation will not introduce abnormal bias; the number of imputation iterations and convergence criteria are recorded; the purpose of standardization is to eliminate the differences in the scale of different indicators and facilitate LPA and subsequent comparisons, each observation indicator must be uniformly converted into a standardized value.

[0049] The fit indices are weighted based on hierarchical clustering analysis, and the optimal number of potential categories is determined based on the weighted comprehensive evaluation value. The fit indices are AIC, BIC, AWE, CLC, and KIC.

[0050] Latent-Profile Analysis (LPA) was used to assess the comprehensive spinal morphological characteristics of children of different sexes. The main steps included determining the optimal model, descriptive parameters, and category naming. The first step was to determine the optimal model by calculating model fit indices for a range of 1-5 latent categories. Then, based on Analytic Hierarchy-Process (AHP) analysis, the five model fit indices were weighted to determine the optimal number of latent categories. The five fit indices used for weighting and their weights are as follows:

[0051] 1) Akaike Information Criterion (AIC), with a weight of 0.2323;

[0052] 2) Bayesian Information Criterion (BIC), with a weight of 0.2525;

[0053] 3) Approximate Weight of Evidence (AWE), with a weight of 0.1129;

[0054] 4) Classification-Likelihood-Criterion (CLC), with a weight of 0.0922;

[0055] 5) Kullback-Information-Criterion (KIC), with a weight of 0.3101.

[0056] The hierarchical clustering method is used to perform data-driven weighting of the relative importance of fitting indices. The steps are as follows: Constructing indicator behavior vectors: The normalized index column is regarded as a behavior vector, that is, the index is the object and the performance of model m is the vector component of the object in each model; Calculating similarity / distance between indices: The similarity (e.g., Pearson correlation coefficient) or distance is calculated for each pair of indices to quantify the consistency of the performance patterns of the two indices in different models; Hierarchical clustering implementation: Hierarchical clustering is implemented on the distance matrix of the indices, for example, using the Ward method or the average connectivity method, to generate a dendrogram of the indices; Determining the clustering scheme: Based on the number of clusters obtained by the elbow method or the pruning dendrogram, the 5 fitting indices are divided into several clusters. Each cluster index is considered a subset with similar behavior; intra-cluster variability is calculated for the indices, for example, the sum of the variances of the indices across all models. The smaller the intra-cluster variability, the more stable the cluster index is in model evaluation and the more reliable it is as a source of evaluation. The index weights are assigned as shown above, and a weighted summation method is used to obtain the comprehensive score. The candidate models are ranked according to the comprehensive score, and the model with the highest comprehensive score is selected as the optimal number of potential categories.

[0057] Analysis of variance was used to detect the differences between children's spinal morphology and each of the optimal potential categories. Based on the direction and effect size of the differences of each indicator among different potential categories, the phenotypic characteristics of each potential category were summarized and named. The prior probability and potential category probability of the test subjects were obtained. The prior probability reflects the relative proportion of each potential category in the total population, and the sum of the prior probabilities of all potential categories is 1. The potential category probability reflects the probability that an individual belongs to each potential category under the conditions of the observed data, and the sum of the probabilities of all potential categories of each individual is 1.

[0058] Overall framework of the steps: (1) Input: the standardized value matrix of each observation index and the LPA optimal model partitioning result, the number of categories given by the model and the posterior probability of each category corresponding to each observation record.

[0059] (2) Test: Perform ANOVA on each observation indicator among the potential categories to determine the state of significant differences between groups; after the test, summarize the phenotypic characteristics of each potential category according to the direction of difference and effect size and name them.

[0060] (3) Probability calculation: Calculate the prior probability (group proportion) of each potential category and the potential category probability of each examinee (posterior membership probability output by LPA), and output the above results in tabular form and archive them;

[0061] The second step of LPA is to describe the parameters and name the categories. First, analysis of variance is used to assess the differences in spinal morphology scores across different latent categories, summarizing the characteristics of each category and naming them. Second, prior probabilities and latent category probabilities are calculated. Prior probabilities reflect the relative proportion of each latent category in the total population; the sum of the prior probabilities of all latent categories is 1. A higher prior probability indicates that the latent category is more prevalent in the sample. The latent category with the highest prior probability is usually considered a typical group, but typicality needs further judgment based on feature distribution. Latent category probabilities reflect the likelihood that an individual belongs to each latent category given the observed data; the sum of the probabilities of all latent categories for each individual is 1. A higher latent category probability indicates that the individual's characteristics better match the distribution characteristics of that latent category.

[0062] Categorical Feature Induction and Naming Rules: For each potential category, features are induction in the following order:

[0063] List the measures that are significantly higher in this class than in the general population or other classes (after post-hoc comparison) (and give the effect size); list the measures that are significantly lower than in the general population or other classes; assess the biological / clinical meaning of these measures and summarize them with short descriptive names, such as "excessive flexion" or "pelvic tilt type";

[0064] Naming rules: If there are ≥2 significant key morphological indicators with consistent direction (both positive or negative) and the effect size is ≥ moderate, then it is named "Type X (main feature)".

[0065] If only a single indicator is significant and has a small effect size, add "mild / borderline" when naming it to avoid medical diagnostic language;

[0066] Before naming, it is advisable to have at least two experts in the field (such as pediatrics / orthopedics or ergonomics experts) independently review and reach a consensus.

[0067] Covariates for the subjects were obtained, including maternal education level, parity, gestational age at delivery, child sex, child's birth weight, weaning age, age at spinal follow-up, and weekly duration of vigorous physical activity at the time of spinal follow-up. The covariates were imputed using a chain equation multiple imputation algorithm to generate a preset number of imputed datasets. The statistical results of all imputed datasets were merged using Rubin's rule to calculate the overall parameter estimate, standard error, and preset confidence interval. A generalized linear regression model was then used to analyze the association between umbilical cord blood PFAS exposure and children's spinal morphology based on the covariates.

[0068] Input data and variable definitions:

[0069] Independent variable (X): single PFAS concentration (umbilical cord blood measurement value). If the sample contains multiple PFAS, then each PFAS should be analyzed individually.

[0070] Outcome variable (Y): Pediatric spinal morphology index, which is a standardized continuous index (z-score) or vector (terminal regression). It is recommended to use z-score as the outcome to be compatible with LPA input.

[0071] Set of covariates (Z): The covariates explicitly listed and used in the examples include: maternal education level, parity, gestational week at delivery, child sex, child birth weight, weaning age, age at spinal follow-up (accurate to the year), and weekly duration of vigorous physical activity at spinal follow-up.

[0072] Covariate collection and preprocessing: Coding rules: Define a coding table for categorical covariates (e.g., mother's education level); Preparation before missing value handling: Perform preliminary logical checks and outlier labeling on all variables used for imputation to ensure the stability of the imputation model;

[0073] GLM Model Details and Implementation Recommendations: To improve the skewed distribution of PFAS, a logarithmic transformation of X is preferred; Model Selection: Use linear regression (ordinary least squares) for continuous outcomes, and employ robust standard errors (Huber-White) when heteroscedasticity exists; use Poisson or negative binomial regression for count outcomes; Interaction and Stratification: Perform stratified analysis based on children's gender. Implementation methods include two types: Implementation A (Stratification): Split the sample into two groups by gender and perform repeated imputation and regression analysis; Implementation B (Interaction Term): Add an interaction term X×sex to a single model and test the interaction coefficient (if significant, perform stratified interpretation). Model Diagnosis: Check residual normality, heteroscedasticity, and high leverage points for linear models; check goodness of fit and influence points for logistic models.

[0074] A generalized linear regression model was used to analyze the association between umbilical cord blood PFAS exposure and spinal morphology in children. Stratified analysis was performed based on child sex. Based on previous literature, the corrected covariates in the model included: maternal education level, parity, gestational age at delivery, child sex, birth weight, weaning age, age at spinal follow-up (accurate to the year), and weekly duration of vigorous physical activity at spinal follow-up. The covariates were imputed using a chain equation multiple imputation algorithm, generating 20 imputed datasets. The statistical results of all imputed datasets were combined using Rubin's rule, and the overall parameter estimates, standard errors, and 95% confidence intervals were calculated. To assess the robustness of the results, a sensitivity analysis was conducted to examine the impact of age-specific BMI as a potential confounding variable on the association between PFAS exposure and spinal morphology in children. Based on the above model, age-specific BMI was further controlled for, and a multiple linear regression model was constructed to compare the differences in the effects of PFAS exposure on spinal morphology.

[0075] The concentrations of a predetermined number of PFAS types in the test sample are preprocessed and multiple imputed. The level of each PFAS is quantified by quantile, and a weighted quantile and regression model is constructed. The mixed effects of positive and negative PFAS are estimated in the weighted quantile and regression model. The weight of each PFAS is obtained by repeated sampling or cross-validation. The mixed exposure index is calculated using the weights and is included as an independent variable in a linear regression or generalized linear regression model to evaluate the overall effect of mixed exposure on spinal morphology outcomes. The overall effect estimate of mixed exposure, the relative weight ranking of each component, and the significant outcome indicators under the positive or negative model are output.

[0076] The combined exposure effects of eight PFAS mixtures were calculated using weighted quantiles and regression. WQS, a weighted quartile summation method combined with linear regression, not only estimates the combined positive and negative effects of the mixtures but also calculates the importance weights of each exposure component based on multiple imputation. All statistical analyses were performed in R 4.4.2. Multiple imputation used the "mice" package, data cleaning used the "tidyverse" framework, latent profile analysis used the "tidyLPA" and "mclust" packages, linear regression analysis used the "tidymodels" framework, and combined exposure analysis used the "gWQS" and "miWQS" packages. Plotting was done using the "ggplot2" package, and table creation used the "gt" and "gtsummary" packages. Statistical tests were two-tailed, with a significance level of 0.05.

[0077] In the model investigating the effect of umbilical cord blood PFAS exposure on fecal bile acid metabolism in children, the corrected covariates included: maternal education level, parity, gestational age at delivery, child sex, birth weight, and weaning age. In the model investigating the effect of fecal bile acid metabolism on spinal morphology in children, the corrected covariates were the aforementioned covariates. The mediating effect of fecal bile acid metabolism on the relationship between umbilical cord blood PFAS exposure and spinal morphology in children was tested and evaluated using structural equation modeling. The mediating effect was then categorized based on the test results.

[0078] The relationship between variables in the mediation effect model can be described as follows:

[0079] ;

[0080] ;

[0081] ;

[0082] In the formula, As the independent variable; The dependent variable; as independent variable For dependent variable The overall effect; To reflect the independent variable For mediating variables Its function; In order to control the independent variable In the case of mediator variables For dependent variable The impact; To control for mediating variables In the case of independent variable For dependent variable The direct effect of residuals; express;

[0083] In the mediation effect model, the mediation effect is equivalent to the indirect effect, and its magnitude is... and product The relationship between the coefficients can be described as follows:

[0084] .

[0085] The mediation effect test in this study consisted of the following three steps:

[0086] 1) Test the direct effect (coefficient c), i.e. the impact of umbilical cord blood PFAS exposure on spinal morphology. This test was completed in the first part of the study. Significant associations were considered mediating effects, and other associations were considered masking effects, all of which were included in the subsequent analysis.

[0087] 2) The coefficients a and b were tested sequentially, namely the effects of umbilical cord blood PFAS exposure on fecal bile acid metabolism and the effects of fecal bile acid exposure on spinal morphology. This test was performed in the GLM model studied in this section.

[0088] 3) If there is at least one insignificant association between coefficients a and b, use the Bootstrap test to test ab, and then test the coefficients. And confirm ab and Whether the signs are the same. This test is performed using structural equation modeling. Based on the test results and the basic premise in step 1, the mediation effect is classified into no-mediation effect, complete mediation effect, partial mediation effect, and masking effect.

[0089] like Figure 2 As shown, the three mediation effects described above describe different mechanisms by which the independent variable X influences the dependent variable Y. Among them, the complete mediation effect means that the effect of X on Y is entirely transmitted through the mediating variable M, which can be understood as "X must affect Y through M". In this case, the direct effect of X on Y disappears after controlling for M. The partial mediation effect indicates that X both directly affects Y and indirectly affects Y through M, which can be understood as "X has both direct and indirect influence paths on Y". The masking effect means that the introduction of M enhances or reverses the direct effect of X on Y, that is, "M masks the true effect of X", and may even make the originally insignificant XY relationship significant.

[0090] The mediation effect was tested using the following steps:

[0091] Testing direct effects The effect of umbilical cord blood PFAS exposure on spinal morphology;

[0092] Test coefficient and The effects of umbilical cord blood PFAS exposure on fecal bile acid metabolism and fecal bile acid exposure on spinal morphology were investigated using the GLM model.

[0093] For coefficients and There is at least one non-significant association; use the Bootstrap test. Then test the coefficients. and confirm and The same sign state is achieved through structural equation modeling;

[0094] Based on the test results, the mediation effect is classified into four categories: no mediation effect, complete mediation effect, partial mediation effect, and masking effect. The complete mediation effect is... right The effect is entirely determined by the mediating variable Transmission, control at this time back right The direct effect disappears; the partial mediating effect indicates It directly affects And through Indirect impact The masking effect is The introduction of enhancement or reversal right The direct effects.

[0095] Based on the assessment results of the child's spine, the child's daily life trajectory data is obtained, which includes the child's movement altitude data sequence. An altitude correction term is obtained from the movement altitude data sequence using an altitude correction term formula, and then removed from the child's spine assessment results to obtain an altitude-corrected child spine dataset. The altitude correction term formula is as follows:

[0096] ;

[0097] In the formula, For the first The target variable for each child after altitude correction; For the first The original target measurements for each child; For adjustment coefficients; For the first The average altitude of the children in the exercise altitude data sequence; This is the standard altitude.

[0098] To adjust the coefficients, by Obtain, in the formula For the first Other covariate matrices for each child, such as sex and birth weight; The corresponding coefficients were obtained through historical experience; This is the coefficient for the altitude deviation term; To obtain the value that minimizes the sum of squared residuals value; The standard altitude is set by region when there is significant heterogeneity in the region or population.

[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A method for assessing and detecting spinal morphology in children based on multi-source data fusion, characterized in that, Includes the following steps: S1: Obtain the observation indicators of the children to be tested and perform preprocessing; S2: Construct a potential profile analysis model based on standardized values, obtain the fitting index of each model, and perform weighted discrimination on the fitting index to determine the optimal number of potential categories; S3: Perform ANOVA on the standardized values ​​of each observation indicator for each potential category to test for differences, and summarize and name the phenotypic characteristics of each potential category based on the direction and magnitude of the differences between each indicator and different potential categories. S4: The generalized linear regression model was used to assess the single pollutant exposure effect of PFAS, the weighted quantile and regression model was used to assess the combined exposure effect of PFAS, and the GLM model was used to assess the effects of different types of umbilical cord blood PFAS exposure on fecal bile acid metabolism in children and the assessment results of fecal bile acid metabolism in children on spinal morphology. S5: Correct the spine based on the assessment results of the child's spine and obtain a dataset of the corrected child's spine.

2. The method for assessing and detecting pediatric spinal morphology based on multi-source data fusion as described in claim 1, characterized in that, The process of acquiring and preprocessing the observation indicators of the children to be tested includes: collecting multiple morphological observation indicators of the spine and pelvis of the children to be tested, preprocessing and standardizing each observation indicator, and obtaining the standardized value of each observation indicator.

3. A method for assessing and detecting pediatric spinal morphology based on multi-source data fusion as described in claim 2, characterized in that, The process of constructing a potential profile analysis model based on standardized values, obtaining a fitting index for each model, and weighting the fitting index to determine the optimal number of potential categories includes: weighting the fitting index based on hierarchical clustering analysis, and determining the optimal number of potential categories based on the weighted comprehensive evaluation value, wherein the fitting index is AIC, BIC, AWE, CLC, and KIC.

4. A method for assessing and detecting pediatric spinal morphology based on multi-source data fusion as described in claim 1, characterized in that, The process of performing ANOVA on the standardized values ​​of each observation indicator for each potential category to test for differences, and summarizing and naming the phenotypic characteristics of each potential category based on the direction and effect size of the differences between each indicator and different potential categories, further includes: using ANOVA to detect the difference between the child's spinal morphology and each of the optimal potential categories; summarizing and naming the phenotypic characteristics of each potential category based on the direction and effect size of the differences between each indicator and different potential categories; and obtaining the prior probability and potential category probability of the test subject. The prior probability reflects the relative proportion of each potential category in the total population, and the sum of the prior probabilities of all potential categories is 1. The potential category probability reflects the probability that an individual belongs to each potential category under the conditions of the observed data, and the sum of the probabilities of all potential categories for each individual is 1.

5. A method for assessing and detecting pediatric spinal morphology based on multi-source data fusion as described in claim 1, characterized in that, The method for assessing the single-pollutant exposure effect of PFAS using a generalized linear regression model includes: obtaining covariates for the subjects, including maternal education level, parity, gestational age at delivery, child sex, child's birth weight, weaning age, age at spinal follow-up, and weekly duration of vigorous physical activity at spinal follow-up; imputing the covariates using a chain equation multiple imputation algorithm to generate a preset number of imputed datasets; merging the statistical results of all imputed datasets using Rubin's rule; calculating the overall parameter estimate, standard error, and preset confidence interval; and analyzing the association between umbilical cord blood PFAS exposure and children's spinal morphology using a generalized linear regression model based on the covariates.

6. A method for assessing and detecting pediatric spinal morphology based on multi-source data fusion as described in claim 1, characterized in that, The method of assessing the combined exposure effect of PFAS using weighted quantiles and regression includes: preprocessing and multiple imputation of the concentrations of a predetermined number of PFAS types in the sample to be tested; quantifying the level of each PFAS by quantiles; constructing a weighted quantile and regression model; estimating the mixed effects of positive and negative types in the weighted quantile and regression model; obtaining the weight of each PFAS through repeated sampling or cross-validation; calculating the mixed exposure index using the weights; and incorporating the mixed exposure index as an independent variable into a linear regression or generalized linear regression model to assess the overall effect of mixed exposure on spinal morphological outcomes; and outputting the overall effect estimate of mixed exposure, the relative weight ranking of each component, and the significant outcome indicators under the positive or negative models, respectively.

7. A method for assessing and detecting pediatric spinal morphology based on multi-source data fusion as described in claim 1, characterized in that, The method of using a GLM model to evaluate the effects of different types of umbilical cord blood PFAS exposure on fecal bile acid metabolism in children and the evaluation results of the effects of fecal bile acid metabolism on spinal morphology in children includes: in the model of the effect of umbilical cord blood PFAS exposure on fecal bile acid metabolism in children, the corrected covariates include: maternal education level, parity, gestational age at delivery, child sex, child's birth weight, and weaning age; in the model of the effect of fecal bile acid metabolism on spinal morphology in children, the corrected covariates are the aforementioned covariates; the method of using structural equation modeling to test and evaluate the mediating role of fecal bile acid metabolism between umbilical cord blood PFAS exposure and spinal morphology in children; and the method of classifying the mediating effect based on the test results.

8. A method for assessing and detecting pediatric spinal morphology based on multi-source data fusion as described in claim 7, characterized in that, The structural equation model-based examination and evaluation of the mediating role of fecal bile acid metabolism in the relationship between umbilical cord blood PFAS exposure and spinal morphology in children includes: the relationship between variables in the mediation effect model is described as follows: ; ; ; In the formula, As the independent variable; The dependent variable; as independent variable For dependent variable The overall effect; To reflect the independent variable For mediating variables Its function; In order to control the independent variable In the case of mediator variables For dependent variable The impact; To control for mediating variables In the case of independent variable For dependent variable The direct effect of residuals; express; In the mediation effect model, the mediation effect is equivalent to the indirect effect, and its magnitude is... and product The relationship between the coefficients can be described as follows: 。 9. A method for assessing and detecting pediatric spinal morphology based on multi-source data fusion as described in claim 7, characterized in that, The process of classifying mediation effects based on test results includes: conducting mediation effect tests through the following steps: Testing direct effects The effect of umbilical cord blood PFAS exposure on spinal morphology; Test coefficient and The effects of umbilical cord blood PFAS exposure on fecal bile acid metabolism and fecal bile acid exposure on spinal morphology were investigated using the GLM model. For coefficients and There is at least one non-significant association; use the Bootstrap test. Then test the coefficients. and confirm and The same sign state is achieved through structural equation modeling; Based on the test results, the mediation effect is classified into four categories: no mediation effect, complete mediation effect, partial mediation effect, and masking effect. The complete mediation effect is... right The effect is entirely determined by the mediating variable Transmission, control at this time back right The direct effect disappears; the partial mediating effect indicates It directly affects And through Indirect impact The masking effect is The introduction of enhancement or reversal right The direct effects.

10. A method for assessing and detecting pediatric spinal morphology based on multi-source data fusion as described in claim 1, characterized in that, The step of correcting the child's spine based on the assessment results to obtain a corrected child spine dataset includes: acquiring the child's daily life trajectory data based on the assessment results, the daily life trajectory data including the child's movement altitude data sequence; obtaining an altitude correction term based on the movement altitude data sequence using an altitude correction term formula, and removing it from the child's spine assessment results to obtain an altitude-corrected child spine dataset, wherein the altitude correction term formula is: ; In the formula, For the first The target variable for each child after altitude correction; For the first The original target measurements for each child; For adjustment coefficients; For the first The average altitude of the children in the exercise altitude data sequence; This is the standard altitude.

Citation Information

Patent Citations

  • Spine monitoring method and device for teenagers and children, electronic equipment and storage medium

    CN115153514A

  • Large-scale mixed exposure data analysis method based on machine learning

    CN116738172A

  • Intelligent spine diagnosis and evaluation method and device based on GLM

    CN116824265A