ASD phenotype recognition method based on multi-gene and multi-environment risk scores
By employing a multi-gene and multi-environment risk scoring method, combined with traditional scoring and AI models, data preprocessing and feature screening were performed to establish a fusion-enhanced model. This approach solved the problem of limited predictive ability in ASD phenotypic identification, enabling accurate identification of ASD developmental trajectory, phenotypic severity, and subtype classification, thus providing support for personalized intervention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-03-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the identification of ASD phenotypes, existing technologies have limited predictive ability of multi-environment risk models for continuous phenotypic indicators, machine learning phenotyping methods have not been fully applied in clinical practice, multi-gene risk scoring does not fully consider environmental factors and gender interaction, gene-environment interaction models need to be strengthened in explaining ASD phenotypic heterogeneity, and there is a lack of personalized intervention strategies.
We employ a multi-gene and multi-environment risk scoring method, using the PRS-CSx Bayesian algorithm to calculate the ASD multi-gene risk score. We combine traditional scoring models and pure AI models for data preprocessing and feature selection, and utilize extreme gradient boosting trees and random survival forest algorithms for regression and survival prediction. We establish a fusion enhancement model for comprehensive loss optimization and output multi-dimensional risk assessment results.
It significantly improves the predictive accuracy of ASD developmental trajectory, phenotypic severity, and subtype classification, provides multi-dimensional risk assessment, offers direct decision support for clinicians to develop personalized intervention plans, and improves the robustness and accuracy of the model.
Smart Images

Figure CN121709232A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of ASD phenotype identification methods, and more specifically, to an ASD phenotype identification method based on multi-gene and multi-environment risk scores. Background Technology
[0002] Autism spectrum disorder (ASD) is a common and complex neurodevelopmental disorder with a prevalence of approximately 1 in 59 children in the United States, and a similar prevalence globally. In recent years, with the broadening of diagnostic criteria, the number of ASD diagnoses has risen rapidly, and patients exhibit greater heterogeneity in phenotype and genetics. ASD has a broad symptom spectrum, with significant differences in the severity and manifestation of core symptoms among individuals, and is often accompanied by unique features such as cognitive, language, and comorbid mental disorders. my country attaches great importance to the prevention and treatment of ASD, and national and local governments have successively introduced policies to promote early screening, diagnosis, rehabilitation treatment, and the construction of social support systems. With the widespread implementation of ASD screening for newborns and preschool children, epidemiological studies have found a significantly higher incidence of ASD in populations exposed to high-risk environmental factors, suggesting that cumulative environmental risks play an important role in the development and progression of ASD.
[0003] Application of the Multiple Environmental Risk Model (PERS) in ASD Risk Prediction: Environmental factors are one of the important influencing factors of ASD occurrence. In recent years, researchers have attempted to construct comprehensive environmental risk prediction models to improve the accuracy of risk assessment. A multicenter study collected 57 maternal risk factors from preconception, perinatal, and postpartum periods and constructed five machine learning models in cohorts across 10 cities to predict ASD incidence. The results showed that the Extreme Gradient Boosting Decision Tree (XGBoost) model performed best, achieving a prediction accuracy of 66.2% in an external validation cohort across 3 cities. Based on the predicted probabilities of the optimal model, individuals were divided into low, medium, and high-risk groups. Further analysis revealed that children in the high-risk group had a significantly higher risk of developing ASD than those in the low-risk group. Shapley value interpretation showed that unstable pregnancy mood and lack of multivitamin supplementation were the most contributing risk factors. However, most current PERS models are based on binary environmental exposure variables, and their predictive ability for some continuous phenotypic indicators (such as the severity of social impairment and cognitive ability scores) is limited.
[0004] Exploring Machine Learning in ASD Phenotyping: Due to the significant differences within the ASD population, researchers have attempted to use data-driven methods to subtype ASD. A study published in Nature Genetics in 2025 used a person-centered computational model, based on more than 230 phenotypic characteristics of over 5,000 children with ASD, to cluster individuals with ASD into four distinct clinical and biological subtypes: "Social and Behavioral Challenges," "Mixed ASD with Developmental Delays," "Moderately Challenges," and "Pervasively Affected." Each subtype showed significant differences in developmental trajectory, physical illnesses, and behavioral and psychological comorbidities, and was associated with specific patterns of genetic variation. For example, the "Pervasively Affected" subtype carried the highest proportion of de novo detrimental mutations and exhibited comprehensive developmental delays and psychological comorbidities; while the "Social and Behavioral Challenges" subtype, despite having prominent core autism symptoms, had developmental milestones close to normal levels and often had comorbidities such as ADHD and anxiety.
[0005] The impact of polygenic risk scores (PRS) on ASD phenotypic heterogeneity: Large-scale genome-wide analysis revealed that the cumulative effect of common genetic variations (PRS) is associated with some phenotypic traits of ASD. Thomas et al. analyzed data from 6064 ASD patients in the SFARI database and found that PRS for education level was significantly negatively correlated with the severity of ASD symptoms, while PRS for attention deficit hyperactivity disorder (ADHD) and major depressive disorder was positively correlated with most autism phenotypes. For example, higher PRS for education level often corresponds to lower scores on multiple behavioral scales in ASD children (milder symptoms), while higher PRS for ADHD and depression corresponds to higher scores on scales such as social impairment and stereotyped behaviors. The study also revealed gender-differenced genetic effects: in female ASD patients, autism PRS was negatively correlated with symptoms such as stereotyped behaviors (suggesting a possible genetic protective effect in females), while no significant correlation was observed in male patients.
[0006] The Role of Gene-Environment Interactions in ASD Phenotypes: Individual genetic or environmental factors alone cannot fully explain the heterogeneity of ASD, and their interaction may be a key driving factor. A study based on the Simons Simple Autism Cohort (SSC) found a significant interaction effect between genetic copy number variation (CNV) and prenatal environmental exposure. Individuals carrying ASD-related CNVs whose mothers had severe infections during pregnancy had significantly higher scores for social communication impairment and stereotyped behaviors than those with only a single risk factor. In other words, genetic susceptibility combined with environmental stimuli such as prenatal infections leads to more severe core ASD symptoms. However, for non-core phenotypes such as cognitive function and adaptive abilities, the above gene-environment interaction effect is not significant.
[0007] Regarding multi-environment risk models, most current PERS models are based on binary environmental exposure variables, and their predictive ability for some continuous phenotypic indicators (such as the severity of social impairment and cognitive ability scores) is limited. This, to some extent, restricts the formulation of individualized intervention strategies for different phenotypic characteristics. In terms of machine learning phenotyping, although existing studies have used data-driven methods to classify ASD subtypes, further exploration is needed on how to more accurately identify and classify different subtypes, and how to apply these subtype classifications to clinical practice. Regarding multi-gene risk scoring, while PRS helps quantify the influence of genetics on phenotype, further in-depth research is needed on the interaction between PRS and environmental factors and gender, and these interactions should be incorporated into models to improve the explanatory power of ASD phenotypic heterogeneity. Regarding gene-environment interaction models, these models can explain some phenotypic variations, but future research needs to strengthen the identification and quantification of key interactions to better understand the causal mechanisms of ASD phenotypic heterogeneity. Meanwhile, in the practical application of community screening, how to comprehensively consider the model prediction performance and intervention resource constraints, optimize the early warning threshold and net benefit, and determine the screening intervention strategy under different risk levels to obtain the optimal cost-effectiveness ratio is also an urgent problem to be solved. Summary of the Invention
[0008] The technical problem to be solved by this invention is how to achieve accurate identification and prediction of high-risk ASD phenotypes. In order to overcome the defects of the above-mentioned existing technologies (or related technologies), this invention provides an ASD phenotype identification method based on multi-gene and multi-environment risk scores. This invention provides a method for ASD phenotype identification based on multi-gene and multi-environment risk scores, comprising the following steps: Step S1: Collect multi-dimensional environmental data, multi-dimensional indicator data, and developmental milestone data for each individual patient in the ASD patient cluster. Preprocess the multi-dimensional environmental data to obtain preprocessed data. Construct a multi-dimensional phenotypic matrix based on the multi-dimensional indicator data and the milestone data. Step S2: The PRS-CSx Bayesian algorithm is used to calculate the ASD polygenic risk score for each patient based on the preprocessed data. The linear risk score is obtained by screening and scoring each ASD polygenic risk score according to a traditional scoring model based on multiple pre-set environmental risk factors. Step S3: Using a pure AI model, perform continuous phenotype regression prediction, milestone survival prediction, and four high-risk phenotype classification prediction on the multidimensional phenotype matrix to obtain continuous phenotype prediction features, survival prediction features, and classification prediction features. Step S4: Input the linear risk score, the continuous phenotype prediction feature, the survival prediction feature, and the classification prediction feature into the fusion enhancement model to perform target fusion optimization for comprehensive loss, and obtain the continuous phenotype estimate, the risk of developmental milestone events, and the probability distribution of the four subtypes for each patient as the fusion prediction result.
[0009] The ASD phenotype identification method based on multi-gene and multi-environment risk scoring of this invention has the following advantages compared with the prior art: This invention acquires multi-dimensional environmental data and multi-dimensional indicator data, and constructs a multi-dimensional phenotypic matrix in step S1; calculates linear risk scores in step S2; predicts continuous phenotypic prediction features, survival prediction features, and classification prediction features in step S3; and generates fusion prediction results in step S4. By integrating environmental data, genetic data, and complex phenotypic data, and employing a three-stage modeling strategy of traditional scoring models + pure AI models + fusion enhancement models, this invention solves the problem of limited predictive ability of single models and significantly improves the accuracy of predicting ASD developmental trajectory, phenotypic severity, and subtype classification. It also addresses the common issue of interval censoring in milestone data by using survival analysis for prediction, resulting in data formats more closely aligned with clinical practice and more reliable prediction results. Finally, the output fusion prediction results are no longer a single diagnostic label but provide a multi-dimensional risk assessment, offering direct and rich decision support for clinicians to develop personalized early intervention plans.
[0010] In one possible implementation, in step S1, environmental exposure variables, genetic information, phenotypic measurements, and key developmental times of multiple ASD children aged 4-18 years are obtained from an autism professional database as the multidimensional environmental data.
[0011] Compared with existing technologies, the above technical solution can clearly identify that the multi-dimensional environmental data comes from a high-quality, large-scale autism professional database, ensuring the authority, standardization and consistency of the data used in the prediction stage, and laying a solid foundation for the reliability and generalization ability of the model to finally generate fusion prediction results.
[0012] In one possible implementation, in step S1, the multi-maintenance environment data is sequentially subjected to missing value imputation, outlier correction, and variable encoding unification to obtain the preprocessed data.
[0013] Compared with existing technologies, the above-mentioned technical solution can significantly improve data quality through a systematic preprocessing process, reduce the interference of noise and bias on subsequent model analysis, and is a key preliminary step to ensure the robustness and accuracy of the model.
[0014] In one possible implementation, in step S1, the adaptive behavior scale VABS-II, verbal intelligence VIQ, nonverbal intelligence NVQ, ADI-R social domain score, ADOS social influence score, SRS scale, ADI-R verbal communication score, ADOS communication score, ADI-R repetitive behavior score, ADOS stereotyped behavior score, and RBS-R scale are collected from multiple ASD children aged 4-18 as the multidimensional indicator data.
[0015] Compared with existing technologies, the above-mentioned technical solution uses internationally recognized, comprehensive and diverse gold standard assessment scales as input features, ensuring that the constructed multidimensional phenotypic matrix can comprehensively and accurately characterize the core symptoms and functional levels of ASD, thereby making the model's prediction and classification results highly clinically effective and significant.
[0016] In one possible implementation, step S2, after calculating the ASD polygenic risk score for each individual patient, further includes: The calculated ASD polygenic risk scores are standardized using Z-scores to ensure that each ASD polygenic risk score is on the same dimension as each environmental risk factor.
[0017] Compared with existing technologies, the above technical solution can standardize the ASD polygenic risk score and environmental risk factors through Z-score standardization, thereby eliminating the improper impact of the difference in the scale between variables on the model weight allocation, ensuring that data from different modalities can participate in the model calculation fairly and effectively, and improving the convergence speed and performance of the model.
[0018] In one possible implementation, in step S2, the scores of each of the ASD polygenic risk scores are accumulated to obtain a first risk score; based on the effect size of each environmental risk factor on the ASD polygenic risk score, the ASD polygenic risk scores that meet the effect size index are selected and accumulated to obtain a second risk score; based on the effect size and evidence level of each environmental risk factor on the ASD polygenic risk score, the ASD polygenic risk scores that meet the effect size index and evidence level index are selected and double-weightedly accumulated to obtain a third risk score; and based on the correlation between the first risk score, the second risk score, the third risk score and the core ASD phenotypic index, the linear risk score is obtained.
[0019] Compared with existing technologies, the above-mentioned technical solution provides a variety of scientifically validated linear scoring construction methods. By screening based on the correlation with the core phenotype, it ensures that the selected linear risk scores are the most representative and predictive comprehensive indicators of environmental risk, which provides a stable and effective baseline feature for subsequent fusion enhancement models.
[0020] In one possible implementation, in step S3, the pure AI model uses extreme gradient boosting trees to perform regression prediction on the cognitive ability quantitative phenotype and social ability quantitative phenotype in the multidimensional phenotype matrix to obtain the continuous phenotype prediction features.
[0021] Compared with existing technologies, the above-mentioned technical solution can leverage the advantages of extreme gradient boosting trees in handling complex nonlinear relationships, and can more accurately predict continuous phenotypes such as cognitive and social abilities, capturing complex patterns that traditional linear models may overlook, thereby providing a more refined assessment of individual functional levels.
[0022] In one possible implementation, in step S3, the pure AI model uses a random survival forest algorithm to predict the risk function of developmental events as the survival prediction feature for the milestone data with interval censoring in the multidimensional phenotypic matrix.
[0023] Compared with existing technologies, the above-mentioned technical solution can use the random survival forest algorithm to process interval censored milestone data, avoiding the inaccuracies of traditional methods in processing such data. It can more accurately predict the time risk of specific developmental events, which is crucial for assessing the degree of individual developmental delay.
[0024] In one possible implementation, in step S3, the pure AI model establishes a multi-classification model using XGBoost or LightGBM to analyze the discrete subtype labels in the multidimensional phenotypic matrix and obtain the predicted probability of each patient individual belonging to the social behavior disorder type, mixed developmental delay type, moderate challenge type, or widely restricted type as the classification prediction feature.
[0025] Compared with existing technologies, the above-mentioned technical solution can adopt an advanced gradient boosting tree multi-classification model and output a probability distribution. It can not only achieve accurate classification of the four ASD subtypes, but also provide the probability of belonging to each subtype. This reflects the heterogeneous nature of ASD better than hard classification, helps to identify mixed features or uncertain factors, and provides a more detailed basis for stratified intervention.
[0026] In one possible implementation, in step S4, the linear risk score is used as an additional feature, and the continuous phenotype prediction feature, the survival prediction feature, and the classification prediction feature are used as original input features along with the additional features. These are then fed into a secondary learner, the LightGBM second-order model, for training. The goal is to simultaneously optimize the comprehensive loss of the original input features. The output for each patient is the estimated value of the continuous phenotype, the risk of the developmental milestone event, and the probability distribution of the four subtypes as the fusion prediction result.
[0027] Compared with existing technologies, the above-mentioned technical solution, by using linear risk scores as features and inputting them together with the output of pure AI models into a secondary learner, achieves deep complementarity and enhancement of the advantages of different models. With the goal of optimizing the comprehensive loss, it means that its original design intention is to perform regression, survival prediction and classification tasks simultaneously, thereby outputting a set of internally consistent and mutually synergistic multi-dimensional prediction results, which greatly improves the comprehensive performance and practical value of the final fused prediction results. Attached Figure Description
[0028] Figure 1 This is a flowchart of the steps of the present invention. Detailed Implementation
[0029] First, those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention. Those skilled in the art can make adjustments as needed to adapt to specific application scenarios.
[0030] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0031] See Figure 1This invention discloses a method for identifying ASD phenotypes based on multi-gene and multi-environment risk scores, comprising four steps: S1, S2, S3, and S4. In step S1, multi-dimensional environmental data, multi-dimensional indicator data, and developmental milestone data are collected for each individual ASD patient in a cluster. The multi-dimensional environmental data is preprocessed to obtain preprocessed data, and a multi-dimensional phenotype matrix is constructed based on the multi-dimensional indicator data and the milestone data. In step S2, the PRS-CSx Bayesian algorithm is used to calculate the ASD multi-gene risk score for each individual patient based on the preprocessed data, and a traditional scoring model is applied. Based on a set of pre-defined environmental risk factors, the ASD multigene risk scores are screened and scored to obtain linear risk scores. In step S3, a pure AI model is used to perform continuous phenotype regression prediction, milestone survival prediction, and four-type high-risk phenotype classification prediction on the multidimensional phenotype matrix to obtain continuous phenotype prediction features, survival prediction features, and classification prediction features. In step S4, the linear risk scores, continuous phenotype prediction features, survival prediction features, and classification prediction features are input into the fusion enhancement model for target fusion optimization of comprehensive loss to obtain the continuous phenotype estimate, the risk of developmental milestone events, and the probability distribution of the four subtypes for each patient as the fusion prediction results.
[0032] In this embodiment of the invention, in step S1, the multidimensional environmental data relies on the Simons Foundation's autism database (mainly SSC), from which multidimensional environmental data of more than 2,500 ASD patients aged 4-18 years are obtained. These data cover multiple aspects such as environmental exposure variables, genetic information, phenotypic measurements, and key developmental events, providing a rich and comprehensive data foundation for subsequent research.
[0033] In this embodiment of the invention, based on the latest relevant literature, 18 highly credible environmental risk factors related to ASD were carefully selected. These factors include parental age, maternal infection or medication during pregnancy, perinatal complications, early nutrition and parenting style, etc. These factors have been verified by a large number of studies to be closely related to ASD and are of great significance for subsequent analysis.
[0034] In this embodiment of the invention, the preprocessing method in step S1 includes missing value handling, outlier correction, and variable coding standardization. Missing value handling involves filling in missing values in environmental exposure variables using multiple imputation (MICE), a relatively advanced and effective method for handling missing values. MICE involves simulating the possible values of missing values multiple times to more accurately estimate the true situation of the data. Outlier correction involves logically checking and appropriately correcting outliers in environmental exposure variables. Logical checking ensures the rationality and accuracy of the data and avoids interference from outliers in subsequent analysis results. Variable coding standardization involves unifying the coding format for different types of variables, converting qualitative exposures into binary or ordinal categories. For example, some environmental exposure variables are converted into binary "yes" or "no" variables, or divided according to certain levels. Continuous variables are standardized so that different variables can be analyzed on the same scale, facilitating subsequent modeling and comparison.
[0035] In this embodiment of the invention, when performing ASD polygenic risk scoring in step S2, genome-wide association studies (GWAS) are used to summarize statistics, and the PRS-CSx Bayesian algorithm is used to calculate the ASD polygenic risk score (PRS) for each patient. The PRS-CSx Bayesian algorithm is an algorithm based on Bayesian theory, which can more accurately calculate the ASD polygenic risk score for each patient, providing important genetic information for subsequent analysis. Subsequently, the calculated ASD polygenic risk score is standardized by Z-score, so that the ASD polygenic risk score and environmental risk factors can be analyzed on the same dimension. This Z-score standardization is a commonly used standardization method, which can transform the data into a standard normal distribution with a mean of 0 and a standard deviation of 1, facilitating comparison and analysis between different variables.
[0036] In this embodiment of the invention, in step S1, multiple cognitive ability indicators, social ability indicators, language ability indicators, and repetitive and stereotyped behavior indicators are collected for each individual patient. These indicators include the Adaptive Behavior Scale VABS-II, Verbal Intelligence Quantity (VIQ), Nonverbal Intelligence Quantity (NVIQ), ADI-R Social Domain Score, ADOS Social Influence Score, SRS Scale, ADI-R Verbal Communication Score, ADOS Communication Score, ADI-R Repetitive Behavior Score, ADOS Stereotyped Behavior Score, RBS-R Scale, etc. These indicators can comprehensively characterize the core symptom spectrum of ASD.
[0037] In this embodiment of the invention, in step S1, data on several important milestone events in the development of each patient are compiled, including the age at which milestones such as social / language ability, fine / gross motor skills, independent feeding, and toilet training are reached. These milestone data often have the characteristic of interval censoring, that is, only whether they occurred and the approximate time period are recorded. During preprocessing, these data are converted into an interval format suitable for survival analysis so that more accurate analysis can be performed later.
[0038] In this embodiment of the invention, a constructed multidimensional phenotypic matrix is used to perform latent trait / latent profile analysis (LPA / LCA) to initially extract data-driven phenotypic classifications. This analysis method can uncover potential structures and patterns in the data, providing a foundation for subsequent classification. Combining the typical four-type features in existing literature, semi-supervised learning is used to anchor the clustering results, thereby determining the subtype label for each individual patient, including social behavioral disorder, mixed developmental delay, moderate challenge, and widespread restriction. For samples with unstable labels during repeated sampling or model fitting, they are temporarily designated as "undefined / mixed" and incorporated into the model training with soft labels to reduce the impact of label noise. This approach can improve the stability and accuracy of the model.
[0039] In this embodiment of the invention, in step S2, when calculating the risk score using a traditional scoring model, three PERS scores are calculated based on the 18 selected environmental risk factors according to different weighting schemes: Esum-ASD, Emeta-ASD, and ES-ASD. In the Esum-ASD scoring method, the scores of each ASD polygenic risk score are simply summed to obtain the first risk score. This method is simple and intuitive, and can quickly obtain a preliminary risk score. In the Emeta-ASD scoring method, the effect size of each environmental risk factor on the ASD polygenic risk score (e.g., OR) is used to calculate the risk score. The second risk score is obtained by weighting and accumulating the values (or correlation coefficients) of different factors on the ASD polygenic risk score, making the score more scientific and reasonable. In the ES-ASD scoring method, the evidence level and effect size are comprehensively considered, and the ASD polygenic risk scores that meet the effect size and evidence level indicators are double-weighted and accumulated, which further improves the accuracy and reliability of the score. Then, the correlation between these three PERS scores and the core phenotype indicators of ASD is compared, and the best-performing one is selected as the linear risk score. Through correlation analysis, it can be determined which scoring method can most accurately reflect the core phenotype of ASD, providing a more reliable basis for subsequent early warning.
[0040] In this embodiment of the invention, Youden's J and decision curve analysis (DCA) can also be used to determine the high-risk threshold recommendations corresponding to the linear risk score. Youden's J can measure the diagnostic accuracy of the model, and decision curve analysis can help determine the optimal threshold, thereby providing a clear standard for subsequent screening and early warning.
[0041] In this embodiment of the invention, the analysis and processing of the pure AI model includes continuous phenotypic regression prediction, milestone survival prediction, and four-type high-risk phenotypic classification prediction. Specifically, in continuous phenotypic regression prediction, Extreme Gradient Boosting Tree (XGBoost) is used to regress quantitative phenotypic traits such as cognitive and social abilities. XGBoost is a powerful machine learning algorithm characterized by high efficiency and accuracy, and it can handle nonlinear relationships well. Simultaneously, SHAP values are used to analyze the importance and interaction of each feature (including environmental factors and PRS) in the model. SHAP values can intuitively show the contribution of each feature to the prediction results, helping us to better understand the model's decision-making process. In milestone survival prediction, for interval-censored milestone event data, the Random Survival Forest (RSF) model is used to predict the risk function of developmental events. Random Survival Forest is a machine learning algorithm suitable for survival analysis, capable of handling interval-censored data. The time-dependent C-index and Prediction Error Curve (PEC) are calculated to evaluate the model's performance. Yes, and examine the model's calibration over time. The time-dependent C-index can measure the model's predictive accuracy at different time points, and the prediction error curve can intuitively show how the model's error changes over time. In the classification prediction of four high-risk phenotypes, for discrete subtype labels, a multi-class model of XGBoost or LightGBM is established (a one-vs-rest strategy can be used to handle class imbalance), outputting the predicted probability of each patient belonging to the four subtypes. To avoid the model being overconfident or biased, the output probability is calibrated (e.g., Platt scaling or equivalent scaling) to ensure that the probability value matches the actual frequency. In the classification model, focus on indicators such as the area under the ROC curve (AUROC) and the area under the mean precision-recall curve (AUPRC) of macro and micro averages, and calculate precision, recall, F1 index, etc. for selected thresholds to evaluate the model's ability to identify a few high-risk categories. These indicators can comprehensively evaluate the model's classification performance and help us select the most suitable model.
[0042] In this embodiment of the invention, in step S4, the linear risk score is treated as an additional feature and input together with the original input features containing continuous phenotypic prediction features, survival prediction features, and classification prediction features into the secondary learner—the LightGBM second-order model—for training. This fusion method can fully utilize the interpretability of traditional scoring models and the nonlinear learning ability of pure AI models to improve the predictive performance of the fusion-enhanced model. With the goal of simultaneously optimizing the comprehensive loss of continuous phenotypic prediction features, survival prediction features, and classification prediction features, the fusion prediction result is output, which includes the continuous phenotypic estimate, the risk of developmental milestone events, and the probability distribution of the four subtypes for each patient. By optimizing the comprehensive loss, the fusion-enhanced model can achieve better performance on multiple tasks. During model training, five-fold cross-validation is used to evaluate the improvement of the fusion-enhanced model compared to the baseline model and the single AI model. If the improvement of the main performance indicators (such as the C-index or AUROC) exceeds 5%, the fusion-enhanced strategy is considered to have achieved significant benefits. Five-fold cross-validation can more accurately evaluate the performance of the fusion-enhanced model and avoid the problem of overfitting.
[0043] In this embodiment of the invention, the performance of the three-track parallel model can also be compared on an independent validation set. Appropriate metrics are used for evaluation based on different task types. For continuous phenotypic prediction, the percentage reduction in C-index, cumulative Brier score (IBS), and time-dependent AUC are reported, along with calibration curves and net benefit analysis. The C-index measures the model's predictive accuracy, the cumulative Brier score assesses the model's prediction error, the time-dependent AUC demonstrates the model's predictive performance at different time points, the calibration curve checks the model's calibration degree, and the net benefit analysis evaluates the model's benefits in practical applications. For developmental milestone survival prediction, the C-index and IBS of the RSF model are calculated, and a prediction error curve is plotted during the follow-up period to observe the error change over time. These metrics provide a comprehensive evaluation of the model's performance in developmental milestone survival prediction. For the four-type classification task, macro / micro average AUROC and AUP are reported. The RC value is used to plot the calibration curve and calculate the expected calibration error (ECE). Simultaneously, the recall and precision of the model at the Top N high-risk individuals (e.g., Recall@Top5%, PPV@Top1%) are assessed to evaluate its feasibility for screening and early warning. These metrics evaluate the model's performance in the four-type classification task and its practicality in screening and early warning. Regarding the final model selection, the advantages of the three types of models—traditional scoring, pure AI, and fusion enhancement—are comprehensively compared using the above metrics. If the fusion enhancement model outperforms other models in all major metrics (with an improvement exceeding a preset threshold, such as 5%), then the fusion enhancement model is selected as the final solution. For clinical net benefit validation, the DCA decision curve is used to verify the final model's clinical net benefit at different thresholds, ensuring its practical value in public health applications. The decision curve helps determine the optimal threshold, thereby maximizing the model's net benefit in practical applications.
[0044] In this embodiment of the invention, genetic-environment interaction analysis can also be performed, which is divided into traditional statistical analysis and machine learning-enhanced analysis. In traditional statistical analysis, for continuous phenotype regression models, a linear regression or generalized linear model is established, with continuous phenotype as the dependent variable and PERS, PRS, and their product interaction terms as independent variables. Necessary covariates are included, and the regression coefficients and significance of the interaction terms are estimated to determine whether the synergistic effect of genes and environment on phenotype is statistically significant. Through this analytical method, the synergistic effect of gene and environmental risk factors on continuous phenotype can be clarified, providing important statistical basis for a deeper understanding of the pathogenesis of ASD. For subtype label regression models, for four-category subtype labels, a multinomial logistic regression model (or ordered logit model, depending on the relationship between subtypes) is used to calculate the relative risk ratio of PERS, PRS, and the PERS×PRS interaction term for each subtype classification. The ratio method identifies which environmental or genetic factors exhibit enhancing or weakening effects in which subtypes. This analytical approach helps us understand the relationship between different subtypes of ASD and genetic and environmental factors, providing a basis for personalized treatment and intervention. In machine learning augmentation analysis, regarding feature interaction quantification, trained XGBoost, LightGBM, and RSF models are used to extract SHAP values and interaction SHAP values, quantifying the contribution of pairwise feature interactions to the prediction results. Emphasis is placed on the interaction SHAP between PERS and PRS, its value in different tasks (regression, survival, classification), and its distribution differences among individuals of different subtypes. This analytical method provides a deeper understanding of the impact of interactions between genetic and environmental risk factors on fusion prediction results, providing important information for model optimization. Regarding the discovery of key interaction pairs, for survival models, the minimum depth of variables in a random survival forest and the importance of variables based on permutations are calculated to discover potential key gene-environment interaction pairs (e.g., a certain environmental risk factor and a specific PRS). The combination of these factors significantly contributes to the prediction of specific subtypes. This analytical method can identify gene-environment interaction pairs that have a significant impact on the prediction of specific subtypes, providing direction for further research and intervention.
[0045] In this invention, a three-track parallel technical route of "traditional scoring model - pure AI model - fusion enhancement model" was designed. Combining multi-environment risk score (PERS) and multi-gene risk score (PRS) with machine learning technology, it achieves accurate identification and early warning of high-risk ASD phenotypes. At the same time, through the analysis of genetic-environment interactions, the complex mechanism of high-risk phenotype formation is deeply understood. In model validation, if the main performance indicators of the fusion enhancement model improve by more than 5%, the fusion enhancement strategy is considered to have achieved significant benefits. It effectively solves the problems of existing models in continuous phenotype indicator prediction, subtype classification accuracy, factor interaction research, and community application, improves the explanatory power of ASD phenotype heterogeneity and the practicality of the model in community scenarios, and provides a more effective strategy and method for ASD screening and intervention, which helps to obtain the optimal cost-effectiveness ratio.
[0046] In the description of this invention, the references to "one embodiment," "some embodiments," "in this embodiment," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0047] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for ASD phenotypic identification based on multi-gene and multi-environment risk scores, characterized in that, Includes the following steps: Step S1: Collect multi-dimensional environmental data, multi-dimensional indicator data, and developmental milestone data for each individual patient in the ASD patient cluster. Preprocess the multi-dimensional environmental data to obtain preprocessed data. Construct a multi-dimensional phenotypic matrix based on the multi-dimensional indicator data and the milestone data. Step S2: The PRS-CSx Bayesian algorithm is used to calculate the ASD polygenic risk score for each patient based on the preprocessed data. The linear risk score is obtained by screening and scoring each ASD polygenic risk score according to a traditional scoring model based on multiple pre-set environmental risk factors. Step S3: Using a pure AI model, perform continuous phenotype regression prediction, milestone survival prediction, and four high-risk phenotype classification prediction on the multidimensional phenotype matrix to obtain continuous phenotype prediction features, survival prediction features, and classification prediction features. Step S4: Input the linear risk score, the continuous phenotype prediction feature, the survival prediction feature, and the classification prediction feature into the fusion enhancement model to perform target fusion optimization for comprehensive loss, and obtain the continuous phenotype estimate, the risk of developmental milestone events, and the probability distribution of the four subtypes for each patient as the fusion prediction result.
2. The ASD phenotype recognition method according to claim 1, characterized in that, In step S1, environmental exposure variables, genetic information, phenotypic measurements, and key developmental times of multiple ASD children aged 4-18 years are obtained from a specialized autism database as the multidimensional environmental data.
3. The ASD phenotype recognition method according to claim 1, characterized in that, In step S1, the multi-maintenance environment data is sequentially subjected to missing value imputation, outlier correction, and variable encoding to obtain the preprocessed data.
4. The ASD phenotype recognition method according to claim 2, characterized in that, In step S1, the adaptive behavior scale VABS-II, verbal intelligence VIQ, nonverbal intelligence VIQ, ADI-R social domain score, ADOS social influence score, SRS scale, ADI-R verbal communication score, ADOS communication score, ADI-R repetitive behavior score, ADOS stereotyped behavior score, and RBS-R scale of multiple ASD children aged 4-18 years are collected as the multidimensional indicator data.
5. The ASD phenotype recognition method according to claim 1, characterized in that, In step S2, after calculating the ASD polygenic risk score for each individual patient, the following is also included: The calculated ASD polygenic risk scores are standardized using Z-scores to ensure that each ASD polygenic risk score is on the same dimension as each environmental risk factor.
6. The ASD phenotype recognition method according to claim 1, characterized in that, In step S2, the scores of each ASD polygenic risk score are accumulated to obtain a first risk score. Based on the effect size of each environmental risk factor on the ASD polygenic risk score, the ASD polygenic risk scores that meet the effect size index are selected and accumulated to obtain a second risk score. Based on the effect size and evidence level of each environmental risk factor on the ASD polygenic risk score, the ASD polygenic risk scores that meet the effect size index and evidence level index are selected and double-weightedly accumulated to obtain a third risk score. Based on the correlation between the first risk score, the second risk score, the third risk score and the core ASD phenotypic index, the linear risk score is obtained.
7. The ASD phenotype recognition method according to claim 1, characterized in that, In step S3, the pure AI model uses extreme gradient boosting trees to perform regression prediction on the cognitive ability quantitative phenotype and social ability quantitative phenotype in the multidimensional phenotype matrix to obtain the continuous phenotype prediction features.
8. The ASD phenotype recognition method according to claim 1, characterized in that, In step S3, the pure AI model uses the random survival forest algorithm to predict the risk function of developmental events as the survival prediction feature for the milestone data with interval censoring in the multidimensional phenotypic matrix.
9. The ASD phenotype recognition method according to claim 1, characterized in that, In step S3, the pure AI model establishes a multi-classification model using XGBoost or LightGBM to analyze the discrete subtype labels in the multidimensional phenotypic matrix and obtains the predicted probability of each patient belonging to the social behavior disorder type, mixed developmental delay type, moderate challenge type, or widespread restriction type as the classification prediction feature.
10. The ASD phenotype recognition method according to claim 1, characterized in that, In step S4, the linear risk score is used as an additional feature. The continuous phenotype prediction feature, the survival prediction feature, and the classification prediction feature are used as the original input features, along with the additional features, and are input into the secondary learner LightGBM second-order model for training. The goal is to simultaneously optimize the comprehensive loss of the original input features. The output of the continuous phenotype estimate, the risk of occurrence of the developmental milestone event, and the probability distribution of the four subtypes for each patient are used as the fusion prediction result.