Construction method of offspring individual autism spectrum disorder risk prediction model and risk prediction system
By using prospective cohort data and relative transcriptional order to identify risk markers in the blood of mothers during pregnancy, an individualized ASD risk prediction model was constructed, which solved the problem of low accuracy in pregnancy ASD risk assessment in the prior art, and achieved accurate early warning and individualized detection of ASD during pregnancy.
Patent Information
- Application Number
- CN202510085837.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to individually evaluate the risk of progeny autism spectrum disorder (ASD) through maternal blood during pregnancy, and the existing models have low prediction accuracy and cannot identify real biomarkers.
Through prospective cohort data, based on transcriptional relative order, an individualized ASD risk prediction model was constructed, and the order and machine learning algorithms were evaluated using expression characteristics.
It realizes the accurate assessment of the risk of ASD in the offspring of MIA during pregnancy, improves the accuracy of prediction, can conduct individualized testing, and accurately alert ASD to pregnancy.
Smart Images

Figure CN120015286A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of prenatal diagnosis of individualized risk assessment of autism spectrum disorders in offspring, and in particular to a method for constructing a risk prediction model for individual autism spectrum disorders in offspring and a risk prediction system. Background Art
[0002] Autism spectrum disorder (ASD) is a neurodevelopmental disorder characterized by social communication disorders and repetitive stereotyped behaviors. Its etiology is complex and is affected by the combined effects of genetics, environment and their interactions. It is highly heterogeneous and the core symptoms persist throughout life. ASD is widespread and common in children and adults. The global prevalence rate is 1-2%, which is increasing year by year. It has become a major public health and socioeconomic issue with global consensus.
[0003] At present, the age of diagnosis of ASD is 2 to 12 years old. The effect of intervention decreases significantly with age. Severe cases will become lifelong disabled, seriously endangering the quality of life of individuals and their families. However, the construction of ASD risk assessment models in the past has relied on multi-omics big data of brain tissue from autopsies many years after the onset of the disease or blood samples when the disease course develops to a clinically diagnosable stage. The timeliness of the risk genes obtained is poor, and there are few risk markers in maternal blood during pregnancy for offspring ASD risk assessment. JianjunOu et al. used stepwise logistic regression analysis to detect the effects of 28 candidate risk factors related to pre-pregnancy, pregnancy, perinatal and postpartum on the risk of ASD in offspring. The results showed that influenza-like illness during pregnancy, pregnancy stressors, maternal allergies / autoimmune diseases, cesarean sections and hypoxia were significantly associated with the risk of ASD, and the prediction accuracy of the constructed risk model was 71.1%. Kathryn Hollowood et al. conducted a metabolic profile test during pregnancy on mothers who already had an ASD child, and were able to accurately predict that the risk of ASD in the second child was about 18.7%, while the risk of the general population was about 1.7%. However, the models constructed by the above methods are all designed based on the test population level, and often fail to identify real biomarkers due to batch effects produced by different omics sequencing methods, resulting in low prediction accuracy and inability to perform personalized testing. There is an urgent need to use prospective cohorts to identify new and reliable key risk genes during pregnancy and develop a stable personalized risk prediction model for MIA-induced ASD in offspring.
[0004] A large number of epidemiological studies and animal model experiments have shown that maternal immune activation (MIA) during pregnancy can disrupt the dynamic balance of immunity between the maternal and fetal environments, have a profound and lasting impact on neurodevelopment, and thus increase the risk of ASD after birth. However, due to clinical ethics, prospective studies cannot be conducted, and there are delays in "diagnosis, intervention, and family" in clinical practice, which makes ASD difficult to recover and control, and has a high risk of disability. Therefore, it is urgent to strengthen early screening and early intervention. Summary of the invention
[0005] Based on the above technical problems, the present invention uses prospective cohort data to identify risk markers in the blood of pregnant mothers that are associated with ASD in offspring based on the relative order of transcription, and proposes a more effective model for assessing the risk of individual ASD in offspring. This model can achieve accurate assessment of the risk of ASD in offspring by MIA during pregnancy, and accurately predict the early warning of ASD during pregnancy, which has certain social value.
[0006] In order to achieve the above object, the technical solution provided by the present invention is as follows:
[0007] The first aspect of the present invention provides a method for constructing an offspring individual ASD risk prediction model, comprising the following steps:
[0008] Identify genes ranked in the top 10% by absolute fold change (FC) in blood transcriptome data from a prospective cohort of mothers with ASD and typical development (TD) offspring;
[0009] Compare the expression levels of any two genes among several genes to obtain the expression feature pair rank (FPR) spectrum;
[0010] Based on the expression signature pair order profiles and compared with the maternal blood during pregnancy associated with typical development of the offspring, identifying risk gene pairs with significantly reversed expression signature pair order in the maternal blood during pregnancy associated with the risk of ASD in the offspring;
[0011] A mixed feature selection is performed on the risk gene pairs, and the optimal parameters are obtained using cross-validation. The effect value is calculated according to the optimal parameters, and an individualized risk prediction model is constructed based on the effect value and the results after the mixed feature selection.
[0012] Furthermore, the blood transcriptional data of pregnant mothers whose offspring were ASD and TD were derived from the transcriptomics data provided by GSE148450.
[0013] Furthermore, the method for obtaining the order of expression features includes:
[0014] Select any two genes from the plurality of genes, compare the expression levels of the two genes, and construct an order symbol scoring function to obtain the expression feature pair order spectra after the gene expression profiles of all samples of the blood of pregnant mothers whose offsprings have autism spectrum disorders and whose offsprings have typical development are converted;
[0015] The obtained expression feature pair order spectrum is: if the expression level of one gene is greater than that of another gene, the expression feature pair order is recorded as 1; if the expression level of one gene is less than that of another gene, the expression feature pair order is recorded as -1.
[0016] Furthermore, the risk gene pairs are identified by:
[0017] Fisher's exact test was performed for each gene pair between the blood groups of mothers with ASD offspring and those with TD offspring during pregnancy;
[0018] The P values obtained by Fisher's exact test were corrected for multiple testing, and genes with P < 0.05 after correction were considered risk gene pairs.
[0019] Furthermore, the risk gene pairs whose expression characteristics in the blood of pregnant mothers during pregnancy are significantly reversed and are associated with the risk of ASD in offspring are identified, which specifically includes the following steps:
[0020] The number of samples that met the different expression magnitude relationships in the blood of pregnant mothers whose offspring were ASD and whose offspring were TD were counted respectively;
[0021] Based on the number of samples that met the different expression magnitude relationships, Fisher's exact test and multiple testing correction were used to identify risk gene pairs with significantly reversed order of expression characteristics in maternal blood during pregnancy that were associated with ASD risk in offspring.
[0022] Furthermore, the risk gene pairs screened were AC009229.5-RP5-1158E12.1, AC087491.2-GPT 、C1orf177-LOC286189, C1orf177-RP11-75L1.1, CYB5R2-KLK15, DKFZP434K028-LAMP5, IGFL1-TRIM60, LINC00635-RFTN2, LI NC01010-RP11-69C17.3, LOC100506384-TRIM60, LOC101060277-PRB4, LOC286189-LOC400620, MEF2A-RAB24, OR4C13-TBCA, RFTN2-RP11-104E19.1, RIMBP2-RNU1-131P, RNU6-798P-ZIC5, RNU6-597P-RNU6-788P, MAGI1-TRIM60, LOC100130373-TMEM2 53. KCNH1-TRIM60, GHRHR-SNORD114-17, FARP-AS1-MIR548M, FANCD2OS-RP11-272D20.2, CHST10-ZNF668, AMHR2-LOC286189.
[0023] Furthermore, hybrid feature selection is performed using the minimum absolute value shrinkage and selection operator.
[0024] Furthermore, risk gene pairs with effect values greater than 0 were selected, and the individualized risk prediction model was constructed using the expression weighted effect value. Among them, the risk gene pairs with effect values greater than 0 were AC009229.5-RP5-1158E12.1, AC087491.2-GPT 、 C1orf177-LOC286189, C1orf177-RP11-75L1.1, CYB5R2-KLK15, DKFZP434K028-LAMP5, IGFL1-TRIM60, LINC00635-RFTN2, LINC01010-RP11-69C17.3, LOC10 0506384-TRIM60, LOC101060277-PRB4, LOC286189-LOC400620, MEF2A-RAB24, OR4C13-TBCA, RFTN2-RP11-104E19.1, RIMBP2-RNU1-131P, RNU6-798P-ZIC5.
[0025] In a second aspect, the present invention provides a system for predicting the risk of ASD in offspring individuals during pregnancy and before birth, which includes a data collection unit, a data analysis unit, a risk scoring unit, and a result output unit;
[0026] The data collection unit is used to collect the expression feature pair order (FPR) of the risk gene pairs in the blood samples of each pregnant mother whose offspring are ASD and TD;
[0027] The data analysis unit is used to use the expression feature pairs of risk gene pairs to identify prenatal risk factors associated with offspring ASD risk based on a machine learning algorithm;
[0028] The risk scoring unit is used to construct an individualized risk score for the occurrence of ASD in offspring during pregnancy using the individualized risk prediction model constructed according to the above method, and to perform grade classification according to the risk score result;
[0029] The result output unit is used to generate an assessment report based on the risk scoring result.
[0030] Furthermore, the data collection unit is used to detect the expression levels of risk gene pairs in peripheral blood during pregnancy, and calculate the expression feature order based on the expression levels.
[0031] Furthermore, the data analysis unit is used for:
[0032] Blood samples of pregnant mothers whose offspring were ASD and TD were divided into training sets and validation sets. The least absolute shrinkage and selection operator (Lasso) algorithm was used for hybrid feature selection in the training set, and the optimal parameter λ was obtained in combination with 10-fold cross validation. The effect value was calculated based on the optimal parameter, and an individualized risk prediction model was constructed based on the effect value and the hybrid feature selection results.
[0033] The ROC curve was drawn to verify the constructed individualized risk prediction model in the validation set, and combined with Bootstrap resampling, the average AUC and confidence interval were obtained to evaluate the individualized risk prediction model.
[0034] Furthermore, the risk scoring unit is used to classify pregnant women into three levels: high risk (5.06-8.22), medium risk (-1.82-4.25) and low risk (-5.46--2.35) using a Gaussian mixture model. The larger the risk score of the pregnant woman, the higher the risk of the offspring suffering from ASD.
[0035] Furthermore, the result output unit is used to provide individualized health management suggestions and recommend specific medical intervention measures according to the risk level of the pregnant woman.
[0036] Compared with the prior art, the present invention has the following technical effects:
[0037] Based on the prospective cohort data of blood transcriptomics of pregnant mothers with ASD offspring, the present invention combines the order of gene expression feature pairs in a single sample and machine learning to construct an individualized ASD risk prediction model for offspring of MIA during pregnancy. It not only takes into account the heterogeneity of ASD and the prevalence of micro-effect genes, but also takes into account the interaction between the genome and environmental risk factors during pregnancy, effectively overcoming the difficulties in obtaining materials and ethical constraints. The screening process of this model is rigorous and reliable, and the risk of individual offspring can be evaluated by the order of relative expression feature pairs of risk gene pairs in blood samples during pregnancy.
[0038] The present invention can achieve accurate assessment of the risk of ASD in offspring by MIA during pregnancy, and accurately predict the early warning of ASD during pregnancy, which is beneficial for early warning during pregnancy and early intervention after birth, thereby achieving the purpose of personalized precision medicine, and has certain social value.
[0039] The present invention identifies risk markers associated with ASD in offspring in the blood of pregnant mothers through expression feature pair order (FPR), and proposes a more effective model for assessing the risk of ASD in offspring individuals: FPRscore = (-0.980) × sgn (AC009229.5-RP5-1158E12.1) + 0.131 × sgn (AC087491.2-GPT) + (-0.625) × sgn (C1orf177-LOC286189) + (-0.100) × sgn (C1orf177-RP11-75L1.1) + 0.292 × sgn (CYB5R2-KLK15) + (-1.487) × sgn (DKFZP434K028-LAMP5) + (-0.331) × sgn (IGFL1-TRIM60) + 0.15 0×sgn(LINC00635-RFTN2)+(-0.486)×sgn(LINC01010-RP11-69C17.3)+(-0.242)×sgn(L OC100506384-TRIM60)+(-1.306)×sgn(LOC101060277-PRB4)+0.801×sgn(LOC286189-LO C400620)+(-0.774)×sgn(MEF2A-RAB24)+0.858×sgn(OR4C13-TBCA)+(-0.322)×sgn(RFT N2-RP11-104E19.1)+(-1.249)×sgn(RIMBP2-RNU1-131P)+0.264×sgn(RNU6-798P-ZIC5).
[0040] The present invention constructs the first ASD risk prediction model for offspring individuals based on prenatal diagnosis of the pregnant woman's blood risk gene FPR, and the AUC reaches 0.0.821. Compared with previous risk prediction models, it can perform individualized assessments without being affected by batches. For pregnant women who already have first- or second-degree relatives with ASD, the accuracy of the present invention in prenatal diagnosis of ASD risk in offspring individuals is greatly improved. In addition, the constructed FPRscore risk score can divide pregnant women into three groups: high risk, medium risk and low risk. The higher the FPRscore value, the higher the risk. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 Based on a prospective cohort, risk markers in the blood of pregnant mothers that are associated with ASD in their offspring were successfully identified. (A) The heat map shows the relative expression characteristics of 26 risk gene pairs in the blood of pregnant mothers whose offspring are ASD and TD. (B) The pie chart shows the gene function types of 45 risk genes. (C) The percentage bar chart shows the similarities and differences in the chromosomal location of the risk gene pairs. (D) The violin plot shows the differential expression of risk genes between the two groups. (E) The bubble chart shows the degree of correlation of the expression of risk gene pairs in each group.
[0042] Figure 2 The individualized risk prediction model for MIA-induced ASD in offspring was successfully developed. (A) Lasso analysis of 26 risk gene pairs. (B) Effect values of 26 risk gene pairs. (C) There are differences in the risk assessment values of maternal blood during pregnancy for offspring ASD and TD in the training set (left) and validation set (right). (D) ROC curves of the training set and validation set. (E) AUC distribution of 1000 bootstrap resamplings.
[0043] Figure 3 FPRscore is an effective method for assessing the individualized risk of ASD in offspring induced by MIA during pregnancy. (A) Violin plot shows that FPRscore differs between the two groups. (B) ROC curve (left) and AUC (right) of the FPRscore risk score model after 1000 bootstrap resampling. (C) Heat map shows that the Gaussian mixture model divides the samples of pregnant mothers with ASD in offspring into high, medium, and low risk groups. (D) Volcano plot shows differentially expressed genes in high and low risk groups.
[0044] Figure 4 The ROC curve (left) and AUC after 1000 bootstrap resampling of the FPRscore risk score model in umbilical cord blood of offspring with ASD (right). DETAILED DESCRIPTION
[0045] The technical scheme in the embodiment of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiment of the present invention. The given embodiment is only for illustrating the present invention, but not for limiting the scope of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, but not all of the embodiments.
[0046] Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in the field without making any creative work shall fall within the scope of protection of the present invention.
[0047] Autism spectrum disorder (ASD) is a neurodevelopmental disorder with social communication disorders and repetitive stereotyped behaviors. Its etiology is complex and is affected by the combined effects of genetics, environment and their interactions. It is highly heterogeneous and the core symptoms persist throughout life.
[0048] In the past, the construction of ASD risk assessment models relied on multi-omics big data of brain tissue from autopsies many years after the onset of the disease or blood samples when the disease progressed to a clinically diagnosable stage. The timeliness of the risk genes obtained was poor, and there were few risk assessments of ASD in offspring based on risk markers in maternal blood during pregnancy. Moreover, they were all designed based on the detection population level, and often due to batch effects produced by different omics sequencing methods, it was impossible to identify real biomarkers, resulting in low prediction accuracy of the existing risk models and the inability to conduct individualized testing.
[0049] Based on this, the present invention provides a method for constructing a risk prediction model for autism spectrum disorder in offspring individuals, comprising the following steps:
[0050] Identify genes ranked in the top 10% by absolute fold difference in maternal blood transcriptome data from a prospective cohort of offspring with autism spectrum disorder and typical development;
[0051] Compare the expression levels of any two genes among several genes to obtain the order spectrum of expression feature pairs;
[0052] Based on the order profile of the expression signature pairs, and compared with the maternal blood during pregnancy associated with typical development of the offspring, identify risk gene pairs with significantly reversed order of expression signature pairs in maternal blood during pregnancy that are associated with the risk of autism spectrum disorder in the offspring;
[0053] Hybrid feature selection was performed on the identified risk gene pairs, and cross-validation was used to obtain the optimal parameters. The effect value was calculated based on the optimal parameters, and an individualized risk prediction model was constructed based on the obtained effect value and the results after hybrid feature selection.
[0054] Embodiment 1:
[0055] The embodiment of the present invention provides a method for constructing a model for predicting the risk of ASD in offspring based on prenatal assessment of the expression feature pair order (FPR) of risk gene pairs in the blood of pregnant women, comprising the following steps:
[0056] 1. Identify risk markers associated with ASD in offspring based on maternal blood transcriptional data during pregnancy
[0057] Methods: A prospective cohort data on blood transcriptomics of pregnant mothers with offspring with ASD and typical development (TD, i.e., healthy controls) included in the GEO database (GSE148450; 56 offspring with ASD (m) and 103 offspring with TD (Mm)) were analyzed. The data had been approved by the Human Genetics Ethics Committee of the data source institution, and the inclusion criteria were pregnant women who had first- or second-degree relatives with ASD.
[0058] The RMA function of the ROLigo package (version 1.60.0) was used to remove low-expression genes in the transcriptome data of maternal blood samples during pregnancy, and 33,870 genes were obtained.
[0059] The fold change (FC) of each gene expression in the blood transcriptome data of pregnant mothers whose offspring developed ASD relative to that in the blood transcriptome data of pregnant mothers whose offspring developed TD was calculated. Genes ranked in the top 10% by FC absolute value were selected for pairwise comparison of expression levels, and an order symbol scoring function was constructed to convert the gene expression profile into a feature pair rank (FPR) profile.
[0060] Among them, the order symbol scoring function is constructed according to the following steps:
[0061] Assume that any two genes in the k-th mother sample (k=1,…,m,…,M) are denoted by G i,k and Considering the relationship between the expression levels of two genes (E), an order symbol scoring function is constructed:
[0062]
[0063] That is, in the kth sample, the expression level E of the i-th gene i,k Greater than the expression level E of the jth gene j,k , let the order ranki,j,k be 1, otherwise ranki,j,k is -1.
[0064] Thus, the expression feature pair order (FPR; vector) of the kth sample is obtained, as well as the FPR spectrum (matrix) after the gene expression profiles of all samples of pregnant mothers with risk of autism spectrum disorder in offspring and pregnant mothers with typical development are transformed.
[0065] Then, for each gene pair (G i,k and G j,k ), do the following:
[0066] (1) Statistics of the expression levels of genes in the blood of pregnant mothers whose offspring have ASD i,k ≥E j,k The number of samples a i,j , the calculation formula is as follows:
[0067]
[0068] Where, m: the number of mothers whose offspring are ASD.
[0069] (2) Statistics of the expression levels in the blood of pregnant mothers whose offspring have ASD i,k <E j,k The number of samples b i,j , the calculation formula is as follows:
[0070]
[0071] Where, m: the number of mothers whose offspring are ASD.
[0072] (3) Statistical analysis of the expression levels in the blood of pregnant mothers with TD offspring i,k ≥E j,k The number of samples c i,j , the calculation formula is as follows:
[0073]
[0074] Wherein, m: the number of pregnant mothers whose offspring are ASD; M: the total number of pregnant mothers whose offspring are ASD and TD.
[0075] (4) Count the expression levels in the blood of pregnant mothers whose offspring are TD i,k <E j,k The number of samples d i,j , the calculation formula is as follows:
[0076]
[0077] Wherein, m: the number of pregnant mothers whose offspring are ASD; M: the total number of pregnant mothers whose offspring are ASD and TD.
[0078] Subsequently, Fisher's exact test and multiple testing correction (FDR < 0.05) were used to compare the order of maternal blood expression signatures associated with typical offspring development to identify risk gene pairs with significantly reversed order of maternal blood expression signatures associated with offspring ASD risk. The specific method is as follows:
[0079] Based on the above statistical results and performing Fisher's exact test between the two groups for each gene pair, the calculation formula is as follows:
[0080]
[0081] Multiple testing correction P BH <0.05 was used as the significant threshold, and 26 risk gene pairs were identified ( Figure 1 Interestingly, these gene pairs contained a total of 45 genes, including protein-coding genes and other non-protein-coding genes ( Figure 1 Middle B), most of which are not on the same chromosome ( Figure 1 Two independent sample T-tests were used for differential analysis, and the results showed that most genes (33) were significantly differentially expressed between the two groups ( Figure 1 D), and the pairwise Pearson correlation test showed that most of the risk gene pairs (15) were significantly correlated in at least one group ( Figure 1 Middle E), indicating that the FPR strategy can identify more and more reliable risk markers.
[0082] 2. Construction and validation of a risk prediction model for ASD in offspring based on prenatal assessment of FPR based on risk genes in maternal blood
[0083] Using the 26 risk gene pairs identified above, a risk model for offspring ASD based on maternal blood during pregnancy was constructed. Specifically, it includes:
[0084] The least absolute shrinkage and selection operator (LASSO) was used for hybrid feature selection, and the optimal parameter λ = 0.01171899 ( Figure 2 A), LASSO analysis can select features and shrink the regression coefficients corresponding to some unimportant features to 0 through the penalty coefficient λ. The optimal parameters are selected through cross-validation, and the final effect value is calculated based on the optimal parameters to ensure that the model retains important features while reducing noise and controlling overfitting. The individualized risk prediction model ( Figure 2 B) is as follows:
[0085] Y~(-0.980)×sgn(AC009229.5-RP5-1158E12.1)+0.131×sgn(AC087491.2-GPT)+(-0.625)×sgn(C1orf177-LOC 286189)+(-0.100)×sgn(C1orf177-RP11-75L1.1)+0.292×sgn(CYB5R2-KLK15)+(-1.487)×sgn(DKFZP434K028 -LAMP5)+(-0.331)×sgn(IGFL1-TRIM60)+0.150×sgn(LINC00635-RFTN2)+(-0.486)×sgn(LINC01010-RP11-69 C17.3)+(-0.242)×sgn(LOC100506384-TRIM60)+(-1.306)×sgn(LOC101060277-PRB4)+0.801×sgn(LOC286189- LOC400620)+(-0.774)×sgn(MEF2A-RAB24)+0.858×sgn(OR4C13-TBCA)+(-0.322)×sgn(RFTN2-RP11-104E19.1 )+(-1.249)×sgn(RIMBP2-RNU1-131P)+0.264×sgn(RNU6-798P-ZIC5)+0×sgn(RNU6-597P-RNU6-788P)+0×sgn( MAGI1-TRIM60)+0×sgn(LOC100130373-TMEM253)+0×sgn(KCNH1-TRIM60)+0×sgn(GHRHR-SNORD114-17)+0×sgn (FARP-AS1-MIR548M)+0×sgn(FANCD2OS-RP11-272D20.2)+0×sgn(CHST10-ZNF668)+0×sgn(AMHR2-LOC286189);
[0086] Among them, the coefficients of functions with different signs are the corresponding effect values of the 26 risk gene pairs.
[0087] Random stratified sampling was used to divide the blood samples of pregnant mothers whose offspring had ASD and TD into a training set (70%) and a validation set (30%) for risk assessment. The Wilcoxon rank sum test results showed that the prediction values were different between the blood groups of pregnant mothers whose offspring had ASD and TD ( Figure 2 C).
[0088] The ROC curve was drawn based on the prediction results. The results showed that the AUC of the training set reached 0.9993 and the AUC of the validation set reached 0.9648 ( Figure 2 D).
[0089] Bootstrap resampling was performed on the validation set 1000 times, and the AUC range was 0.8298~1( Figure 2 E), confirming that the individualized risk prediction model for MIA-induced ASD in offspring constructed by the present invention has stable, reliable and accurate prediction capabilities.
[0090] 3. Combined with prenatal transcriptomics data to assess the individual risk of ASD in offspring
[0091] Combined with the individualized risk prediction model constructed above, 17 risk gene pairs with effect values greater than 0 were selected, and the final individualized risk prediction model was constructed using the expression weighted effect value. The scores are as follows:
[0092] Risk score FPRscore=(-0.980)×sgn(AC009229.5-RP5-1158E12.1)+0.131×sgn(AC087491.2-GPT)+(-0.625)×sgn(C1orf177-LOC286189)+(-0.100)×sgn(C1orf177-RP11-75L1.1)+0.292×sgn(CYB5R2-KLK15)+(-1.487)×sgn(DKFZP434K028-LAMP5)+(-0.331)×sgn(IGFL1-TRIM60)+0.150×sgn(LINC00635-RFTN2)+(-0.486) ×sgn(LINC01010-RP11-69C17.3)+(-0.242)×sgn(LOC100506384-TRIM60)+(-1.306)×sgn(LOC101060277-PRB4)+0.801×sgn(LOC286189-LOC400620)+(-0.774 )×sgn(MEF2A-RAB24)+0.858×sgn(OR4C13-TBCA)+(-0.322)×sgn(RFTN2-RP11-104E19.1)+(-1.249)×sgn(RIMBP2-RNU1-131P)+0.264×sgn(RNU6-798P-ZIC5).
[0093] Pearson correlation analysis confirmed that 17 risk gene pairs were strongly associated with FPRscore (FDR<0.05), and T test found significant differences in FPRscore between maternal blood groups in offspring with ASD and TD ( Figure 3 At the same time, stratified sampling was performed to construct the risk scoring model, and combined with 1000 Bootstrap resampling, the validation set AUC was 0.821 ( Figure 3 (middle B).
[0094] The Gaussian mixture model was used to classify pregnant mothers with ASD offspring into three levels according to FPRscore: high risk (5.06-8.22), medium risk (-1.82-4.25), and low risk (-5.46--2.35). Figure 3 Middle C), the greater the maternal risk score, the higher the risk of ASD in the offspring. The differential analysis results showed that there were a large number of differentially expressed genes between the high-risk and low-risk groups (80 up-regulated and 20 down-regulated; Figure 3 (middle D).
[0095] Embodiment 2:
[0096] In order to demonstrate the reliability and stability of the individualized risk prediction model for offspring ASD during prenatal diagnosis constructed by the present invention, the present invention uses the constructed FPRscore to score the external independent ASD and matched control umbilical cord blood data, and combines resampling and drawing ROC curves to evaluate the model performance.
[0097] The results showed that in the umbilical cord blood dataset of ASD and matched controls (GSE123302; 51 ASD cases and 91 TD cases), the personalized risk prediction model of the present invention achieved an AUC of 0.713 (95% CI = 0.624-0.794; Figure 4 ).
[0098] Embodiment three:
[0099] A prenatal offspring individual ASD risk prediction system, comprising a data collection unit, a data analysis unit, a risk scoring unit and a result output unit;
[0100] The data collection unit is used to collect the expression feature order (FPR) of n' risk gene pairs in each pregnancy maternal blood sample whose offspring are ASD and TD;
[0101] The data analysis unit is used to use the FPRs of the n' risk gene pairs and adopt a machine learning algorithm to identify prenatal risk factors associated with the risk of ASD in offspring;
[0102] The risk scoring unit is used to construct an individualized risk score for the occurrence of ASD in offspring induced by MIA during pregnancy based on the obtained risk prediction model and the effect value in the risk prediction model, and to classify the risk into levels;
[0103] The result output unit is used to automatically generate a detailed assessment report based on the risk score and provide it to medical personnel, pregnant women and their accompanying persons.
[0104] Specifically:
[0105] Data collection unit: Detect the expression of n' risk gene pairs in peripheral blood during pregnancy and calculate the relative expression feature order (FPR). The blood transcription data of pregnant mothers with ASD and TD offspring can be derived from the transcriptomics data provided by GSE148450;
[0106] Data analysis unit: Blood samples of pregnant mothers whose offspring were ASD and TD were divided into training set (70%) and validation set (30%). The minimum absolute value shrinkage and selection operator (Lasso) algorithm was used for mixed feature selection in the training set, and the optimal parameter λ was obtained in combination with 10-fold cross validation to establish an individualized risk optimal assessment model. The ROC curve was drawn for verification in the validation set, and combined with 1000 Bootstrap resampling, the average AUC and 95% confidence interval were obtained. The risk prediction model is as follows:
[0107] Risk score FPRscore = (-0.980) × sgn (AC009229.5-RP5-1158E12.1) + 0.131 ×
[0108] sgn(AC087491.2-GPT)+(-0.625)×sgn(C1orf177-LOC286189)+(-0.100)×sgn(C1orf177-RP11-75L1.1)+0.292×sgn(CYB5R2-KLK15)+(-1.487)×
[0109] sgn(DKFZP434K028-LAMP5)+(-0.331)×sgn(IGFL1-TRIM60)+0.150×
[0110] sgn(LINC00635-RFTN2)+(-0.486)×sgn(LINC01010-RP11-69C17.3)+(-0.242)×sgn(LOC100506384-TRIM60)+( -1.306)×sgn(LOC101060277-PRB4)+0.801×sgn(LOC286189-LOC400620)+(-0.774)×sgn(MEF2A-RAB24)+0.858×
[0111] sgn(OR4C13-TBCA)+(-0.322)×sgn(RFTN2-RP11-104E19.1)+(-1.249)×
[0112] sgn(RIMBP2-RNU1-131P)+0.264×sgn(RNU6-798P-ZIC5); where the symbol function in the model risk score formula is the order of the relative expression feature pairs of gene pairs of a single sample.
[0113] Risk scoring unit: A Gaussian mixture model is used to classify pregnant women into three levels: high risk (5.06-8.22), medium risk (-1.82-4.25), and low risk (-5.46--2.35). The larger the risk score of the pregnant woman, the higher the risk of ASD in the offspring.
[0114] Result output unit: Provide individualized health management advice and recommend specific medical interventions based on the risk level of the pregnant woman.
[0115] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A method for constructing a risk prediction model for autism spectrum disorder in offspring individuals, characterized in that: The steps include: Identify genes ranked in the top 10% by absolute fold difference in maternal blood transcriptome data from a prospective cohort of offspring with autism spectrum disorder and typical development; Compare the expression levels of any two genes among several genes to obtain the order spectrum of expression feature pairs; Based on the expression signature pair order profiles and compared with the maternal blood during pregnancy associated with typical development of the offspring, identifying risk gene pairs with significantly reversed expression signature pair order in the maternal blood during pregnancy associated with autism spectrum disorder in the offspring; A mixed feature selection is performed on the risk gene pairs, and the optimal parameters are obtained using cross-validation. The effect value is calculated according to the optimal parameters, and an individualized risk prediction model is constructed based on the effect value and the results after the mixed feature selection.
2. The construction method according to claim 1, characterized in that: The blood transcriptional data of mothers with ASD and typical offspring during pregnancy were derived from the transcriptomics data provided by GSE148450.
3. The construction method according to claim 1, characterized in that: The method for obtaining the order spectrum of expression features specifically includes: Select any two genes from the plurality of genes, compare the expression levels, and construct an order symbol scoring function based on the comparison result to obtain an order spectrum of expression feature pairs after the gene expression profiles of blood samples of pregnant mothers whose offspring have autism spectrum disorder and whose offspring have typical development are converted; The obtained expression feature pair order spectrum is: if the expression level of one gene is greater than that of another gene, the expression feature pair order is recorded as 1; if the expression level of one gene is less than that of another gene, the expression feature pair order is recorded as -1.
4. The construction method according to claim 1, characterized in that: The risk gene pairs were identified by: Fisher's exact test was performed for each gene pair between the blood panels of mothers with offspring with ASD and offspring with typical development; The P values obtained by Fisher's exact test were corrected for multiple testing, and the gene pairs with P < 0.05 after correction were considered risk gene pairs.
5. The construction method according to claim 1, characterized in that: The risk gene pairs identified are AC009229.5-RP5-1158E12.1, AC087491.2-GPT 、 C1orf177-LOC286189, C1orf177-RP11-75L1.1, CYB5R2-KLK15, DKFZP434K028-LAMP5, IGFL1-TRIM60, LINC00635-RFTN2, LI NC01010-RP11-69C17.3, LOC100506384-TRIM60, LOC101060277-PRB4, LOC286189-LOC400620, MEF2A-RAB24, OR4C13-TBCA, RFTN2-RP11-104E19.1, RIMBP2-RNU1-131P, RNU6-798P-ZIC5, RNU6-597P-RNU6-788P, MAGI1-TRIM60, LOC100130373-TMEM2 53. KCNH1-TRIM60, GHRHR-SNORD114-17, FARP-AS1-MIR548M, FANCD2OS-RP11-272D20.2, CHST10-ZNF668, AMHR2-LOC286189.
6. The construction method according to claim 1, characterized in that: Hybrid feature selection is performed using the minimum absolute value shrinkage and selection operators.
7. The construction method according to claim 1, characterized in that: The risk gene pairs with effect values greater than 0 are selected, and the individualized risk prediction model is constructed using the expression weighted effect values.
8. A system for predicting the risk of autism spectrum disorder in offspring individuals during pregnancy and before birth, characterized in that: It includes a data collection unit, a data analysis unit, a risk scoring unit and a result output unit; The data collection unit is used to collect the order of expression feature pairs of risk gene pairs in the blood samples of mothers in each pregnancy whose offsprings have autism spectrum disorders and whose offsprings have typical development; The data analysis unit is used to use the expression feature pair order of the risk gene pair and identify the prenatal risk factors associated with the risk of autism spectrum disorder in the offspring based on the machine learning algorithm; The risk scoring unit is used to construct an individualized risk score for the occurrence of autism spectrum disorder in offspring during pregnancy using the individualized risk prediction model constructed according to the method according to any one of claims 1 to 7, and to perform grade classification according to the risk score result; The result output unit is used to generate an assessment report based on the risk scoring result.
9. The risk prediction system according to claim 8, characterized in that: The data collection unit is used to collect the expression levels of risk gene pairs in peripheral blood during pregnancy, and calculate the order of expression feature pairs based on the expression levels.
10. The risk prediction system according to claim 8, characterized in that: The data analysis unit is used for: Blood samples of pregnant mothers whose offspring have autism spectrum disorder and whose offspring have typical development are divided into a training set and a validation set, hybrid feature selection is performed on the training set using the minimum absolute value shrinkage and selection operator algorithm, and the optimal parameters are obtained in combination with cross-validation, the effect value is calculated according to the optimal parameters, and an individualized risk prediction model is constructed based on the effect value and the hybrid feature selection results; The ROC curve was drawn to verify the constructed individualized risk prediction model in the validation set, and combined with resampling, the average AUC value and confidence interval were obtained to evaluate the individualized risk prediction model.