A classification, diagnostic, and predictive method, system, and electronic device for congenital neurodevelopmental disorders.

By using whole-genome methylation signatures and customized SVM models, the challenge of accurate diagnosis of congenital neurodevelopmental disorders has been solved, enabling one-stop differential diagnosis of multiple mechanisms and diseases, improving diagnostic efficiency and accuracy, and making it suitable for clinical application.

CN122135945APending Publication Date: 2026-06-02THE INTERNATIONAL PEACE MATERNITY & CHILD HEALTH HOSPITAL OF CHINA WELFARE INSTITUTE

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE INTERNATIONAL PEACE MATERNITY & CHILD HEALTH HOSPITAL OF CHINA WELFARE INSTITUTE
Filing Date
2026-04-13
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies are inefficient in the molecular diagnosis of congenital neurodevelopmental disorders, especially for occult variants and deep intron or complex structural variants. They are difficult to accurately classify and intervene early, and existing methods are costly and difficult to implement routinely.

Method used

By combining whole-genome methylation signatures with a customized support vector machine (SVM) model, methylation microarray detection is performed on peripheral blood samples to screen disease-specific probes and construct a methylation signature matrix, enabling one-stop differential diagnosis of multiple mechanisms and diseases.

Benefits of technology

It enables accurate diagnosis of six rare congenital neurodevelopmental disorders, with no false positives or false negatives in the test. The sample is non-invasive, the process is standardized, the cost is low, and it is suitable for routine clinical use, filling a gap in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135945A_ABST
    Figure CN122135945A_ABST
Patent Text Reader

Abstract

This disclosure belongs to the field of medical testing technology, and specifically relates to a classification, diagnostic, and prediction method, system, and electronic device for congenital neurodevelopmental disorders. Addressing the clinical diagnostic challenges of congenital neurodevelopmental disorders, this disclosure establishes for the first time a classification, diagnostic, and prediction method for congenital neurodevelopmental disorders based on whole-genome methylation signatures and customized SVM machine learning. This method achieves one-stop differential diagnosis across multiple mechanisms and diseases, filling a gap in existing technologies and providing a novel and efficient solution for the accurate diagnosis of congenital neurodevelopmental disorders.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical testing technology, and specifically relates to a classification, diagnosis and prediction method, system and electronic equipment for congenital neurodevelopmental disorders. Background Technology

[0002] Studies have shown that more than half of gene (genome) variations can cause congenital neurodevelopmental disorders (NDDs) with or without other systemic abnormalities, leading to impairment of the central and / or peripheral nervous system. Affected patients may present with symptoms such as developmental delay, intellectual disability, language impairment, abnormal muscle tone, epilepsy, and / or ataxia; these diseases have complex phenotypes and mostly lack clinical specificity, making clinical diagnosis more difficult.

[0003] Our team is one of the earliest clinical laboratories in China to apply next-generation sequencing (NGS) on a large scale for the diagnosis of genetic diseases (Genet Med 2018; Clin Chem 2014). We have diagnosed more than 30,000 patients, involving more than 600 diseases. We have established multiple patient cohorts (JCEM 2022; Clin Genet 2019; EurJ Endocrinol 2019, etc.) and developed novel diagnostic strategies for some genetic diseases using transcriptome sequencing and in vitro functional validation (Clin Chim Acta 2022; Exp Hematol 2018). We have identified 6 new pathogenic genes (NEJM 2014; eBioMedicine 2024; J Bone Miner Res 2016, etc.) and revealed a number of novel pathogenic variants / mechanisms (J BoneMiner Res 2023; Nucleic Acids Res 2022; Cell Rep 2022, etc.). A review of diagnosed cases revealed a relatively low molecular diagnostic positivity rate (<30%) for NDD-type genetic diseases, compared to the nearly 50% positivity rate in other systems (such as cardiomyopathy and metabolic diseases). This is primarily due to the presence of numerous variants of undetermined clinical significance (VUS), and the fact that some cases may carry deep introns or complex structural variations undetectable by conventional NGS techniques (such as whole-exome sequencing). This significantly complicates accurate patient typing, early intervention, and family decisions regarding reproduction. Therefore, there is an urgent clinical need to develop new methods to improve the efficiency of molecular diagnosis for these diseases.

[0004] Current research on novel molecular diagnostic strategies for congenital NDD focuses primarily on developing NGS-assisted diagnostic methods to address the challenges of VUS or occult variants (such as deep introns and structural variations). For occult variants, detection methods such as targeted / whole-genome sequencing, third-generation sequencing, and optical genome mapping have been developed; however, due to high testing costs and the need to define candidate gene ranges during data analysis, these methods are currently difficult to routinely implement in clinical practice and are generally used for identifying occult pathogenic loci in specific cases (e.g., in recessive genetic diseases where only one pathogenic locus is detected). Several studies, including our team's, have also explored the clinical efficacy of RNA-seq based on patient peripheral blood samples in genetic molecular diagnosis. This technology identifies candidate pathogenic genes and variants through comprehensive analysis of alternative splicing, expression abnormalities, and fusion gene formation; however, similar to the methods mentioned above, RNA-seq analysis typically requires targeting specific candidate genes. Although studies have shown that RNA-seq can assist in the diagnosis of NGS-negative samples to some extent, its application in the molecular diagnosis of congenital NDD is also limited due to the limited expression of neurodevelopment-related genes in peripheral blood and the significant impact of age on gene expression. Furthermore, in vitro and in vivo functional studies are an effective way to address the pathogenicity of VUS; however, the complexity of studying nerve cell function and the difficulty in accessing patient neural tissue make this method difficult to routinely perform in clinical laboratories.

[0005] DNA methylation is an epigenetic modification that alters DNA structure and chemical properties, thereby affecting chromatin assembly and gene transcription. Genome-wide DNA methylation characteristic alteration analysis has demonstrated its value in the clinical diagnosis, subtyping, and prognostic prediction of various diseases, such as predicting early preeclampsia and analyzing the diagnosis and prognosis of tumors like esophageal squamous cell carcinoma based on changes in cell-free DNA methylation in pregnant women's plasma. Existing research indicates that mutations in genes encoding epigenetic regulatory mechanisms, including those involved in reading, writing, and erasing histone post-translational signals and chromatin remodeling, are closely related to the occurrence and development of various NDD diseases and are often accompanied by genome-wide DNA methylation abnormalities. Therefore, performing genome-wide methylation detection on NDD patients and identifying differentially expressed sites to establish corresponding disease-specific methylation signatures (Epi-signatures) has become a feasible new molecular diagnostic strategy. DNA methylation detection in these patients can not only help solve the problem of VUS pathogenicity, but in some cases, specific DNA methylation signatures may be the only molecular discovery.

[0006] Therefore, there is an urgent clinical need to develop disease classification and diagnostic prediction methods for NDD based on genomic methylation signatures to improve the efficiency of molecular diagnosis of these diseases. Summary of the Invention

[0007] The purpose of this invention is to provide a method, system, and electronic device for the classification, diagnosis, and prediction of NDD based on genomic methylation feature profiles, enabling one-stop differential diagnosis of multiple mechanisms and diseases, filling the gap in existing technologies, and providing a novel and efficient solution for the accurate diagnosis of congenital neurodevelopmental disorders.

[0008] The objective of this invention is achieved through the following technical solution: In a first aspect, the present invention provides a method for classifying, diagnosing, and predicting congenital neurodevelopmental disorders, comprising the following steps: Methylation chip detection and data standardization were performed on peripheral blood samples that had passed collection and quality inspection to obtain a whole-genome methylation analysis matrix of the samples. From the whole genome methylation analysis matrix, methylation-specific probe sets for various congenital neurodevelopmental disorders are screened, and the corresponding methylation signals are extracted to obtain the methylation feature matrix of each disease; Each disease-specific methylation feature matrix is ​​input into the corresponding congenital neurodevelopmental disorder diagnosis and prediction model to calculate the pathogenicity score of methylation variation for each disease. Based on the pathogenicity score of methylation variation corresponding to each disease, the samples were classified and diagnosed as congenital neurodevelopmental disorders. The congenital neurodevelopmental disorder is selected from at least one of Sotos syndrome, Kabuki syndrome type I, CHARGE syndrome, Wiedemann-Steiner syndrome, Williams-Beuren syndrome, and Prader-Willi syndrome.

[0009] In some specific embodiments of the present invention, the methylation chip detection includes: using Illumina Infinium... TM The Methylation V2.0 (935K) chip was used to detect peripheral blood genomic DNA transformed with Bisulfite to obtain raw idat detection data. The raw detection data was read using the R language meffil package to obtain a whole-genome methylation β-value matrix.

[0010] In some specific embodiments of the present invention, the detection data standardization processing includes: After quality filtering of the whole genome methylation β value matrix, quantile normalization and chip batch factor correction were performed. Non-finite values ​​in the whole genome methylation β value matrix are defined as missing values. Probes with a missing ratio >15% are removed. The remaining missing values ​​are iteratively reconstructed and completed using a PCA model with 50 principal components. The completed β values ​​are restricted to the range of 0 to 1. After sample quality control and correction for variations in batch, age, sex, and cell composition, the whole genome methylation analysis matrix was obtained.

[0011] In some specific embodiments of the present invention, the step of screening a set of methylation-specific probes for various congenital neurodevelopmental disorders from the whole-genome methylation analysis matrix includes: Differential methylation sites between various congenital neurodevelopmental disorders and healthy controls were calculated from the genome-wide methylation analysis matrix to obtain a complete set of candidate probes for each disease; From the entire candidate probe set for each disease, a methylation-specific probe set for each disease was selected. Specifically: (1) For Sotos syndrome, Kabuki syndrome type I, CHARGE syndrome or Wiedemann-Steiner syndrome, the candidate probes are subjected to preliminary screening and specific screening in sequence to obtain a set of methylation-specific probes for each disease; (2) For Williams-Beuren syndrome, after the candidate probes are screened for initial screening and specific screening in sequence, the candidate probe set for the disease is screened for intersection with the probe sites related to the disease reported in the literature. The probes obtained from specific screening and the probes obtained from intersection screening together constitute the methylation specific probe set for the disease. (3) For Prader-Willi syndrome, the candidate probe set for the disease was screened by intersection with the probe sites related to the disease reported in the literature, and filtered under a fixed threshold |Δβ|≥0.01 to obtain the methylation-specific probe set for the disease.

[0012] In some embodiments of the present invention, the preliminary screening includes: Significance screening was performed using an adjusted P-value < 0.05, and effect strength filtering was performed using the |Δβ| value determined by a fixed threshold and a distribution adaptive threshold. The initial screening probe count has a lower limit of 300 and an upper limit of 3000. If the number is insufficient, the threshold will be relaxed to make up the difference. If the number is excessive, the probes ranked higher will be selected.

[0013] In some embodiments of the present invention, the specific screening includes: Based on the similarity of probe effect patterns among diseases, identify the set of indistinguishable diseases for the target disease; For each initial screening probe, a comprehensive ranking score is calculated. The comprehensive ranking score consists of a positive support score and an overlap penalty score, and its calculation principle is as follows: in, This represents the overall ranking score of the candidate probes. This represents the positive support score, which includes the basic discovery score and the AUC value for each scenario. This indicates the overlap penalty score, which includes overlap ratings and hit rate. according to Perform conservative filtering and retain only Probes with a strength ≥0.85, wherein, This indicates the conservative discrimination threshold for the probe. This indicates the preset minimum target disease-control discrimination threshold. This represents the minimum threshold for distinguishing between target and non-target diseases. This indicates the lowest conservative discrimination threshold for distinguishing between diseases that are difficult to differentiate. The number of specific probes is limited to a minimum of 100 and an upper limit of 300. If the number is insufficient, the effect strength threshold is relaxed to compensate. If the number is excessive, correlation pruning is performed based on the Pearson correlation coefficient.

[0014] In some embodiments of the present invention, the set of methylation-specific probes for Sotos syndrome is as follows: A total of 150 specific probes; the methylation-specific probe set for Kabuki syndrome type I is as follows: A total of 145 specific probes; the methylation-specific probe set for CHARGE syndrome is as follows: A total of 158 specific probes; the methylation-specific probe set for Wiedemann-Steiner syndrome is as follows: A total of 150 specific probes were identified; the set of methylation-specific probes for Williams-Beuren syndrome is as follows: A total of 232 specific probes were identified; the methylation-specific probe set for Prader-Willi syndrome is as follows: A total of 219 specific probes were used.

[0015] In some embodiments of the present invention, the diagnostic prediction model for congenital neurodevelopmental disorders is constructed through the following steps: (1) For each target congenital neurodevelopmental disorder, a “one-vs-all” binary classification model is constructed, with the disease sample set as the positive class and healthy controls and other non-target disease samples set as the negative class; (2) Using the feature spectrum matrix corresponding to the methylation-specific probe set of each disease as input, a linear kernel support vector machine (SVM) is used for model training. The model is trained by grid search on a logarithmic scale of 10. -5 Up to 10 4 The penalty parameter C was optimized within the range, and repeated stratified cross-validation was used to determine the optimal penalty parameter C for each disease based on the optimal F1 value. (3) In model training, preset empirical weights are assigned to target disease samples, healthy control samples and non-target disease samples respectively to alleviate the class imbalance problem; the Platt scaling method is used to calibrate the output of the basic SVM to a prediction probability of 0-1, which is defined as the methylation variant pathogenicity score (MVP score).

[0016] In some embodiments of the present invention, the classification and diagnosis of congenital neurodevelopmental disorders based on the pathogenicity score of methylation variants corresponding to each disease includes: determining the pathogenicity score of methylation variants (MVP score) according to a preset threshold, wherein an MVP score ≥ 0.5 is determined to be a positive disease, an MVP score < 0.1 is determined to be a negative disease, and 0.1 ≤ MVP score < 0.5 is determined to be an unclear result, thereby achieving the classification and diagnosis of each congenital neurodevelopmental disorder.

[0017] In a second aspect, the present invention provides a classification, diagnostic, and prediction system for congenital neurodevelopmental disorders, comprising: A building module is used to construct the diagnostic and predictive model for the aforementioned congenital neurodevelopmental disorders; The acquisition module is used to acquire the detection data after sample methylation detection, and to standardize the detection data to obtain the whole genome methylation analysis matrix of the sample, and to screen out the disease-specific methylation feature matrix. The model processing module is used to input the disease-specific methylation feature matrix data into the corresponding congenital neurodevelopmental disorder diagnosis and prediction model, and calculate the pathogenicity score of methylation variation for each disease. The evaluation module is used to classify and diagnose congenital neurodevelopmental disorders based on the pathogenicity scores of methylation variations corresponding to each disease.

[0018] In a third aspect, the present invention provides an electronic device comprising: Processor; and, A memory storing computer-executable instructions, which, when executed, cause the processor to perform the method according to any one of the preceding statements.

[0019] In a fourth aspect, the present invention provides a computer-readable storage medium that stores one or more programs that, when executed by a processor, implement the method described in any one of the preceding descriptions.

[0020] The technical solution provided by this invention has the following technical contributions: (1) This disclosure addresses the pain points in the clinical diagnosis of congenital neurodevelopmental disorders by establishing for the first time a disease classification and diagnostic prediction method for NDD based on whole-genome methylation feature profiles and customized SVM machine learning. Its core advantages are: ① Both the probe and the model are customized for the characteristics of small samples of neurodevelopmental disorders and rare diseases; ② Excellent diagnostic performance with no false positives / false negatives in the test cohort, effectively solving the problem of VUS determination; ③ Non-invasive sample preparation, standardized process, low cost, and easy to routinely carry out in clinical practice; ④ Achieves one-stop differential diagnosis of multiple mechanisms and multiple diseases, filling the gap in existing technologies and providing a new and efficient solution for the accurate diagnosis of congenital neurodevelopmental disorders.

[0021] (2) The six diseases of congenital neurodevelopmental disorders disclosed in this invention are all rare diseases with small clinical sample sizes (the training cohort of this scheme has only 20-39 cases of a single disease). Moreover, the pathogenic mechanisms cover single gene point mutations (Sotos / Kabuki), chromosomal fragment deletions (WBS), and abnormal genomic methylation (PWS). Existing technologies are mostly for single-mechanism detection (such as MLPA which only detects copy number variations). However, this scheme can achieve one-stop differential diagnosis of multiple mechanisms and multiple diseases through a whole-genome methylation detection process, which is suitable for the detection needs of complex clinical cases.

[0022] (3) Congenital neurodevelopmental disorders often share common phenotypes such as developmental delay and intellectual disability, making clinical identification difficult by visual inspection. However, the 1052 specific probes selected in this protocol, combined with a one-vs-all SVM model, can accurately distinguish between six diseases. In the test cohort, only two samples were in the gray range, while the rest were accurately matched for diagnosis. Furthermore, it can assist in identifying different disease subtypes caused by the same gene mutation (such as...). CHD7 (Charge syndrome and Kaman syndrome caused by mutations).

[0023] (4) Existing conventional whole exome sequencing and other NGS technologies have a positive diagnostic rate of <30% for this type of disease, while the ROC_AUC of the 6 disease models in this protocol is ≥96.7%, accuracy is ≥96.2%, and precision is ≥94.4%. Among them, Sotos syndrome achieves 100% in AUC, accuracy, precision, and recall. In the test cohort, only 2 out of 89 samples were in the gray range, with no false positive / false negative results. It can also accurately identify chimeric mutation samples and samples with atypical clinical phenotypes, and its diagnostic efficacy is significantly better than existing technologies.

[0024] (5) For VUS samples in diseases such as WSS and Kabuki, this scheme can achieve accurate identification through methylation feature profile and MVP score (e.g., 1 case of VUS in WSS was identified as negative by the model; 2 cases of VUS in Kabuki were identified as negative due to clinical phenotype mismatch), which solves the clinical pain points of VUS causing difficulty in accurate patient classification and family decision-making on re-fertility, provides clear basis for genetic counseling, and becomes an important supplementary scheme to existing gene testing.

[0025] (6) This disclosure establishes a standardized process from peripheral blood DNA collection to model output results, including strict sample quality control, probe quality control, and a unified data analysis process. Peripheral blood samples are used throughout the process. Compared with tissue samples required for RNA-seq and high-purity DNA samples required for third-generation sequencing, it is easier to carry out routinely in clinical laboratories. Sample acquisition is non-invasive and convenient. Attached Figure Description

[0026] Figure 1 This is a flowchart for constructing disease-specific methylation signature profiles.

[0027] Figure 2 This is the process of establishing an SVM classification, diagnosis, and prediction model.

[0028] Figure 3 This is a heatmap showing the differences in methylation between the six disease samples and the control samples.

[0029] Figure 4 This is a test plot of the SVM model for WSS disease. The top part is a box plot of the WSS disease samples, with red dashed lines marking 0.5 (positive threshold) and blue dashed lines marking 0.1 (negative threshold), representing positive (≥0.5), grayscale range (0.1~0.5), and negative (<0.1). Below are the ROC (AUC=0.997, 95% CI 0.990-1.000) plots for the training set and the ROC (AUC=1, 95% CI 1.000-1.000) plots for the test set, respectively.

[0030] Figure 5This is a test plot of the SVM model for Sotos disease. The top part is a box plot of the Sotos disease samples, with red dashed lines marking 0.5 (positive threshold) and blue dashed lines marking 0.1 (negative threshold), representing positive (≥0.5), grayscale range (0.1~0.5), and negative (<0.1). Below are the ROC (AUC=1.000, 95% CI 1.000-1.000) plots for the training set and the test set for Sotos disease, respectively.

[0031] Figure 6 This is a test plot of the SVM model for Kabuki disease type I. The top part is a box plot of the type I Kabuki disease samples, with red dashed lines marking 0.5 (positive threshold) and blue dashed lines marking 0.1 (negative threshold), representing positive (≥0.5), grayscale range (0.1~0.5), and negative (<0.1). Below are the ROC plots for the training set (AUC=0.967, 95% CI 0.888-1.000) and test set (AUC=1.000, 95% CI 1.000-1.000) of the type I Kabuki disease model.

[0032] Figure 7 This is a test plot of the SVM model for WBS disease. The top part is a box plot of the WBS disease samples, with red dashed lines marking 0.5 (positive threshold) and blue dashed lines marking 0.1 (negative threshold), representing positive (≥0.5), grayscale range (0.1~0.5), and negative (<0.1). Below are the ROC (AUC=0.999, 95% CI 0.996-1.000) for the training set and the ROC (AUC=1.000, 95% CI 1.000-1.000) for the test set, respectively.

[0033] Figure 8 This is a test plot of the SVM model for the CHARGE disease. The top part is a box plot of the CHARGE disease samples, with red dashed lines marking 0.5 (positive threshold) and blue dashed lines marking 0.1 (negative threshold), representing positive (≥0.5), grayscale range (0.1~0.5), and negative (<0.1). Below are the ROC (AUC=0.986, 95% CI 0.969-0.998) for the training set and the ROC (AUC=1.000, 95% CI 1.000-1.000) for the test set, respectively.

[0034] Figure 9This is a test plot of the SVM model for PWS disease. The top part is a box plot of the PWS disease samples, with red dashed lines marking 0.5 (positive threshold) and blue dashed lines marking 0.1 (negative threshold), representing positive (≥0.5), grayscale range (0.1~0.5), and negative (<0.1). Below are the ROC (AUC=0.994, 95% CI 0.985-0.999) for the training set and the ROC (AUC=1.000, 95% CI 1.000-1.000) for the test set, respectively.

[0035] Figure 10 This disclosure provides a schematic diagram illustrating the principle of the diagnostic and predictive method for congenital neurodevelopmental disorders.

[0036] Figure 11 This is a schematic diagram of the structure of the diagnostic and prediction system for congenital neurodevelopmental disorders provided in this disclosure.

[0037] Figure 12 This is a schematic diagram of the structure of an electronic device provided in this disclosure.

[0038] Figure 13 This is a schematic diagram of a computer-readable medium provided in this disclosure. Detailed Implementation

[0039] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0040] Unless otherwise specified, experimental methods in the following examples were performed under standard conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or as recommended by the manufacturer. All commonly used reagents used in the examples are commercially available products.

[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0042] In the following embodiments, unless otherwise specified, the terms used are explained as follows: In this disclosure, unless otherwise specified, the term "IDAT file or Intensity Data" refers to the binary raw data format of Illumina chips. It does not directly contain the methylation rate (Beta value), but rather the MeanSignal, Standard Deviation, and Number of Beads for each probe.

[0043] In this disclosure, unless otherwise specified, the term "differentially methylated position (DMP)" refers to a genomic location where the DNA methylation level of a single CpG site differs statistically significantly under different biological conditions (e.g., disease / normal, treatment / control). It is the most fundamental and core analytical object in epigenome association studies (EWAS).

[0044] In this disclosure, unless otherwise specified, the term "differentially methylated region (DMR)" refers to a region of the genome (typically containing multiple CpG sites, generally hundreds of bp to several kb in length) where the overall methylation level differs significantly under different biological states (e.g., disease / normal, different developmental stages, treatment / control). Compared to a single DMP, the DMR better reflects the regional coordinated regulation of epigenetics and is a core object for elucidating gene expression regulation mechanisms.

[0045] In this disclosure, unless otherwise specified, the term "β value" is a quantitative indicator used to represent the degree of methylation at a single CpG site, ranging from 0 to 1 (0% to 100% methylation); the formula is: β = methylation signal / (methylation signal + unmethylation signal + constant).

[0046] In this disclosure, unless otherwise specified, the term "Methylation Difference" refers to the absolute value of the difference in methylation levels between two groups of samples at a certain site or region, which can be expressed as: |Δβ| = |β1 β2|. Where Δβ represents the difference in methylation β values ​​between the two groups of samples. Generally, |Δβ|≥0.1 indicates a significant difference in methylation, |Δβ|≥0.2 indicates a large difference, while |Δβ|<0.05 usually indicates a small difference.

[0047] In this disclosure, unless otherwise specified, the term "Sotos syndrome" refers to a group of congenital neurodevelopmental disorders characterized by excessive growth (significantly increased height / head circumference after birth), distinctive facial features (high forehead, long face, prominent jaw), intellectual disability, premature skeletal maturation, and delayed language / motor development. Approximately 80% of patients exhibit these characteristics. NSD1 Mutations / deletions are predominantly de novo mutations, with a few familial mutations. These mutations / deletions can affect H3K36me2-related epigenetic regulation and are associated with characteristic changes in the genome-wide DNA methylation profile.

[0048] In this disclosure, unless otherwise specified, the term " NSD1 "Refers to the chromosome 5q35" NSD1 The gene encodes histone lysine methyltransferase, specifically the nuclear receptor-binding SET domain protein 1. Mutations in this protein are primarily associated with diseases such as Sotos syndrome and acute myeloid leukemia (AML), with the mechanisms often involving loss of function or abnormal fusion. The core function of the NSD1 protein is to catalyze the modification of histone H3K36me2, regulating chromatin structure and gene expression, and is crucial for embryonic development, cell proliferation, and differentiation. Abnormal function can lead to epigenetic regulatory disorders, resulting in developmental abnormalities or tumors.

[0049] In this disclosure, unless otherwise specified, the term "Kabuki syndrome" (KS, also known as Niikawa-Kuroki syndrome) refers to a rare congenital genetic disorder, primarily caused by... KMT2D or KDM6A Caused by gene mutations, it is characterized by distinctive facial features, intellectual disability, growth retardation, and skeletal abnormalities. It belongs to autosomal dominant or X-linked dominant inheritance, and most of them are de novo mutations.

[0050] In this disclosure, unless otherwise specified, the term " KMT2D "" refers to the gene encoding histone H3K4 methyltransferase, a term " KDM6A "" refers to the gene encoding H3K27 demethylase. Both are involved in epigenetic regulation. Mutations in this gene lead to disordered gene expression and cause multisystemic developmental abnormalities.

[0051] In this disclosure, unless otherwise specified, the term "CHARGE Syndrome" refers to a rare, multisystemic congenital genetic disorder, named from the acronym of its core clinical features. It is a neural crest developmental disorder caused by mutations in the CHD7 gene, exhibiting highly heterogeneous clinical manifestations. Approximately 60%–90% of clinically diagnosed patients have detectable abnormalities. CHD7 Gene mutation.

[0052] In this disclosure, unless otherwise specified, the term " CHD7"(Chromodomain Helicase DNA-binding protein 7)" refers to the protein located on chromosome 8q12.2. CHD7 Gene. CHD7 It encodes an ATP-dependent chromatin remodeling factor belonging to the Snf2 family. It regulates chromatin structure and is crucial for the migration and differentiation of neural crest cells.

[0053] In this disclosure, unless otherwise specified, the term "Wiedemann–Steiner syndrome, OMIM #605130, WSS" refers to a rare autosomal dominant neurodevelopmental disorder caused by… KMT2A (MLL) Caused by gene mutation, the core manifestations are distinctive facial features, intellectual disability, short stature, and excessive hair growth on the elbows.

[0054] In this disclosure, unless otherwise specified, the term " KMT2A (Histone-lysine N-methyltransferase 2A) (formerly known as...) MLL, ALL-1, HRX H3K4 is a key epigenetic regulatory gene located at 11q23.3, which encodes histone H3K4 methyltransferase. Its core function is to regulate embryonic development, hematopoiesis, and neurogenesis. Its germline heterozygous mutation can lead to a rare developmental syndrome (Wiedemann-Steiner syndrome).

[0055] In this disclosure, unless otherwise specified, the term "Williams-Beuren Syndrome (WBS)" refers to a rare neurodevelopmental disorder. Its core cause is a microdeletion in the 7q11.23 region of chromosome 7, resulting in the loss of multiple genes (primarily...) within that region. ELN and LIMK1 The haploid dose of the disease is insufficient. The disease is characterized by distinctive "elfin" facial features, cardiovascular lesions such as supravalvular aortic stenosis (SVAS), a "cocktail" social personality (extremely extroverted but lacking social boundaries), and unique cognitive deficits (poor spatial faculties but relatively preserved language and musical abilities).

[0056] In this disclosure, unless otherwise specified, the term "7q11.23 deletion" refers to a microdeletion in the q11.23 region of chromosome 7, with a deletion fragment approximately 1.5-1.8 Mb in length. More than 95% of patients with Williams syndrome have a 7q11.23 deletion. This region contains highly homologous low-copy repeats (LCRs), which are prone to non-allelic homologous recombination (NAHR) during meiosis, leading to fragment loss.

[0057] In this disclosure, unless otherwise specified, the term "Prader-Willi Syndrome (PWS)," commonly known as "Fatty-Willi Syndrome," refers to a rare and complex neurodevelopmental genetic disorder caused by the loss of expression of a paternal gene in a region on the long arm of chromosome 15 (15q11.2-q13). The maternal gene in this region is usually imprinted and silent (methylated), but when the paternal gene in this region is absent or inactivated, the disease develops. This disease is one of the most common genetic disorders of obesity in humans, and its clinical manifestations show distinct stages with age. Its core features also include intellectual disability, severe hypotonia in infancy, and uncontrollable hyperphagia beginning in childhood.

[0058] In this disclosure, unless otherwise specified, the term "abnormal methylation of the 15q11.2-q13.1 region" refers to the abnormal methylation status of the imprinted region on the long arm of chromosome 15 (15q11.2-q13.1), a core molecular genetic indicator for the diagnosis of Prader-Willi Syndrome (PWS) or Angelman Syndrome (AS). The q11.2-q13 region of chromosome 15 is a classic imprinted region in the human genome.

[0059] The selection and quantity of these reagents fall within the scope of the expertise and routine techniques of those familiar with this technology. The following description, in conjunction with specific examples, illustrates this point.

[0060] Example 1. Methylation signature profiles of six patients with congenital neurodevelopmental disorders 1. Composition of the "Training Queue" Including: Sotos syndrome ( NSD1 Gene mutation), Kabuki syndrome type I ( KMT2D Gene mutation), CHARGE syndrome ( CHD7 Gene mutation), Wiedemann–Steiner syndrome (WSS); KMT2A The study included six diseases: gene mutation, Williams-Beuren syndrome (WBS; copy number deletion in the 7q11.23 region), and Prader-Willi syndrome (PWS; methylation abnormality in the 15q11.2-13.1 region). All cases were clinically diagnosed and had clear gene testing results. The number of cases is shown in Table 1.

[0061] Table 1. Types of diseases and number of cases involved in this application 1.16 Diagnostic Challenges of the Included Diseases The six diseases listed in Table 1 are syndrome-type neurodevelopmental disorders, characterized by strong clinical heterogeneity and a lack of standardized clinical diagnostic criteria, making accurate diagnosis difficult based on clinical features alone. Diagnosis of these diseases heavily relies on molecular testing. However, NSD 1. KMT2A , KMT2D and CHD7 These genes belong to a class of pathogenic genes with large molecular weights, and the probability of rare missense mutations is relatively high. In clinical practice, high-throughput genome sequencing frequently detects de novo (meaning both parents are normal wild-type, but the patient carries the mutation) and indel (meaning whole-frame deletions or insertions that do not affect the amino acid reading frame) mutations in these genes. According to current clinical guidelines, these mutations can often only be assessed as undetermined significance (VUS), leading to numerous adjustments in molecular diagnosis. Therefore, there is an urgent need to develop novel molecular diagnostic methods independent of genomic variant detection to effectively solve the diagnostic challenges of VUS cases.

[0062] 1.2 Molecular basis of genomic DNA methylation alterations in included diseases NSD1 The gene encodes a histone methyltransferase that specifically catalyzes the dimethylation of lysine at position 36 of histone H3 (H3K36me2). This modification can directly recruit DNA methyltransferase 3A (DNMT3A) to mediate DNA methylation modification.

[0063] KMT2A The gene encodes a histone methyltransferase that specifically catalyzes the trimethylation of lysine 4 of histone H3 (H3K4me3). As a marker of transcriptional activation, H3K4me3 can antagonize the recruitment and allosteric activation of the DNA methyltransferase complex DNMT3L-DNMT3A / B, thereby preventing CpG methylation of DNA molecules and gene expression silencing. KMT2D The gene encodes a histone methyltransferase that catalyzes H3K4me1 / me3 modification and acts as a regulator of enhancer activity. By maintaining H3K4me1 labeling, it prevents DNA methyltransferase from depositing and methylating DNA in enhancer regions.

[0064] CHD7 The gene encodes chromatin helicase, an ATP-dependent chromatin remodeling factor. CHD7 physically interacts with DNA methyltransferase 3B (DNMT3B), assisting DNMT3B in establishing DNA methylation at specific genomic sites during neurogenesis by remodeling chromatin.

[0065] Copy number variation in the 7q11.23 region occurs through trans-action ( trans -acting) affects DNA methylation. Transcription factor genes in this region ( GTF2I / GTF2IRD1 Dosage changes can indirectly and genome-wide affect the DNA methylation levels of other genes.

[0066] Methylation abnormalities in the 15q11.2-13.1 region occur directly within this region itself. This region is a typical imprinted gene regulatory region, and its methylation status directly determines the expression of imprinted genes within the region.

[0067] Therefore, the aforementioned gene or genomic variations will cause changes in the methylation level of genomic DNA.

[0068] 2. Methylation detection data acquisition 2.1 Sample Collection and Quality Inspection Using Illumina's Infinium TM The Methylation V2.0 (935K) chip was used to detect methylation in peripheral blood genomic DNA samples from "training cohort patients" and age- and sex-matched healthy controls. The main process included: extracting peripheral blood genomic DNA from patients and healthy controls using a kit (Tiangen Biotech (Beijing) Co., Ltd.; catalog number DP348), purifying and quality-checking the DNA (OD260 / 280=1.8-2.0, DNA concentration ≥50ng / μL) for later use.

[0069] 2.2 Methylation chip detection Bisulfite transformation was performed (Illumina Infinium MethylationEPIC v2.0 Kit, catalog number 20087708); the transformed DNA was then subjected to microarray detection. The overall process included amplification-fracture-precipitation-resuspending-hybridization-extension-washing-scanning, and methylation microarray scanning was performed using iScan (iScan scanner parameters: pixel resolution 0.53 μm, 16-bit acquisition, 532 nm and 658 nm dual laser excitation wavelengths, scan throughput approximately 30 minutes / chip, chip density ≥4 million dots / cm²), to obtain the raw detection data in the idat file.

[0070] The meffil (v1.4.0) package in R (v4.4.1) was used to read the idat raw data file and obtain the whole genome methylation β value matrix data (hereinafter referred to as β value matrix) of each sample, which is the methylation detection data.

[0071] 3. Standardization of test data 3.1 Probe quality filtering and data normalization Based on the preset probe quality control parameters shown in Table 2, the original methylated probes corresponding to the generated β-value matrix are quality filtered, retaining probes that meet the quality control requirements; and the β-value matrix corresponding to the retained probes after filtering is normalized. The normalization uses the meffil package (v1.4.0) quantile normalization method, sets the number of principal components to 10, and corrects for chip batch factors Slide and Array as random effects.

[0072] Table 2. Preset probe quality control parameters 3.2 Probe missing value completion For whole-genome methylation β matrix data of CpG sites that still have missing values ​​after normalization, a method based on Principal Component Analysis (PCA) was used for data reconstruction and completion. Specifically, firstly, non-finite values ​​in the whole-genome methylation β matrix data were defined as missing values. Based on the training set, the missing proportion of each CpG probe was statistically analyzed, and probes with a missing proportion greater than 15% were removed, retaining only probes with a missing proportion not exceeding 15% for subsequent analysis. For the missing values ​​still present in the retained probes, a PCA-based missing value reconstruction model was constructed. The model parameters were set to 50 principal components, and centering and normalization were applied. For residual missing values ​​in the training set, initial imputation was performed using the mean of the corresponding probe in the training set. Then, iterative low-rank reconstruction was performed based on the PCA loading matrix fitted from the training set, with 5 iterations. After reconstruction, the results were inversely transformed back to the original β value space, and the completed β values ​​were restricted to the interval between 0 and 1 to obtain the completed β value matrix data.

[0073] 3.3 Sample quality assessment Furthermore, sample quality is assessed based on preset sample quality control parameters. These parameters include: a gender anomaly threshold of 6 standard deviations, a methylation / non-methylation signal intensity anomaly threshold of 4 standard deviations, and a non-compliant probe proportion threshold of 0.1 based on the detection P-value. The sample quality is comprehensively assessed by considering the sample's detection signal intensity, data integrity, consistency, and anomaly distribution, with samples failing quality control in the training set being removed. Finally, the sources of data variation are evaluated and corrected based on factors such as sample batch, age, gender, and cell composition, generating a genome-wide methylation analysis matrix (hereinafter referred to as analysis matrix data) for subsequent statistical analysis and model construction.

[0074] 4. DMP and DMR Analysis After completing the basic quality control described above, the meffil (v1.4.0) package in R (v4.4.1) was used to calculate the set of differentially methylated positions (DMPs) for the disease and controls based on the analysis matrix data of the training set. Then, based on the statistical results of the differentially methylated positions, DMRcate (v3.2.1) was used to perform differentially methylated regions (DMRs). DMR analysis was used to identify regions where multiple adjacent CpG sites exhibited consistent methylation changes in the same direction. For each target disease, the set of probes constituting the DMP represents the complete set of candidate probes for that disease. Furthermore, candidate probes from DMP sources were further screened (screening process described below). For the probes obtained from the final screening, methylation analysis matrix data corresponding to the disease group and control group in the training cohort were extracted and presented as a heatmap. Figure 3 The distribution of methylation levels at the corresponding sites of the probes in the two groups of samples is shown. The results indicate that there are different degrees of characteristic differences in the methylation levels of the screened probes between the disease group and the control group.

[0075] 5. Probe-based initial screening 5.1 Significance Screening In the actual analysis process, the probes are first screened for significance based on the adjusted P-value after multiple corrections in the DMP analysis. The initial screening threshold is adjusted P-value < 0.05. If the number of probes retained after initial screening for any disease is less than the preset value of 300, the adjusted P-value threshold is relaxed to 0.1 to ensure that there are a sufficient number of probes for subsequent analysis.

[0076] 5.2 Absolute methylation difference effect intensity (|Δβ| value) filtering Based on the significance screening, further filtering is performed using the absolute methylation difference effect strength of the probes (|Δβ| value), which is derived from the difference in methylation intensity between each disease and the control. The threshold screening for |Δβ| value uses a combination of a fixed effect strength threshold (FEST) and a distribution adaptive threshold (DAT) to determine the actual effect size screening threshold (AESST) (i.e., |Δβ| value ≥ AESST). In this scheme, FEST is set between 0.01 and 0.20, and the 25th percentile of the current distribution of significant probe effect strength for each disease is calculated as the adaptive reference threshold (AT). When the number of significant probes for a disease is sufficiently large, the larger of FEST and DAT is used as the AESST for that disease; otherwise, AT is used directly for screening.

[0077] 5.3 Upper and lower limits of the initial screening probe set To avoid excessive differences in the number of initial screening probe sets after primary screening for different diseases, and to balance screening stability and computational efficiency, this protocol further sets a lower and upper limit on the number of initial screening probes retained after primary screening. The lower limit is set at 300, and the upper limit is set at 3000. When the number of initial screening probes after dual filtering by significance and effect strength is lower than the lower limit, the retention is first relaxed according to the fixed effect strength threshold (FEST, 0.01-0.20). If it is still insufficient, it is supplemented to the minimum number of candidates in the order of significance first, followed by effect strength. If the number of initial screening probes exceeds the upper limit, the top-ranked initial screening probes are selected for subsequent analysis in the same priority order.

[0078] 5.4 Primary screening parameters and the number of primary screening probes Table 3. Overview of Initial Screening Parameters Based on the screening rules set in 5.1-5.3, since the degree of methylation change varies in different diseases, the main change in application is to modify the effect strength threshold (see Table 3 for parameter overview), while keeping the other parameters unchanged. Under this setting, each disease has 20 sets of initial screening probes, which are the initial screening probe sets generated at a threshold of 0.01-0.20 (step size of 0.01), which are used to screen or combine the best methylation-specific probe sets for each disease.

[0079] 6. Specific probe screening 6.1 Definition of Disease Relationship After obtaining the probe sets that pass the initial screening for each disease, a further assessment of the similarity and differentiation difficulty between diseases is constructed. Specifically, based on the consistency of probe effect patterns among different diseases, the pairwise similarity between diseases is calculated, and the group of diseases that are most difficult to distinguish from the target disease and the group of diseases that are relatively easy to distinguish are identified accordingly. In this scheme, disease similarity is assessed primarily based on the correlation of probe effect vectors between diseases, and the top 5 diseases most similar to the target disease are selected as the difficult-to-distinguish diseases. For disease combinations that have existing clinical or molecular biological evidence indicating similarity, a priori similarity constraint can be imposed to ensure that the differentiation difficulty is not less than 0.70 (i.e., the correlation coefficient is set to a minimum of 0.7).

[0080] 6.2 Screening strategy for specific probes After the disease relationships are established, a comprehensive ranking score is calculated for each initially screened probe. This score comprehensively reflects the supporting evidence for the probe in the target disease and its overlap penalty in non-target diseases, and serves as the core basis for subsequent ranking and screening of specific probes. Combined with subsequent threshold filtering, salvage, and pruning rules, specific probes corresponding to each disease are further screened.

[0081] (1) Calculation of comprehensive ranking score After the initial screening probes enter the comprehensive ranking step, a basic discovery score is first constructed using the significance p-value and Δβ, which simultaneously reflects the significance and effect strength of the probe. For probes Its basic discovery score Defined as: in, Indicates probe The significance level after multiple adjustment in the target disease. Indicates probe In target disease The methylation differential effect value in This represents the normalization function.

[0082] In practice, the normalization method adopted is minimum-maximum normalization, which is calculated as follows: To evaluate the specificity of the probe relative to non-target diseases, the specificity ratio and specificity difference are further defined. Probe Specificity ratio Defined as: probe Specificity difference Defined as: To measure the recurrence of probes in non-target diseases, a global overlap score and a difficult-to-distinguish disease overlap score are further defined. Probe Global Overlap Score Defined as: set up Indicates the relationship with the target disease The most difficult set of diseases to distinguish is the probe. Difficult-to-distinguish disease overlap scores Defined as: Furthermore, if a probe simultaneously satisfies both the preset significance condition and the preset effect strength condition in a non-target disease, it is recorded as a hit of the probe in that non-target disease. Let the probe... The hit count across all non-target diseases was The total number of non-target diseases is Then its hit rate Defined as: Furthermore, to evaluate the discriminative ability of a single probe against the target disease, this method calculates the area under the receiver operating characteristic (AUC) curve of the probe under different comparison scenarios. Specifically, the target disease sample is designated as the positive class, and the corresponding comparison group sample is designated as the negative class. The methylation value of the probe in the corresponding sample is used as the continuous discrimination score to calculate the corresponding AUC. The comparison group may include normal controls, all other diseases, a set of difficult-to-distinguish diseases, or a set of easily distinguishable diseases.

[0083] in, Indicates probe The area under the original curve calculated under a given comparison scenario The following discriminant ability indicators are further defined: Indicates probe Discriminant ability in comparison between target disease and normal controls; Indicates probe The ability to distinguish between the target disease and all other diseases; Indicates probe Discriminative ability in comparing the target disease with a set of indistinguishable diseases; Indicates probe Discriminative ability in comparing the target disease with a set of easily distinguishable diseases.

[0084] Furthermore, to evaluate the stability of the probe in the most unfavorable comparison of closely related diseases, the minimum AUC obtained when comparing the target disease with each difficult-to-distinguish disease is defined as: in, Indicates the relationship with the target disease The most difficult set of diseases to distinguish Indicates probe In target disease A disease that is difficult to distinguish The AUC obtained when comparing pairs of pairs has the same direction.

[0085] In the calculation of the overall ranking score, the aforementioned discriminative ability indicators are included in the positive support score. (Probe) The positive support score can be represented as: in, Indicates the basic discovery score. Indicates probe In target disease The absolute effect intensity in Indicates the specificity ratio. Indicates the specificity difference. Represents the normalization function. This represents the weighting coefficient of each component, with values ​​between 0 and 1, and no fixed rule. The sum is set to 1, and the rest are assigned values ​​of 0.25, 0.15 and 0.1 based on general experience.

[0086] The overlap score is jointly incorporated into the overlap penalty score, including indistinguishable disease overlap scores, global overlap scores, and hit rate. (Probe) Overlapping penalty score Defined as: in, Indicates difficulty in distinguishing overlapping disease scores. Indicates global overlap score. The percentage of hits is included in the overlap penalty score. This represents the weighting coefficient of each component, with values ​​between 0 and 1. There are no fixed rules, but values ​​of 0.25, 0.15, and 0.1 are generally given based on experience. The main purpose of the penalty term is to reduce the ranking score of probes that only have a single discriminative ability (such as only being able to distinguish between the case group and the control group), rather than to completely exclude them.

[0087] Therefore, for each initial screening probe, the overall ranking score consists of a positive support score and an overlap penalty score, and its calculation principle is as follows: in, This indicates the overall ranking score of the initial screening probes. Indicates positive support score, This indicates the overlapping penalty score.

[0088] The positive support score reflects the probe's support for the target disease, including significance, effect strength, case-control differentiation ability, ability of cases to differentiate other diseases, ability to distinguish difficult-to-distinguish diseases, and specific advantages of the probe for the target disease relative to other diseases. The overlap penalty score reflects the degree of repetition of the probe in non-target diseases, especially overlap in difficult-to-distinguish diseases. Finally, the overall ranking score of the probe is obtained by subtracting the overlap penalty score from the positive support score.

[0089] In this scheme, the positive support score is specifically represented as: Correspondingly, the overlap penalty score is specifically represented as follows: Therefore, a higher overall ranking score indicates that the probe is more likely to simultaneously meet the following characteristics: first, it has high significance and a large effect strength in the target disease; second, it has strong discriminative ability in multiple scenarios, such as case-control, case-to-other diseases, and case-to-difficult-to-distinguish diseases; and third, it has low overlap in other diseases, especially in difficult-to-distinguish diseases. Probe specificity screening and ranking are based on this overall ranking score from high to low, and additional conditions are combined to gradually form the final set of specific probes.

[0090] (2) Conservative discrimination criteria and threshold filtering based on comprehensive ranking scores After calculating the comprehensive ranking score, the comprehensive ranking score is used as the core basis for screening specific probes, and the specific probes are further filtered and ranked.

[0091] First, to prevent specific probes that perform well only in a single comparative scenario from being included in subsequent steps, a conservative discrimination criterion is introduced. This conservative discrimination criterion comprehensively considers the case-control differentiation ability, the case's differentiation ability against all non-target diseases, and the weakest differentiation ability against difficult-to-distinguish diseases, and is set as the minimum of these three factors.

[0092] The conservative criterion is: in, This indicates the conservative discrimination threshold for the probe. This indicates the preset minimum target disease-control discrimination threshold. This represents the minimum threshold for distinguishing between target and non-target diseases. This indicates the lowest conservative discrimination threshold for distinguishing between diseases that are difficult to differentiate. Only probes that simultaneously meet the preset minimum target disease-control discrimination ability, minimum target disease-non-target disease discrimination ability, and minimum conservative discrimination threshold (only probes that meet the requirements of conservative discrimination index ≥ preset conservative comprehensive discrimination threshold) will be retained for subsequent steps.

[0093] In this scheme, the minimum thresholds for case-control discrimination ability and case discrimination ability against other diseases are both set to 0.70, and the conservative comprehensive discrimination threshold is set to 0.85 according to different implementations.

[0094] (3) Remediation when the number of specific probes is insufficient, correlation pruning and generation of disease-specific probe sets ① Remedial measures when the number of specific probes is insufficient If the number of specific probes retained after sorting, screening, conservative discrimination criteria, and threshold filtering for a particular disease is insufficient (the lower limit of the specific probe set is set to 100), this scheme further employs an adaptive salvage strategy. Specifically, using a fixed effect strength threshold (FEST, 0.01-0.20) as the relaxation target, the minimum effect requirement for entering the specific probe set is gradually reduced in a fixed step size, and the sorting results are re-merged after each relaxation until the minimum number of specific probes is reached or the maximum number of relaxations is reached. In this scheme, the step size for each relaxation of the effect strength threshold is set to 0.005, and the maximum number of relaxations is set to 10. If the minimum number is still insufficient after all relaxations are completed, the minimum number of probes is directly supplemented according to the comprehensive sorting results.

[0095] ②Relevance pruning If the number of probes in the specific probe set is too large (the upper limit of the specific probe set is set to 300), to further reduce redundancy and improve the stability of the final specific probe combination, probes with high correlation are pruned based on the correlation between specific probes. Specifically, specific probes for each disease are retained in descending order of their comprehensive ranking scores. If the correlation between a specific probe and the retained specific probes exceeds a preset threshold, the specific probe is identified as a redundant probe and removed. In this scheme, the correlation is measured using the Pearson correlation coefficient, and the initial correlation threshold is set to 0.85. If the final number of specific probes after pruning is lower than the minimum requirement, attempts are made to supplement from the remaining high-scoring probes at the same correlation threshold; if this is still insufficient, the correlation threshold is gradually relaxed, with a relaxation step size of 0.02 and a maximum relaxation number of 5 times, until the minimum probe number requirement or the allowable upper limit is reached.

[0096] ③ Upper and lower limits of specific probe sets Finally, a specific probe set is formed for each disease after significance screening, effect strength screening, comprehensive ranking, threshold filtering, adaptive salvage, and relevance pruning. In this scheme, the lower limit for the number of probes retained for each disease is set to 100, and the upper limit is set to 300. Different embodiments can set specific effect strength thresholds, conservative discrimination thresholds, and other screening parameters for different diseases within the above framework to obtain the optimal probe combination suitable for the disease identification task.

[0097] (4) Specific screening parameters Table 4. Overview of Specific Screening Parameters 6.3 Epi-signature specific probe set (1) Construction of methylation characteristic spectrum of Sotos syndrome With other screening parameters remaining unchanged, a set of probes with an effect intensity threshold set to 0.10 was selected, and 150 specific probes were finally obtained. Table 5. Set of 150 “Epi-signature” probe sites for Sotos syndrome (2) Construction of methylation characteristic spectrum of Kabuki syndrome With other screening parameters remaining unchanged, five probe sets with effect intensity thresholds set to 0.05–0.10 were selected, and the intersection of the obtained probes was taken to finally screen out 145 specific probes.

[0098] Table 6. Set of "Epi-signature" probe sites for Kabuki syndrome (145 sites) (3) Construction of methylation characteristic spectrum of CHARGE syndrome With other screening parameters remaining unchanged, 10 sets of probes with effect intensity thresholds set to 0.01 to 0.10 were selected, and the intersection of the obtained probes was taken to finally screen out 158 ​​specific probes.

[0099] Table 7. Epi-signature probe site set for CHARGE syndrome (158 sites) (4) Construction of methylation characteristic spectrum of WSS syndrome With other screening parameters remaining unchanged, a set of probes with an effect intensity threshold set to 0.10 was selected, and 150 specific probes were finally obtained.

[0100] Table 8. Epi-signature probe site set for Williams–Steiner syndrome (150 sites) (5) Construction of methylation characteristic spectrum of WBS syndrome With other screening parameters remaining unchanged, a probe set with an effect strength threshold set of 0.10 was selected, resulting in 172 specific probes in the first step. Furthermore, given that WBS is a 7q11.23 deletion syndrome, it may exhibit opposite methylation changes to 7q11.23 duplication syndromes in certain specific methylation regions. Therefore, 164 probes with opposite methylation directions reported in the literature were introduced, and their intersection with the DMP set (WBS vs Control) obtained from the training set of this project was screened, ultimately supplementing the results with 60 more specific probes. In summary, a total of 232 specific probes were obtained.

[0101] Table 9. Epi-signature probe site set for Williams-Beuren syndrome (232 sites) (6) Construction of methylation characteristic spectrum of PWS syndrome Given that PWS is a syndrome associated with abnormal imprinting in the 15q11.2–q13 region, and its main molecular basis is the deletion or inactivation of paternally expressed genes in this region, this protocol focuses on probe screening targeting this region. Based on literature reports, we will focus on… SNRPN, MAGEL2, NDN, MKRN3, SNORD115 and SNORD116 Key gene or snoRNA cluster-related sites were identified, and 468 relevant probes supported by literature were compiled. Further, the intersection of these 468 probes with the DMP set (PWS vs Control) obtained from the training set of this project was screened, and then filtered under a fixed threshold |Δβ|≥0.01, ultimately yielding 219 specific probes.

[0102] Table 10. Epi-signature probe site set for Prader-Willi syndrome (219 sites) (7) “Epi-signature” probe set Regarding probe set screening strategies, Sotos, Kabuki, CHARGE, WSS, and WBS all screened from the entire genome, while PWS only screened differentially methylated sites in the 15q11.2-13.1 region. Based on the "training cohort," a total of 1052 probe sites suitable for subsequent machine learning were screened across the six disease groups (Tables 5-10), of which 99.8% of the probes were disease-specific, i.e., from a single disease (Table 11).

[0103] Table 11. Specificity of "Epi-signature" probes for 6 diseases In summary, based on the methylation chip detection results of the "test cohort" samples, β-value matrix data was first obtained. After correction for confounding factors, it was transformed into analytical matrix data. Then, DMP and DMR analyses were used to obtain the complete set of candidate probes. Following initial screening and specificity screening, specific probe sets for various diseases were obtained, and specific methylation characteristic spectra for various diseases were constructed (see...). Figure 1 ).

[0104] Example 2. Constructing a classification and diagnostic prediction model for NDD based on machine learning 1. Construction of diagnostic prediction models for various disease categories The project team initially used different machine learning strategies to build classification and diagnostic prediction models, including Logistic Regression, Random Forest, Decision Tree, and Support Vector Machine (SVM). They found that the SVM model had the best diagnostic performance. SVM is a type of generalized linear classifier that performs binary classification of data using supervised learning. Even for rare diseases with a small number of cases (as low as 5 in extreme cases), it can make effective predictions by analyzing differences in methylation features compared to controls.

[0105] Based on the "training queue" of 6 disease groups in Example 1 and 1 healthy control group, this project constructs a "one-vs-all" binary classification model for each disease. That is, the target disease samples are defined as positive classes, and the healthy controls and other non-target disease samples are defined as negative classes. A consistent modeling process and parameter settings are adopted. Based on the disease-specific probe set obtained by screening the corresponding disease (i.e., the disease-specific probe sets shown in Tables 5-10 of Example 1), the matrix data in the corresponding methylation analysis matrix is ​​extracted, which is called disease methylation feature spectrum matrix data (referred to as feature spectrum matrix data). Linear kernel support vector machine is used for training.

[0106] During training, the SVM diagnostic prediction models for each disease optimize the SVM penalty parameter C through grid search. Specifically, a series of candidate C values ​​are pre-defined, and the candidate range of the penalty parameter C is set on a logarithmic scale to cover... to The order of magnitude. Subsequently, the model performance corresponding to each parameter combination is trained and evaluated under the cross-validation framework. The cross-validation preferably adopts a repeated stratified cross-validation method. The number of cross-validation folds is adaptively determined based on the number of positive and negative samples, and does not exceed 10 folds. The preferred number of repetitions is 5, and the random seed is 20260309. Further, the C value that optimizes the F1 score is selected as the final parameter to improve the model's generalization ability and classification performance. The main difference between the models lies in the different optimal values ​​of the penalty parameter C. The optimal C values ​​for WSS, Sotos, Kabuki, WBS, CHARGE, and PWS are respectively... 10, 10, 100 And 1.

[0107] The SVM diagnostic prediction models for each disease also employ repeated stratified cross-validation based on binary labels for parameter selection and performance evaluation (using the F1 score as the optimal parameter selection metric) to maintain a relatively consistent ratio between target disease samples and non-target samples. Furthermore, to mitigate the impact of class imbalance on model training, target disease samples, healthy control samples, and other non-target disease samples are assigned different classification attributes. , and The empirical weights set by logarithm were used, and by imposing higher penalties on non-target disease samples that are difficult to distinguish, the ability of the SVM diagnostic prediction model to distinguish between target diseases and non-target samples was improved.

[0108] After obtaining the basic support vector machine (SVM) classification and diagnostic prediction models for each disease, the Platt scaling method (Sigmoid calibration) is used to probabilistically calibrate the outputs of each disease's basic SVM classification and diagnostic prediction model. This involves fitting a Sigmoid function to the original decision values ​​of the SVM output and mapping them to predicted probabilities between 0 and 1, thus obtaining the disease-specific SVM diagnostic prediction model. For each test sample, its predicted probability under the disease-specific SVM diagnostic prediction model is calculated, where the predicted probability for each disease is defined as the methylation variant pathogenicity (MVP score). Furthermore, based on clinical experience, a preset threshold can be used to determine the MVP score for each disease, converting it to positive (MVP score ≥ 0.5), negative (MVP score < 0.1), or indeterminate (0.1 ≤ MVP score < 0.5), thereby achieving multi-disease classification and diagnosis of the test samples. The above model training, probability calibration, threshold optimization, and prediction process is preferably implemented using Python and the scikit-learn machine learning library, ultimately obtaining the methylation-specific classification and diagnostic prediction model for each disease. Based on this, the performance of the SVM diagnostic prediction model for each disease was evaluated using the out-of-bounds prediction probability obtained by cross-validation. Indicators such as accuracy, precision, recall, F1 score, and area under the ROC curve (AUC) were calculated. The results are shown in Table 11.

[0109] Table 11. Performance Indicators of SVM Classification Diagnostic Prediction Models for Various Diseases Note: TP true positive count; TN true negative count; FP false positive count; P positive sample count; N negative sample count Accuracy = (TP + TN) / (P + N), representing the percentage of samples that are classified as true positives or true negatives. It is an evaluation metric for the classifier's performance on the overall data.

[0110] Precision = TP / (TP+FP) Precision is an evaluation metric for a classifier on data that is predicted as positive.

[0111] Recall = TPR = sensitivity = TP / (TP+FN), and recall / sensitivity is an evaluation metric for the classifier on the entire positive data.

[0112] F1-score / F1-measure=2*(recall*precision / (recall+precision)) is a reconciliation of two contradictory indicators, precision and recall.

[0113] ROC_AUC is used to judge the quality of a model and represents the area under the curve in the ROC algorithm. The value of AUC is generally between 0.5 and 1.

[0114] 2. Performance testing of SVM diagnostic prediction models for various diseases like Figure 2 As shown, the predictive analysis of the SVM diagnostic prediction models for the six diseases that have been constructed is performed respectively, and the MVP score of each disease is obtained. The highest score among the six results is selected as the final score MaxScore of the sample. The matching diagnostic prediction result is determined according to MaxScore>0.5. The multi-disease classification diagnosis of the sample to be tested is realized by combining the judgment results of all diseases.

[0115] 3. Validate the clinical effectiveness of SVM diagnostic prediction models for each disease based on a "test cohort": The six models trained above were used to test 89 samples. The inclusion criteria were: ① cases with consistent clinical phenotypes and clear pathogenicity of variants; ② or cases with basically consistent clinical phenotypes but unclear pathogenic significance of variants (VUS); or ③ cases with atypical clinical features but strong pathogenicity of variants (e.g., KMT2D Classical splicing variants without clinical abnormalities); exclusion criteria are: no relevant gene or genomic variants related to the declared disease were found in the patient after high-throughput sequencing.

[0116] The Platt scaling method was used to convert the SVM decision values ​​of different diagnostic prediction models into methylation variant pathogenicity (MVP) scores; positive samples were defined as MVP ≥ 0.5, 0.1 ≤ MVP < 0.5 was the gray range, and < 0.1 was negative. The test results showed that only two cases had ambiguous classification in the model, and the rest were correctly predicted. The results are shown in Table 12.

[0117] Table 12. Clinical efficacy of the SVM diagnostic prediction models for various diseases disclosed in this paper. Note: TP (True Positives); TN (True Negatives); FP (False Positives); FN (False Negatives); UN (Number of Uncategorically Identifiable Samples). 3.1 Test Results of WSS's SVM Diagnostic Prediction Model Of the 89 tested samples, there were 3 confirmed cases of WSS and 1 case of clinically atypical but pathogenic variant (WSS_23). The clinical features of the WSS_23 case (male) at 6 months of age during gene sequencing were minimal refusal to eat complementary foods, slow growth, excessive hair growth, and long eyelashes; gene testing showed that the infant possessed… KMT2A The gene c.11429+1G>T heterozygous variant, paradoxically inherited from an asymptomatic mother, was initially classified as VUS (Very Uncertain of Clinical Significance). The methylation diagnostic prediction model identified it as a negative sample. Follow-up at 2 years of age showed only mild developmental delay and no WSS characteristics. Further analysis revealed that the c.11429+1G>T variant is located at the C-terminus of the KMT2A protein, and the resulting amino acid changes likely have minimal impact on protein function. Therefore, the c.11429+1G>T variant was reclassified as a possibly benign variant, ruling out a WSS diagnosis. Thus, the WSS methylation diagnostic prediction model achieved 100% classification accuracy; these results demonstrate that the diagnostic prediction model can differentiate WSS from other diseases.

[0118] 3.2 Test Results of Sotos' SVM Diagnostic Prediction Model Of the 89 test samples, 10 were... NSD1 Of the gene mutation samples, 4 were clinically negative and 6 were clinically positive, with a classification accuracy of 100%. These results indicate that the diagnostic prediction model can achieve differential diagnosis of Sotos and other diseases.

[0119] 3.3 Test Results of the SVM Diagnostic Prediction Model for Type I Kabuki Of the 89 test samples, 17 were... KMT2D Two patients with type I Kabuki positivity due to gene mutation. KMT2D The study included samples of variants with uncertain clinical significance (VUS) and one chimeric mutation patient (29% chimerism rate). Results showed 17 positive cases with 100% classification accuracy. Two VUS cases had clinical characteristics inconsistent with Kabuki, carrying c.16018C>T, p.R5340* and c.4020+2T>A respectively. The former presented with abnormal blood glucose, while the latter was detected in an asymptomatic population during routine checkups. Methylation models classified both variants as negative. Further analysis revealed that the c.16018C>T, p.R5340* variant is located closer to the C-terminus of the protein, potentially causing less functional impairment to KMT2D. The c.4020+2T>A variant was predicted to result in a 38-amino acid deletion, also suggesting minimal functional impairment (KMT2D has a large molecular weight). The chimeric mutation patient was a 9-month-old boy. KMT2D(c.7340del, p.Pro2447Leufs*38), clinically a typical Kabuki syndrome, with an MVP value of 0.15 for its methylation model, which is in the gray range and consistent with the genomic chimerism rate of 29%; the above results indicate that the diagnostic prediction model can achieve differential diagnosis between Kabuki and other diseases.

[0120] 3.4 Test Results of the SVM Diagnostic Prediction Model of WBS Of the 89 test samples, one was negative (7q11.23 del (hg19, chr7:72858312-73483725, excluding ELN) and 12 were positive for WBS, with a classification accuracy of 100%. However, one sample was also incorrectly identified as possibly having CHARGE syndrome (MVP score 0.175). These results indicate that the diagnostic prediction model can differentiate WBS from other diseases.

[0121] 3.5 Test Results of CHARGE's SVM Diagnostic Prediction Model Of the 89 tested samples, one was CHD7 mutation negative and one was negative. CHD7 The study included a mutated Kaman syndrome (CHARGE_9268) and 6 CHARGE-positive samples; the 6 positive and 2 negative samples were accurately classified, achieving a diagnostic accuracy of 100%. These results indicate that the diagnostic prediction model can not only effectively identify CHARGE-positive and negative samples, but also provide clues. CHD7 Different disease subtypes caused by mutations.

[0122] 3.6 Test Results of PWS's SVM Diagnostic Prediction Model Of the 89 test samples, 7 were clinically diagnosed PWS samples; 6 of them were accurately classified, but 1 sample was in the gray range of the MVP value (PWS_8391: 0.436).

[0123] Example 3 like Figure 11 As shown, a classification, diagnostic, and prediction system for congenital neurodevelopmental disorders includes: Module 201 is used to construct the diagnostic prediction model for the aforementioned congenital neurodevelopmental disorders; The acquisition module 202 is used to acquire the detection data after sample methylation detection, and to standardize the detection data to obtain the whole genome methylation analysis matrix of the sample, and to screen out the methylation feature matrix of each disease. The model processing module 203 is used to input the disease-specific methylation feature matrix of each disease into the corresponding congenital neurodevelopmental disorder diagnosis and prediction model, and calculate the pathogenicity score of methylation variation for each disease. Evaluation module 204 is used to classify and diagnose congenital neurodevelopmental disorders based on the pathogenicity score of methylation variants corresponding to each disease.

[0124] It should be noted that the technical essence of Embodiment 3 is the same as that of Embodiments 1-2. If there is anything unclear, please refer to Embodiments 1-2.

[0125] The functions of the apparatus in this embodiment have been described in the above method embodiments. Therefore, for any parts not detailed in this embodiment, please refer to the relevant descriptions in the foregoing embodiments, which will not be repeated here.

[0126] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0127] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0128] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0129] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0130] The above specific embodiments are merely illustrative of the content of this disclosure and do not represent a limitation thereof. For those skilled in the art, various modifications and improvements can be made without departing from the spirit and substance of this disclosure, and these modifications and improvements are also considered to be within the scope of protection of this disclosure.

Claims

1. A method for classifying, diagnosing, and predicting congenital neurodevelopmental disorders, characterized in that, Includes the following steps: Methylation chip detection and data standardization were performed on peripheral blood samples that had passed collection and quality inspection to obtain a whole-genome methylation analysis matrix of the samples. From the whole genome methylation analysis matrix, methylation-specific probe sets for various congenital neurodevelopmental disorders are screened, and the corresponding methylation signals are extracted to obtain the methylation feature matrix of each disease; Each disease-specific methylation feature matrix is ​​input into the corresponding congenital neurodevelopmental disorder diagnosis and prediction model to calculate the pathogenicity score of methylation variation for each disease. Based on the pathogenicity score of methylation variation corresponding to each disease, the samples were classified and diagnosed as congenital neurodevelopmental disorders. The congenital neurodevelopmental disorder is selected from at least one of Sotos syndrome, Kabuki syndrome type I, CHARGE syndrome, Wiedemann-Steiner syndrome, Williams-Beuren syndrome, and Prader-Willi syndrome.

2. The method according to claim 1, characterized in that, The methylation chip detection includes: using Illumina Infinium TM The Methylation V2.0 (935K) chip was used to detect peripheral blood genomic DNA transformed with Bisulfite to obtain raw idat detection data. The raw detection data was read using the R language meffil package to obtain a whole-genome methylation β-value matrix.

3. The method according to claim 2, characterized in that, The standardization processing of the detection data includes: After quality filtering of the whole genome methylation β value matrix, quantile normalization and chip batch factor correction were performed. Non-finite values ​​in the whole genome methylation β value matrix are defined as missing values. Probes with a missing ratio >15% are removed. The remaining missing values ​​are iteratively reconstructed and completed using a PCA model with 50 principal components. The completed β values ​​are restricted to the range of 0 to 1. After sample quality control and correction for variations in batch, age, sex, and cell composition, the whole genome methylation analysis matrix was obtained.

4. The method according to claim 1, characterized in that, The step of screening a set of methylation-specific probes for various congenital neurodevelopmental disorders from the whole-genome methylation analysis matrix includes: Differential methylation sites between various congenital neurodevelopmental disorders and healthy controls were calculated from the genome-wide methylation analysis matrix to obtain a complete set of candidate probes for each disease; From the entire candidate probe set for each disease, a methylation-specific probe set for each disease was selected. Specifically: (1) For Sotos syndrome, Kabuki syndrome type I, CHARGE syndrome or Wiedemann-Steiner syndrome, the candidate probe set is subjected to preliminary screening and specific screening in sequence to obtain the methylation specific probe set for each disease; (2) For Williams-Beuren syndrome, after the candidate probe set is screened for initial screening and specific screening in sequence, the candidate probe set for the disease is screened for intersection with the probe sites related to the disease reported in the literature. The probes obtained from specific screening and the probes obtained from intersection screening together constitute the methylation specific probe set for the disease. (3) For Prader-Willi syndrome, the candidate probe set for the disease was screened by intersection with the probe sites related to the disease reported in the literature, and filtered under a fixed threshold |Δβ|≥0.01 to obtain the methylation-specific probe set for the disease.

5. The method according to claim 4, characterized in that, The initial screening includes: Significance screening was performed using an adjusted P-value < 0.05, and effect strength filtering was performed using the |Δβ| value determined by a fixed threshold and a distribution adaptive threshold. The initial screening probe count has a lower limit of 300 and an upper limit of 3000. If the number is insufficient, the threshold will be relaxed to make up the difference. If the number is excessive, the probes ranked higher will be selected.

6. The method according to claim 4, characterized in that, The specific screening includes: Based on the similarity of probe effect patterns among diseases, identify the set of indistinguishable diseases for the target disease; For each initial screening probe, a comprehensive ranking score is calculated. The comprehensive ranking score consists of a positive support score and an overlap penalty score, and its calculation principle is as follows: in, This represents the overall ranking score of the candidate probes. This represents the positive support score, which includes the basic discovery score and the AUC value for each scenario. This indicates the overlap penalty score, which includes overlap ratings and hit rate. according to Perform conservative filtering and retain only Probes with a strength ≥0.85, wherein, This indicates the conservative discrimination threshold for the probe. This indicates the preset minimum target disease-control discrimination threshold. This represents the minimum threshold for distinguishing between target and non-target diseases. This indicates the lowest conservative discrimination threshold for distinguishing between diseases that are difficult to differentiate. The number of specific probes is limited to a minimum of 100 and an upper limit of 300. If the number is insufficient, the effect strength threshold is relaxed to compensate. If the number is excessive, correlation pruning is performed based on the Pearson correlation coefficient.

7. The method according to claim 4, characterized in that, The set of methylation-specific probes for Sotos syndrome is as follows: , A total of 150 specific probes; the set of methylation-specific probes for Kabuki syndrome type I is as follows: , A total of 145 specific probes; the methylation-specific probe set for CHARGE syndrome is as follows: , A total of 158 specific probes; the set of methylation-specific probes for Wiedemann-Steiner syndrome is as follows: , A total of 150 specific probes; the set of methylation-specific probes for Williams-Beuren syndrome is as follows: , A total of 232 specific probes; the set of methylation-specific probes for Prader-Willi syndrome is as follows: , A total of 219 specific probes.

8. The method according to any one of claims 1-7, characterized in that, The diagnostic prediction model for congenital neurodevelopmental disorders is constructed through the following steps: (1) For each target congenital neurodevelopmental disorder, a "one-vs-all" binary classification model is constructed, with the disease sample set as the positive class and healthy controls and other non-target disease samples set as the negative class; (2) Using the feature spectrum matrix corresponding to the methylation-specific probe set of each disease as input, a linear kernel support vector machine (SVM) is used for model training. The model is trained by grid search on a logarithmic scale of 10. -5 Up to 10 4 The penalty parameter C was optimized within the range, and the model was evaluated using repeated stratified cross-validation. The optimal penalty parameter C for each disease was determined based on the optimal F1 score. (3) In model training, preset empirical weights are assigned to target disease samples, healthy control samples and non-target disease samples respectively to alleviate the class imbalance problem; the Platt scaling method is used to calibrate the output of the basic SVM to a prediction probability of 0-1, which is defined as the methylation variant pathogenicity score (MVP score).

9. The method according to claim 1, characterized in that, The method of classifying and diagnosing congenital neurodevelopmental disorders based on the pathogenicity score of methylation variants corresponding to each disease includes: determining the pathogenicity score of methylation variants (MVP score) according to a preset threshold; an MVP score ≥ 0.5 is considered positive for the disease, an MVP score < 0.1 is considered negative for the disease, and 0.1 ≤ MVP score < 0.5 is considered an unclear result, thereby achieving the classification and diagnosis of various congenital neurodevelopmental disorders.

10. A classification, diagnostic, and predictive system for congenital neurodevelopmental disorders, characterized in that, include: A construction module for constructing the diagnostic prediction model for congenital neurodevelopmental disorders as described in claim 8; The acquisition module is used to acquire the detection data after sample methylation detection, and to standardize the detection data to obtain the whole genome methylation analysis matrix of the sample, and to screen out the disease-specific methylation feature matrix. The model processing module is used to input the disease-specific methylation feature matrix data into the corresponding congenital neurodevelopmental disorder diagnosis and prediction model, and calculate the pathogenicity score of methylation variation for each disease. The evaluation module is used to classify and diagnose congenital neurodevelopmental disorders based on the pathogenicity scores of methylation variations corresponding to each disease.

11. An electronic device, characterized in that, The electronic device includes: Processor; and, A memory storing computer-executable instructions, which, when executed, cause the processor to perform the method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs that, when executed by a processor, implement the method of any one of claims 1-9.