Methods for identification and functional assessment of risk structural variants in schizophrenia based on third generation sequencing
By combining third-generation sequencing and multi-tool joint detection with multi-dimensional data analysis, the problems of inaccurate structural variation detection and insufficient functional assessment in existing technologies have been solved. This has enabled precise screening and functional analysis of high-risk variants in schizophrenia, expanding the applicability of genetic research and the application of precision medicine.
Patent Information
- Application Number
- CN202510189754.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-02-20
AI Technical Summary
In current genetic research on schizophrenia, short-read sequencing is insufficient to accurately detect complex structural variations, functional assessment methods mainly focus on coding regions while ignoring non-coding regulatory elements, and studies are concentrated in European populations, resulting in an incomplete understanding of the genetic basis.
Third-generation sequencing technology was used for whole-genome sequencing. Combined with multi-tool joint detection and multi-sample SV results, transcription factor binding analysis, SCZ drug target data and tissue cell-specific expression data were integrated to evaluate the gene regulatory effects of potential pathogenic SVs and their mechanisms of action in SCZ, and to screen high-risk SCZ risk genes.
This improved the accuracy of structural variation detection and pathogenicity assessment, expanded functional analysis, constructed a systematic SCZ risk gene screening system, and enhanced the applicability of the research and its reference value for precision medicine.
Smart Images

Figure CN120148608B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of molecular biology, and particularly relates to a method for identifying and functionally evaluating schizophrenia risk structural variations based on third-generation sequencing. BACKGROUND
[0002] Schizophrenia (SCZ) is a severe neuropsychiatric disorder affecting approximately 1% of the global population. The disease has high genetic susceptibility and complex etiology, involving the interaction of multiple genes and environmental factors. Although genetic research on SCZ has made some progress, a number of genome-wide association studies (GWAS) have identified a number of single nucleotide polymorphisms (SNPs) and copy number variations (CNVs) associated with SCZ, but the functional mechanisms of these variations have not been fully elucidated. In addition, due to the significant genetic heterogeneity of SCZ, research relying solely on common variations cannot fully explain the genetic basis of SCZ, so it is necessary to further explore rare variations with greater biological impact.
[0003] In recent years, structural variations (SVs) as a type of genome variation with greater impact have gradually attracted attention in SCZ genetic research. SVs include large fragment insertion (INS), deletion (DEL), inversion (INV), repeat amplification (DUP), and tandem repeat expansion (TR-Expansions), which can affect gene coding regions or regulatory elements, and thus alter gene expression. Previous studies have shown that some SCZ patients carry rare SVs, and these SVs often affect genes related to neural development and synaptic function. However, most SCZ-related SV studies are still mainly based on short-read sequencing (SRS) data, and the read length limitation of SRS makes it difficult to accurately detect complex SVs, especially those involving repetitive sequences or genomic rearrangements. For example, Lee et al. conducted a family study on 10 SCZ patients and 5 healthy controls using LRS sequencing, identifying 88 medium-sized (50-2000 bp) SVs that may explain part of the "missing heritability" of schizophrenia. In addition, current genetic research on SCZ mainly focuses on European populations, while SV research on non-European populations is less common, limiting the applicability of existing research and making it difficult to fully analyze the genetic basis of SCZ.
[0004] Current research still has several shortcomings in the identification and functional assessment of SCZ-related variants (SVs). First, due to the limitations of the SRS (Self-Rating System), large-scale studies have a high false-negative rate in SV identification, particularly in detecting SVs involving repetitive sequences and complex structural regions. Second, even when some studies have identified SCZ-related SVs, existing functional assessment methods mainly focus on coding region variations, neglecting the impact of SVs on non-coding regulatory elements (such as promoters, enhancers, and transcription factor binding sites), leading to insufficient understanding of the potential pathogenicity of SVs. Finally, current research rarely integrates multidimensional data (such as single-cell transcriptomics, gene regulatory networks, and SCZ drug target data) to deeply analyze the functional impact of SVs, leaving the pathogenic mechanisms of some high-risk SVs unclear. Summary of the Invention
[0005] The purpose of this invention is to provide a method for identifying and functionally assessing structural variants (SVs) in schizophrenia risk based on third-generation sequencing, in order to address the problems of incomplete SV detection, insufficient functional assessment, and lack of population studies in current SCZ genetic research.
[0006] To address the aforementioned problems, a first aspect of the present invention provides a method for identifying and functionally assessing structural variants in schizophrenia risk based on third-generation sequencing, comprising:
[0007] Step 1: Extract DNA from peripheral blood samples of schizophrenia patients and perform whole-genome sequencing using third-generation sequencing technology;
[0008] Step 2: Conduct multi-tool joint detection of SV, and integrate the SV results from multiple tools and multiple samples to generate a high-confidence SV set;
[0009] Step 3: Screen patient-specific SVs by cross-cohort comparison, prioritize them using SVJudge, identify high-risk potential pathogenic SVs, and record SVJudge scores;
[0010] Step 4: Combining transcription factor binding analysis, SCZ drug target data, and tissue- and cell-specific expression data, assess the gene regulatory effects of potential pathogenic SVs and their mechanisms of action in SCZ.
[0011] Step 5: Screen for SCZ risk genes based on SVJudge scores and patient carrier status, and verify their genetic risk and pathogenic mechanism through pathway enrichment, protein interaction network and functional module analysis.
[0012] Current short-read sequencing technologies are difficult to accurately analyze complex SVs, resulting in some potential pathogenic variants being missed, and traditional functional evaluation methods mainly focus on coding regions, and the impact of SVs on regulatory elements is not considered. In addition, genetic studies of SCZ have mainly focused on European populations, while SV studies in non-European populations are less common, affecting the general applicability of research results.
[0013] Preferably, DNA in the peripheral blood sample of a patient with schizophrenia is extracted, and whole genome sequencing is performed using third-generation sequencing technology, including:
[0014] DNA in the peripheral blood sample of a patient with schizophrenia is extracted; wherein the patient with schizophrenia refers to a patient in the population who is clinically diagnosed as having first-episode schizophrenia, and has no history of other neuropsychiatric or mental disorders, such as epilepsy or mental retardation.
[0015] Whole genome sequencing is performed using third-generation sequencing technology; wherein the third-generation sequencing technology is: PacBio Continuous Long Reads (CLR);
[0016] Preferably, SVs are detected by multiple tools, and the SV results of multiple tools and multiple samples are integrated to generate a high-confidence SV set, including:
[0017] S1: using the NGMLR tool to align long reads with the GRCh37 human reference genome, and then using cuteSV, pbsv and Sniffles to detect SVs from the alignment results;
[0018] S2: using the Jasmine tool to perform SV merging of multiple detection tools for the same sample, and retaining high-confidence SVs supported by at least two tools;
[0019] S3: using the Jasmine tool to perform multi-sample SV merging to obtain high-confidence SVs at the population level.
[0020] It should be noted that "long reads" refer to data fragments generated after the sequencing process, which represent a small segment of the original DNA or cDNA molecule. Unlike short-read sequencing (SRS), which produces shorter fragments (usually 100-300 bp), long-read sequencing (LRS) can generate continuous sequence fragments of several thousand to millions of bases. These long reads retain more complete genomic structure information, making them have significant advantages in complex structural variation (SV) detection, repeat sequence analysis, haplotype assembly, epigenetic modification analysis, etc.
[0021] NGMLR is a read aligner designed for long read sequencing data. It can handle high error rate long read data and effectively identify complex structural variations (SVs). NGMLR achieves efficient and accurate sequence alignment across the whole genome by considering insertion, deletion, inversion, etc. It is one of the pre-processing steps for structural variation analysis.
[0022] cuteSV, pbsv and Sniffles are three structural variation (SV) detection tools developed for long read sequencing (PacBio / Nanopore) data. cuteSV uses a segment matching and soft clipping based strategy, with high efficiency and low false positive rate, suitable for large-scale SV analysis; pbsv is an official tool for PacBio, optimized for HiFi and CLR data, accurately detecting insertions and deletions; Sniffles identifies multiple SV types through alignment information, soft clipping and alignment spacing, supporting single sample and multi-sample analysis. These tools can be used together to improve the coverage and accuracy of SV detection.
[0023] Jasmine is a tool for SV integration and deduplication. It can aggregate SV results from different detection tools, merge based on coordinates, variant types, sequence similarity, etc. to generate a high-confidence SV dataset.
[0024] Preferably, patient-specific SVs are screened across cohorts, prioritized by SVJudge, and identified as rare and high-risk potential pathogenic SVs in the normal population, located in coding regions, promoters, enhancers, including:
[0025] Using SVJudge, potential harmful SVs are identified based on genomic annotation results and public population SV datasets, and SVJudge scores are recorded:
[0026] S1: Analyze the location of SV in the genome, determine its association with key functional regions, and establish the relationship between SV and affected genes through genomic annotation or enhancer-promoter interaction information.
[0027] S2: Functional region weight assignment. Evaluate whether the SV falls within CDS, UTR, promoter, enhancer, TAD boundary, and chromatin open region, and assign different weights based on the importance of different regions. For coding region SV, analyze whether it causes frameshift or premature stop codon;
[0028] S3: Population tolerance analysis, calculate the enrichment degree of SV in the normal population, use the gwRVIS framework to evaluate whether the SV is located in the evolutionarily conserved region, and combine the LOEUF value of gnomAD to evaluate the knockout sensitivity of the affected gene. If the SV is located in a low tolerance region (LOEUF is low), it may have higher pathogenicity.
[0029] S4: Gene dosage sensitivity evaluation, integrating HI (haploinsufficiency) and TS (trisomy sensitivity) data to evaluate whether SVs can cause dosage abnormalities;
[0030] S5: Population background correction, using the background occurrence rate of SVs in the normal population to correct the tolerance weight of different functional regions.
[0031] S6: Repeat region re-filtering, if a SV overlaps with a repeat region more than 50%, it is filtered based on the length of the repeat unit, even if it does not meet the overlap threshold, as long as the number of repeat units caused is similar, it is considered as a consistent SV;
[0032] S7: Calculate the final pathogenicity score of SV, considering the importance of functional regions, the frequency of SV in the population, the tolerance of SV in the population, and other factors such as gene tolerance to evaluate the pathogenicity of SV and give a pathogenicity score.
[0033] S8: Verify the accuracy of SVJudge using ClinVar dataset and evaluate its prediction performance by AUC.
[0034] It should be noted that the key functional regions of the genome include: CDS, UTR, promoter, enhancer, TAD boundary and chromatin development region. CDS (coding sequence) is the region of a gene that can be transcribed and translated, determining the amino acid sequence of a protein; UTR (untranslated region) includes 5' UTR and 3' UTR, which regulates the stability and translation efficiency of mRNA; the promoter is located upstream of the gene and contains the RNA polymerase binding site, which determines the transcription initiation of the gene; the enhancer is a long-distance regulatory element that enhances gene expression by binding to transcription factors; the TAD boundary limits the interaction of genes and regulatory elements within the chromatin, maintaining the spatial organization structure of gene expression; the chromatin open region is the unpackaged chromatin region that allows transcription factors and RNA polymerase to bind and regulate gene expression.
[0035] gwRVIS (Genome-wide Residual Variation Intolerance Score) is an index for evaluating the tolerance of genomic regions to structural variations (SV). This method calculates the variation enrichment degree of each genomic region based on population data, and estimates the sensitivity of the region to SV by adjusting the background mutation rate. Regions with lower gwRVIS values are generally less tolerant to variations, indicating that SVs in this region are more likely to be pathogenic. This index is often used for SV prioritization, helping to identify potential high-risk structural variations.
[0036] gnomAD (Genome Aggregation Database) is a large-scale human genome variation database that integrates variation data from multiple large-scale genome and exome sequencing projects, aiming to provide a comprehensive resource to explore the frequency and pattern of human genome variations; LOEUF (Loss-of-function Observed / Expected Upper bound Fraction) is an index provided by gnomAD to measure the tolerance of genes to loss-of-function mutations. Genes with low LOEUF values are sensitive to LoF mutations and may cause diseases; genes with high values are tolerant to LoF mutations and have less impact. This index is used to assess the potential impact of SV (such as deletion) on gene function.
[0037] HI (Haploinsufficiency): Measures whether a gene is sensitive to copy number reduction (such as deletion), genes with high HI may cause loss of function and disease in single copy state.
[0038] TS (Triplosensitivity): Evaluates the tolerance of genes to copy number increase (such as repeat amplification), genes with high TS may cause abnormal expression and pathological changes in multiple copy state
[0039] ClinVar is a public clinical variation database maintained by NCBI, storing and sharing information on the clinical significance of gene variations and their association with diseases. The database integrates variation data from clinical laboratories, research institutions and database projects (such as OMIM, dbSNP, gnomAD). The database covers SV and classifies it according to its impact on disease as pathogenic, potentially pathogenic, benign, potentially benign, uncertain significance (VUS), etc. ClinVar is widely used in genetic research, disease diagnosis and precision medicine, and is an important reference database for evaluating the pathogenicity of variations.
[0040] Preferably, combined with transcription factor binding analysis, SCZ drug target data and tissue cell-specific expression data, the regulatory effect of potential pathogenic SV on genes and its mechanism of action in SCZ are evaluated, including:
[0041] S1: Analyze the impact of SV on transcription factor binding. Using the conserved transcription factor binding site data of UCSC Table Browser and the PWM matrix of CIS-BP database, calculate the change in transcription factor binding strength before and after SV. Use TFM-PVALUE to evaluate the destruction or enhancement of binding sites and screen SV with significant impact.
[0042] S2: Assess the impact of SV on SCZ drug target data. Integrate SCZ drug target information from DrugBank and CTD databases, screen target genes affected by SV and their regulatory factors, and assess the potential impact of SV on SCZ drug action targets.
[0043] S3: Assess the regulatory effect of SV based on tissue cell-specific expression regulation data. Combine hippocampus, adult brain and fetal brain transcriptional regulation networks to analyze the tissue-specific regulatory relationship of SV-affected genes. At the same time, use PsychENCODE single-cell transcriptome data to assess the expression changes of these genes in different brain cell types (such as neurons, astrocytes, oligodendrocytes).
[0044] It should be noted that the UCSC Table Browser is a tool provided by the UCSC Genome Database, which can query and extract genomic annotation data, including genes, transcripts, regulatory elements and variation information, etc.
[0045] CIS-BP database (Catalog of Inferred Sequence Binding Preferences) is a database that collects transcription factor binding sites (TFBS) and their position weight matrix (PWM), which is used to predict the binding of transcription factors and DNA sequences.
[0046] TFM-PVALUE is a tool based on PWM matrix to calculate the impact of mutations on transcription factor binding sites, which can be used to assess the potential impact of SV on transcriptional regulation.
[0047] DrugBank is a comprehensive database containing FDA-approved drugs, experimental drugs and their target gene information, widely used in drug-gene interaction analysis.
[0048] PsychENCODE is a research project that integrates multi-omics data such as transcriptome, epigenetic and single-cell RNA sequencing related to human brain development and mental illness, commonly used in molecular mechanism research of neuropsychiatric diseases such as schizophrenia.
[0049] Preferably, SCZ risk genes are screened based on SVJudge scores and patient carrying conditions, including:
[0050] This method screens high-confidence SCZ risk genes based on the pathogenicity score of potential pathogenic SV and the patient carrying frequency.
[0051] S1: Calculate the impact score of each gene, accumulate the pathogenicity score of all potential pathogenic SVs carried by all patients in the SCZ patient cohort, and ensure that all SVs affecting the gene contribute to it.
[0052] S2: Set the screening threshold, take the median of all gene scores plus standard deviation as the benchmark, screen out genes significantly affected by SV, highlight genes with greater impact in the patient population;
[0053] S3: Only keep high-score genes carried in at least two patients to exclude individual-specific effects and ensure higher reliability of the screened candidate SCZ risk genes.
[0054] Preferably, the genetic risk and potential pathogenic mechanism are verified by pathway enrichment, protein interaction network and gene function module analysis, including:
[0055] S1: Use ToppGene and SynGO tools to perform functional enrichment on candidate SCZ risk genes. If enrichment is found in known SCZ-related pathways, it indicates that the candidate gene captures the biological mechanism of SCZ.
[0056] S2: Establish a protein-protein interaction network for SCZ candidate genes based on the human gene connection group database. If the newly discovered genes in the network have more significant connections with known SCZ genes, it indicates that the newly discovered genes through SV have potential risk for SCZ.
[0057] S3: Use the Louvain method to detect gene modules in the protein interaction network. If the gene modules are enriched in known SCZ-related pathways, it indicates that the candidate gene captures the biological mechanism of SCZ.
[0058] It should be noted that ToppGene is a gene function analysis and candidate gene prioritization tool that can be used for gene enrichment analysis, pathway analysis, protein interaction analysis and candidate gene screening, and is commonly used for identification of disease-related genes.
[0059] SynGO is a gene ontology database focusing on synaptic function, which annotates genes related to synapses based on experimentally verified data, and is commonly used in the study of nervous system diseases.
[0060] The SCZ-related pathways include neurotransmitter pathways (glutamate, dopamine, GABA transmission), neural development pathways (neuron migration, synaptic plasticity), immune and inflammatory pathways, and metabolic and mitochondrial function pathways.
[0061] The known SCZ genes in the network refer to genes previously statistically associated with SCZ through large-scale genome-wide association studies (GWAS), differential expression analysis, differential methylation analysis, etc. New genes refer to potential related genes that have not been clearly associated with SCZ but have been identified in this study.
[0062] The Louvain method is a community detection algorithm based on modularity optimization, which is used to find densely connected subgroups (communities) in complex networks. This method iteratively optimizes modularity by first assigning nodes to the best community and then merging communities to build a new network until the modularity no longer significantly improves. The Louvain method is computationally efficient and suitable for large-scale network analysis, such as functional module detection in gene co-expression networks and protein interaction networks.
[0063] The beneficial effects of the present application are:
[0064] 1. Improve the accuracy of structural variant (SV) detection and overcome the limitations of short-read sequencing. By using high-precision sequencing technology and optimized analysis strategy, the present application can more comprehensively detect complex structural variants, including insertions, deletions, inversions, and repeat expansions, especially in high-repetition and structurally complex genomic regions, significantly improving the accuracy of variant identification. This improvement helps to reduce false negatives and false positives, ensuring that critical pathogenic variants are not missed.
[0065] 2. Optimize the pathogenicity assessment of SV and improve the screening ability of high-risk variants. The present application combines multi-dimensional genetic and functional information to systematically assess the potential pathogenicity of SV, overcoming the limitations of traditional methods that rely only on coding region variants. By comprehensively analyzing the impact of SV on gene function, regulatory elements, and genomic stability, the prioritization ability of pathogenic variants is improved, enabling more accurate identification of high-risk SV.
[0066] 3. Expand the functional analysis of SV in SCZ and explore its mechanism in depth. Existing researches mainly focus on small-scale gene mutations, while the present application integrates multi-omics data to systematically assess the potential impact of SV on gene regulation, not only covering coding regions but also focusing on functional variants in non-coding regulatory regions. This strategy can more comprehensively reveal the impact of SV on the key molecular mechanisms of SCZ, providing a deeper perspective for understanding the genetic basis of the disease.
[0067] 4. Construct a systematic SCZ risk gene screening system to improve the biological credibility of candidate genes. By combining gene function, regulatory network, and disease pathway analysis, the present application can more reliably screen high-confidence candidate genes related to SCZ, compared to traditional methods that rely only on individual variant analysis. This method can more effectively identify pathogenic genes with important biological significance. This screening system helps to promote SCZ genetic risk assessment and provides theoretical support for the development of potential biomarkers.
[0068] 5. Improve the applicability of SCZ genetic research in human populations and promote the development of precision medicine. Existing SCZ genetic research mainly focuses on European populations, while the present invention can be based on genetic data from non-European populations, filling the gap in SCZ research with different genetic backgrounds in different populations. This progress not only improves the broad applicability of research results, but also provides important reference for future personalized medicine and precise disease prediction in different populations. BRIEF DESCRIPTION OF DRAWINGS
[0069] The present invention will be further described below with reference to the accompanying drawings.
[0070] Figure 1 A schematic diagram of the method steps of the embodiments of the present invention is shown in the figure.
[0071] Figure 2 The figure shows the regulatory strength of the screened potential pathogenic SVs in different tissues, which are regulated by transcription factors.
[0072] Figure 3 The figure shows the enrichment of SCZ risk genes in synaptic structure. DETAILED DESCRIPTION
[0073] The technical solutions in the embodiments of the present invention will be described below in detail, and it is obvious that the described embodiments are only a part of the embodiments of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present invention.
[0074] Referring to Figure 1 The first aspect of the present invention provides a method for identifying and evaluating the function of schizophrenia risk structural variations based on third-generation sequencing, which comprises:
[0075] Step 1: Extract DNA from peripheral blood samples of schizophrenia patients, and perform whole genome sequencing using third-generation sequencing technology.
[0076] Step 2: Perform multi-tool joint detection of SVs, and integrate multi-tool and multi-sample SV results to generate a high-confidence SV set.
[0077] Step 3: Cross-coupling comparison to screen patient-specific SVs, combined with SVJudge for prioritization, identification of high-risk potential pathogenic SVs, and recording of SVJudge scores.
[0078] Step 4: Combined with transcription factor binding analysis, SCZ drug target data and tissue cell-specific expression data, evaluate the gene regulatory effect of potential pathogenic SVs and their mechanism of action in SCZ.
[0079] Step five: screening SCZ risk genes based on SVJudge scores and patient carrying status, and verifying their genetic risk and pathogenic mechanism through pathway enrichment, protein interaction network and functional module analysis.
[0080] Extracting DNA from the peripheral blood samples of schizophrenia patients, using a third-generation sequencing platform for whole genome sequencing, including:
[0081] S1: Extracting DNA from the peripheral blood samples of schizophrenia patients; wherein the schizophrenia patients refer to patients in the Chinese population who have been clinically diagnosed with first-episode schizophrenia and have no history of other neuropsychiatric or mental disorders, such as epilepsy or mental retardation.
[0082] S2: Using a third-generation sequencing technology for whole genome sequencing; wherein the third-generation sequencing technology is: PacBio Continuous Long Reads (CLR);
[0083] Carrying out multi-tool joint detection of SV and integrating multi-tool and multi-sample SV results to generate a high-confidence SV set, including:
[0084] S1: Using the NGMLR tool to align long reads with the GRCh37 human reference genome, and then using cuteSV, pbsv and Sniffles to detect SVs from the alignment results;
[0085] S2: Using the Jasmine tool to carry out SV merging of multiple detection tools for the same sample, retaining high-confidence SVs supported by at least two tools;
[0086] S3: Using the Jasmine tool to carry out multi-sample SV merging to obtain a high-confidence SV set at the population level.
[0087] Cross-queue comparison to screen patient-specific SVs, combined with SVJudge for prioritization, to identify high-risk potential pathogenic SVs, including:
[0088] Using SVJudge, based on genomic annotation results and public public population SV datasets, to identify potentially harmful SVs, and record SVJudge scores:
[0089] S1: Analyzing the location of SVs in the genome to determine their association with key functional regions, and establishing the relationship between SVs and affected genes through genomic annotation or enhancer-promoter interaction information.
[0090] S2: Functional region weight assignment. Evaluate whether SV falls into CDS, UTR, promoter, enhancer, TAD boundary, chromatin open region, and assign different weights based on the importance of different regions. For coding region SV, analyze whether it leads to frame shift or premature stop codon;
[0091] S3: Population tolerance analysis, calculate the enrichment degree of SV in the normal population, use the gwRVIS framework to evaluate whether the SV is located in the evolutionarily conserved region, and combine the LOEUF value of gnomAD to evaluate the knockout sensitivity of the affected gene. If the SV is located in a low tolerance region (LOEUF is low), it may have higher pathogenicity.
[0092] S4: Gene dosage sensitivity evaluation, integrate HI (haploinsufficiency) and TS (trisomy sensitivity) data to evaluate whether the SV may cause dosage abnormalities;
[0093] S5: Population background correction, use the background occurrence rate of SV in the normal population to correct the tolerance weight of different functional regions.
[0094] S6: Repeat region re-filtering, if the SV overlaps with the repeat region more than 50%, filter based on the repeat unit length, even if it does not meet the overlap threshold, as long as the number of repeat units caused is similar, it is considered as consistent SV;
[0095] S7: Calculate the final pathogenicity score of SV, consider the importance of functional regions, the frequency of SV in the population, the tolerance of SV in the population, and the gene tolerance, etc. Comprehensive evaluation of SV pathogenicity and give pathogenicity score.
[0096] S8: Use ClinVar dataset to verify the accuracy of SVJudge, and evaluate its prediction performance by AUC.
[0097] Combined with transcription factor binding analysis, SCZ drug target data and tissue cell specific expression data, evaluate the gene regulation effect of potential pathogenic SV and its mechanism of action in SCZ, including:
[0098] S1: Analyze the effect of SV on transcription factor binding. Use the conserved transcription factor binding site data of UCSC Table Browser and the PWM matrix of CIS-BP database to calculate the change of transcription factor binding strength before and after SV. Use TFM-PVALUE to evaluate the destruction or enhancement of binding sites, and screen SV with significant impact.
[0099] S2: Evaluate the impact of SV combined with SCZ drug target data. Integrate SCZ drug target information in DrugBank and CTD databases, screen target genes and their regulators affected by SV, and evaluate the potential impact of SV on SCZ drug action targets.
[0100] S3: Evaluate the SV regulatory effect based on the tissue cell-specific expression regulatory data. Combined with the hippocampus, adult brain and fetal brain transcriptional regulatory network, analyze the tissue-specific regulatory relationship of SV-affected genes. At the same time, using PsychENCODE single-cell transcriptome data, evaluate the expression changes of these genes in different brain cell types (such as neurons, astrocytes, oligodendrocytes).
[0101] Screening of SCZ risk genes based on SVJudge score and patient carrying status, including:
[0102] This method screens high-confidence SCZ risk genes based on the pathogenicity score of SV and the patient carrying frequency.
[0103] S1: Calculate the impact score of each gene, accumulate the pathogenicity score of all SVs carried by each patient in the SCZ patient cohort, and ensure that all SVs affecting the gene are included;
[0104] S2: Set a screening threshold, take the median plus standard deviation of all gene scores as the benchmark, and screen out genes significantly affected by SV to highlight genes with greater impact in the patient population;
[0105] S3: Only keep high-score genes carried in at least two patients to exclude individual-specific effects and ensure that the screened candidate SCZ risk genes have higher reliability.
[0106] The genetic risk and potential pathogenic mechanism are verified by pathway enrichment, protein interaction network and gene function module analysis, including:
[0107] S1: Use ToppGene and SynGO tools to perform functional enrichment on candidate SCZ risk genes. If enrichment is found in known SCZ-related pathways, it indicates that the candidate gene captures the biological mechanism of SCZ;
[0108] S2: Establish a protein-protein interaction network for SCZ candidate genes based on the human gene connection database. If the newly discovered genes in this network have more significant connections with known SCZ genes, it indicates that the newly discovered genes through SV have potential risk for SCZ;
[0109] S3: Use Louvain method to detect gene modules in protein interaction network. If the gene modules are enriched in known SCZ-related pathways, it indicates that the candidate gene captures the biological mechanism of SCZ.
[0110] For example: This example is called ZIB_SCZ cohort, a total of 141 patients with schizophrenia, all clinically diagnosed as first-episode schizophrenia (or psychotic clinical high syndrome), the inclusion criteria are as follows:
[0111] (1) Age, 16-40 years old (gender unlimited);
[0112] (2) No other neuropsychiatric disorders (mental retardation, epilepsy or substance dependence);
[0113] (3) No history of other nervous system diseases;
[0114] (4) No history and family history of epilepsy;
[0115] (5) No mental retardation;
[0116] All patients need to take blood samples for blood routine and gene sequencing. Blood samples are collected by the center, 5mL per person.
[0117] Step one: extract DNA from the peripheral blood samples of patients with schizophrenia, and use third-generation sequencing technology for whole genome sequencing;
[0118] For example: DNA sequencing samples are extracted from whole blood and purified using a modified CTAB method. Sample quality is evaluated by agarose gel electrophoresis, Nanodrop 2000 spectrophotometer and Qubit fluorescent dye method. ExpressTemplate Prep Kit 2.0 (PacBio, #100-938-900) is used to construct a 20kb SMRTbell library, and Qubit 2.0 is used for preliminary quantification. Pulse field electrophoresis is used to verify fragment distribution to ensure that the library insert size is about 20kb, the concentration is not less than 70ng / μL, and the total amount is not less than 7μg. Sequencing uses the PacBio Sequel II platform, using CLR mode and diffusion loading (Diffusion loading) method. The original data is initially Polymerreads, after quality filtering, remove sequences with length less than 50bp, quality value less than 0.8 and self-connection or adapter, finally obtain Subreads as subsequent analysis data.
[0119] Step two: carry out multi-tool joint detection of SV, and integrate multi-tool and multi-sample SV results to generate a high-confidence SV set.
[0120] For example: This example designs a high-confidence SV detection pipeline to extract high-confidence structural variations.
[0121] Firstly, long read alignment was performed using NGMLR (version 0.2.7) to align PacBio CLR sequencing data to the GRCh37 reference genome, and the alignment results were sorted using SAMtools to generate a BAM file. The alignment parameters were set according to the PacBio CLR recommended settings to optimize the alignment accuracy of long read sequences.
[0122] Next, three SV detection tools were used for SV identification. The parameter settings of cuteSV (version 1.0.11) were as follows: the minimum number of sequencing reads supporting the variation was 4, the maximum cluster deviation for insertion variations was set to 100 bases, the merging ratio for insertion variations was set to 0.3, the maximum cluster deviation for deletion variations was set to 200 bases, and the merging ratio for deletion variations was set to 0.5. The parameters of pbsv (version 2.4.1) included: the minimum number of sequencing reads supporting the variation in a single sample was 4, the minimum SV length was set to 30 bases, and the minimum number of sequencing reads supporting the breakpoint when calling structural variations across samples was 4. The parameter settings of Sniffles (version 1.0.12) were as follows: the minimum number of sequencing reads supporting the variation was 4, and the minimum sequence length was set to 500 bases. The parameters not shown were set to default.
[0123] The SV detection results were integrated using Jasmine (version 1.1.5), and the SVs with the same type and overlapping degree exceeding 50% and the breakpoint distance within 1000 bases were defined as consistent SVs. The specific parameters included: the maximum linear distance ratio was set to 0.5, the maximum breakpoint distance was set to 1000 bases, the minimum overlap ratio was set to 50%, the positive and negative strand information was ignored, and the genotyping data was output. Finally, only the high-confidence SV set supported by at least two tools was retained, and the cuteSV results were used as the main selection, and pbsv was used as the secondary selection because it had been verified to have high reliability at 15-20 times PacBio sequencing depth.
[0124] When screening SVs, the following criteria were set:
[0125] (1) Supported by at least 4 sequencing reads to ensure the stability of the variation in the sequencing data.
[0126] (2) The length of insertion and deletion variations is less than 1 million bases to avoid false positives caused by detecting too long variations.
[0127] (3) Remove variations in mitotic regions or intergenic regions annotated in the UCSC database to reduce noise and improve specificity.
[0128] (4) Remove variations in GRCh37 that fall in regions with homology.
[0129] The median number of SVs identified in each individual in this example was 14,392, ranging from 11,658 to 16,876, with deletions and insertions being the most common, accounting for 49.0% and 40.4%, respectively. The length distribution of SVs was consistent with the published pattern of third-generation sequencing data, and the overlap rate with Chinese population data was the highest compared with different population data, showing obvious population specificity, further supporting the reliability of SV identification in this study.
[0130] Step three: Cross-cascade comparison of patient-specific SVs, combined with SVJudge for prioritization, to identify high-risk potential pathogenic SVs;
[0131] For example: using SVJudge, optimize the pathogenicity evaluation of SVs, and improve the screening ability of high-risk variants in SCZ.
[0132] Based on the developed SVJudge, first annotate the SVs, evaluate their impact on the coding and regulatory regions, and combine the distribution frequency of the variants in the population to score, in order to reduce the errors caused by inaccurate breakpoints. Through an adjustable scoring matrix, prioritize variants that affect low-tolerant exons, promoters and enhancers, and focus on structural variations that may significantly affect gene expression.
[0133] The effectiveness of SVJudge was verified by the ClinVar dataset, and it performed better than other SV pathogenicity prediction tools (StrVCTVRE and CADD-SV), with an AUC of 0.89. Based on the prediction score, high-risk insertions, deletions, duplications and inversions were screened, different score thresholds were set to control the false positive rate to be no more than 0.1, and the variants in translocation and genomic unstable regions were removed to reduce misjudgment.
[0134] In summary, we identified 352 high-confidence potential pathogenic structural variants, among which 95 involved tandem repeat expansion, 257 located in non-repetitive regions, and collectively affected 502 genes in coding, promoter or enhancer regions. These variants were enriched in schizophrenia-associated regulatory regions, suggesting that they might affect disease susceptibility through gene regulation. Notably, most of the variants involving tandem repeat expansion were insertions, suggesting that they might affect gene function through repeat expansion mechanisms, such as the 5-HTTLPR VNTR in the promoter of SLC6A4 gene, whose long allele was significantly more frequently carried in schizophrenia patients than other datasets. In addition, some variants affected the coding regions of known schizophrenia risk genes, such as the inversion of exons 23-25 of CACNA1H gene, which encodes a voltage-gated calcium channel that has been reported as a significant risk gene for schizophrenia. Other variants involved schizophrenia-associated pathways such as neuronal generation, cell signal regulation and cell adhesion, further supporting the potential pathogenic role of the screened SVs in schizophrenia.
[0135] Step four: Evaluate the gene regulatory effects of potential pathogenic SVs and their mechanisms of action in SCZ by combining transcription factor binding analysis, SCZ drug target data and tissue cell-specific expression data.
[0136] For example: As a large number of potential pathogenic SVs fall into non-coding regions, to further explore their functional mechanisms, this embodiment focuses on analyzing their interference with transcription factor binding and the regulatory effects on SCZ-related genes. First, based on the transcription factor binding site data of the GRCh37 reference genome, combined with the information of 1,898 transcription factors in the CIS-BP database, the change in binding affinity of the sequence before and after mutation was calculated using the PWM (Position Weight Matrix) method, and the significance was evaluated using TFM-PVALUE, and 157 SVs that disrupt transcription factor binding were screened, among which 65 affect known SCZ-related genes or SCZ drug targets. Further analysis showed that 94.9% (149) of SVs affect SCZ-related transcription factors or gene expression, of which 101 meet both conditions.
[0137] Referring to Figure 2 In hippocampus, adult brain and fetal brain tissues, the 51 genes affected by SVs form 120 regulatory relationships, with significantly higher regulatory strength than other tissues. By consulting single-cell expression data, we found that these genes were significantly enriched in astrocytes of SCZ patients and oligodendrocytes of bipolar disorder patients, suggesting that they might affect synaptic signaling and neural metabolism.
[0138] In addition, in the drug target screening of SCZ and related brain diseases, 7 affected transcription factor genes were found to be SCZ approved drug targets, 5 were drug targets for other mental illnesses (such as personality disorders and bipolar disorders), and 21 genes regulated by SV were SCZ treatment targets. These results suggest that SV may affect the risk of SCZ by regulating gene expression and provide potential targets for individualized treatment.
[0139] Step five: screening of SCZ risk genes based on SVJudge scores and patient carrying status, and verification of genetic risk and pathogenic mechanism through pathway enrichment, protein interaction network and functional module analysis.
[0140] For example: this embodiment further screens potential pathogenic SV-driven SCZ risk genes.
[0141] First, genes affected by SVs carried by at least two patients and higher than the set threshold of SVJudge scores were selected, and finally 130 SCZ risk genes were identified, of which 85 were affected by tandem repeat-related SVs and 45 were only affected by non-tandem repeat-related SVs. Based on pathway analysis, these genes are highly enriched in MHC-related cellular components, synaptic-related functions and neural connection pathways, indicating their potential role in SCZ.
[0142] Referring to Figure 3 Among the 22 genes (14 of which are affected by tandem repeat-related structural variations) screened from the SynGO database for synaptic structure and function annotations, 22 genes (14 of which are affected by tandem repeat-related structural variations) are enriched in synaptic-related pathways, mainly involving presynaptic processes (q = 1.63 x 10 -3 ) and synaptic vesicle cycle (p = 3.22 x 10 -3 ). In addition, these genes have significant functional associations with known SCZ susceptibility genes, especially in the gene network related to regulating synaptic function and neural development, further verifying their potential role in the pathogenesis of SCZ.
[0143] Subsequently, the protein interaction network (PPI) of these genes was constructed using the STRING database, and six significant functional modules were identified, which covered known biological processes closely related to SCZ such as cytoskeleton and neural tissue, stress response and neuronal apoptosis, immune response and cell adhesion, and calcium ion transport. In addition, two less studied modules are involved in cell division regulation, intracellular transport and mitochondrial membrane structure, suggesting that they may play a role in SCZ.
[0144] Among these risk genes, some are located at known SCZ susceptibility loci, such as RAC1 and CREB5, which affect neuronal dendritic spine development and regulation of cAMP signaling pathway, respectively, and are associated with drugs approved for SCZ treatment. Overall, the functional enrichment and network analysis results of these genes suggest that they play an important role in the genetic mechanism of SCZ and can provide new targets to help precision medicine research.
[0145] It should be noted that the relational terms herein, such as first and second, and the like, are used solely to distinguish one from another entity or action without necessarily requiring or implying any actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0146] While embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, combinations, and variations of the embodiments can be undertaken without departing from the spirit and scope of the present application, which is defined solely by the claims and their equivalents.
Claims
1. A method for identifying and functionally assessing structural variations in schizophrenia risk based on third-generation sequencing, characterized in that, include: Step 1: Extract DNA from peripheral blood samples of schizophrenia patients and perform whole-genome sequencing using third-generation sequencing technology; Step 2: Conduct multi-tool joint detection of SV, and integrate the SV results from multiple tools and multiple samples to generate a high-confidence SV set; Step 3: Screen patient-specific SVs by cross-cohort comparison, prioritize them using SVJudge, identify high-risk potential pathogenic SVs, and record SVJudge scores; SVJudge was used to identify potentially harmful SVs based on genome annotation results and publicly available SV datasets, and SVJudge scores were recorded. S1: Determine the location of SV in the genome, identify its association with key functional regions, and establish the connection between SV and affected genes through genome annotation or enhancer-promoter interaction information; S2: Assigning weights to functional regions; Evaluate whether SVs fall into CDS, UTR, promoters, enhancers, TAD boundaries, or open chromatin regions, and assign different weights based on the importance of different regions; For coding region SVs, analyze whether they lead to code shifts or premature codon termination. S3: Population tolerance analysis: Calculate the enrichment of SV in the normal population, use the gwRVIS framework to assess whether SV is located in an evolutionarily conserved region, and combine the LOEUF value of gnomAD to assess the knockout sensitivity of the affected gene. S4: Gene dose sensitivity assessment: Integrating HI and TS data to assess whether SV causes dose abnormalities; S5: Population background correction: Use the background incidence of SV in the normal population to correct the tolerance weights of different functional areas; S6: Repeated Region Re-filtering: For SVs with serial repetition and high repetition regions, use UCSC Table Browser, RepeatMasker and Tandem Repeat Finder for annotation; S7: Calculate the final pathogenicity score of SV, comprehensively consider the importance of functional regions, the frequency of SV in the population, the tolerance of SV in the population, and genetic tolerance factors to comprehensively assess the pathogenicity of SV and give a pathogenicity score. S8: Validate the accuracy of SVJudge using the ClinVar dataset and evaluate its predictive performance using AUC; Step 4: Combining transcription factor binding analysis, SCZ drug target data, and tissue- and cell-specific expression data, assess the gene regulatory effects of potential pathogenic SVs and their mechanisms of action in SCZ. Step 5: Screen for SCZ risk genes based on SVJudge scores and patient carrier status, and verify their genetic risk and pathogenic mechanism through pathway enrichment, protein interaction network and functional module analysis.
2. The method for identifying and functionally assessing structural variations in schizophrenia risk based on third-generation sequencing according to claim 1, characterized in that, In step one, a schizophrenia patient refers to a patient who has been clinically diagnosed with first-episode schizophrenia and has no prior history of other neuropsychiatric or mental illnesses; the third-generation sequencing technology mentioned is PacBio Continuous Long Reads.
3. The method for identifying and functionally assessing structural variations in schizophrenia risk based on third-generation sequencing according to claim 1, characterized in that, The multi-tool joint detection of SVs refers to using the NGMLR tool to align long reads with the GRCh37 human reference genome, and then using cuteSV, pbsv and Sniffles to compare the alignment results to perform SV detection.
4. The method for identifying and functionally assessing structural variations in schizophrenia risk based on third-generation sequencing according to claim 1, characterized in that, Integrating SV results from multiple tools and samples to generate a high-confidence SV set refers to using the Jasmine tool to merge SVs from multiple detection tools for the same sample, retaining high-confidence SVs supported by at least two tools; The Jasmine tool was used to perform multi-sample SV merging to obtain a high-confidence SV set at the population level.
5. The method for identifying and functionally assessing structural variations in schizophrenia risk based on third-generation sequencing according to claim 1, characterized in that, If SV is located in the low tolerance area in S3, it is pathogenic; In S6, if the SV overlaps with the repeating region by more than 50%, it is filtered based on the length of the repeating unit. Even if the overlap threshold is not met, as long as the number of repeating units is similar, it is considered a consistent SV.
6. The method for identifying and functionally assessing structural variations in schizophrenia risk based on third-generation sequencing according to claim 1, characterized in that, The combined analysis of transcription factor binding, SCZ drug target data, and tissue- and cell-specific expression data was used to assess the gene regulatory effects of potential pathogenic SVs and their mechanisms of action in SCZ, including: S1: Analyze the effect of SV on transcription factor binding; use conserved transcription factor binding site data from the UCSC table browser and the PWM matrix from the CIS-BP database to calculate the changes in transcription factor binding strength before and after SV; use TFM-PVALUE to assess the disruption or enhancement of binding sites and screen for SVs with significant effects. S2: Evaluate the impact of SV by combining SCZ drug target data; integrate SCZ drug target information from DrugBank and CTD databases, screen target genes and their regulatory factors affected by SV, and evaluate the potential impact of SV on SCZ drug targets. S3: Evaluate the regulatory effects of SV based on tissue- and cell-specific expression regulation data; analyze the tissue-specific regulatory relationships of genes affected by SV by combining transcriptional regulatory networks in the hippocampus, adult brain, and fetal brain; and evaluate the expression changes of these genes in different brain cell types using PsychENCODE single-cell transcriptome data.
7. The method for identifying and functionally assessing structural variations in schizophrenia risk based on third-generation sequencing according to claim 1, characterized in that, The screening of SCZ risk genes based on SVJudge scores and patient carrier status includes: This method screens for high-confidence SCZ risk genes based on SV pathogenicity scores and patient carrier frequencies; S1: Calculate the influence score for each gene, and accumulate the pathogenicity scores of all SVs carried by all patients in the SCZ patient cohort to ensure that all SV contributions affecting the gene are included. S2: Set a screening threshold and use the median plus standard deviation of all gene scores as a benchmark to screen out genes that are significantly affected by SV, so as to highlight genes that have a large impact on the patient population. S3: Only high-scoring genes carried in at least two patients are retained to exclude individual-specific influences and ensure that the selected candidate SCZ risk genes have higher reliability.
8. The method for identifying and functionally assessing structural variations in schizophrenia risk based on third-generation sequencing according to claim 1, characterized in that, The analysis, including pathway enrichment, protein-protein interaction networks, and gene functional modules, verifies its genetic risk and potential pathogenic mechanisms, including: S1: Functional enrichment of candidate SCZ risk genes was performed using ToppGene and SynGO tools; S2: Establish a protein-protein interaction network for SCZ candidate genes based on the human genome connectome database; S3: Use the Louvain method to detect gene modules in protein-protein interaction networks.
9. The method for identifying and functionally assessing structural variations in schizophrenia risk based on third-generation sequencing according to claim 8, characterized in that, If S1 is enriched in known SCZ-related pathways, it indicates that the candidate gene captures the biological mechanism of SCZ. In S2, if the newly discovered genes in the network have a more significant connection with the known SCZ genes, it indicates that the newly discovered genes through SV pose a potential risk to SCZ. In S3, if the gene module is enriched in known SCZ-related pathways, it indicates that the candidate gene captures the biological mechanism of SCZ.
Citation Information
Patent Citations
Construction method and application of schizophrenia abnormal gene-metabolism regulation network
CN115206420A
Screening method of sick sinus syndrome mutant gene
CN116052780A