Schizophrenia risk structure variation identification and function evaluation method based on three-generation sequencing

Through a three-generation sequencing method, combined with multi-tool joint detection and SVJudge scoring, multi-dimensional data are integrated for identification and functional evaluation of structural variation, which solves the problems of incomplete detection, insufficient functional evaluation and lack of population research in the existing technology, and achieves more efficient structural variation detection and functional evaluation, which promotes the progress of genetic research on schizophrenia.

CN120148608AActive Publication Date: 2025-06-13FUDAN UNIVERSITY

Patent Information

Application Number
CN202510189754.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-13
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The prior art has shortcomings in the detection and functional assessment of schizophrenia-related structural variants, including the limitations of short-read and long sequencing, the singularity of functional assessment methods, and the lack of population research.

Method used

Using a three-generation sequencing method, the DNA of schizophrenia patients was extracted for whole-genome sequencing, combined with multi-tool joint detection and SVJudge score, multi-dimensional data were integrated for identification and functional evaluation of structural variation.

Benefits of technology

It has improved the accuracy of structural variation detection, optimized pathogenicity assessment, expanded functional analysis, built a systematic risk gene screening system, and improved the applicability of research and the development of precision medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148608A_ABST
    Figure CN120148608A_ABST
Patent Text Reader

Abstract

The invention discloses a schizophrenia risk structure variation identification and function evaluation method based on three-generation sequencing, and relates to the field of molecular biology, whole genome sequencing is performed on peripheral blood DNA of a patient through three-generation sequencing, multi-tool joint detection is adopted, multi-sample results are integrated, and a high-confidence SV data set is generated; through cross-queue comparison, the patient specific SV is screened, and the high-risk potential pathogenic SV is identified in combination with an SV priority ordering tool SVJudge. By combining transcription factor binding analysis, SCZ drug target data and histocyte specific expression data, the influence of SV on gene regulation is evaluated, and the potential action mechanism of SV in SCZ is analyzed. The SCZ risk gene is screened based on SVJudge scoring and patient carrying conditions, the genetic risk and pathogenic mechanism of the SCZ are analyzed and verified through pathway enrichment, a protein interaction network and a functional module, and a new technical means and theoretical basis are provided for genetic research of the SCZ.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of molecular biology, and particularly relates to a method for identifying and functionally evaluating schizophrenia risk structural variations based on third-generation sequencing. Background Art

[0002] Schizophrenia (SCZ) is a severe neuropsychiatric disorder that affects approximately 1% of the global population. This disease has a high degree of genetic susceptibility, and its etiology is complex, involving the interaction of multiple genes and environmental factors. Although certain progress has been made in the genetic research of SCZ, and multiple genome-wide association studies (GWAS) have identified multiple single nucleotide polymorphisms (SNPs) and copy number variations (CNVs) associated with SCZ, the functional mechanisms of these variations have not been fully elucidated. In addition, due to the significant genetic heterogeneity exhibited by SCZ, studies relying solely on common variations cannot fully explain the genetic basis of SCZ. Therefore, it is necessary to further explore rare variations with greater biological impact.

[0003] In recent years, structural variations (SVs), as a type of genomic variation with a greater impact, have gradually attracted the attention of SCZ genetic research. SVs include large fragment insertions (INS), deletions (DEL), inversions (INV), duplication amplifications (DUP), and tandem repeat expansions (TR-Expansions), etc., which can affect gene coding regions or regulatory elements, thereby altering gene expression. Previous studies have shown that some SCZ patients carry rare SVs, and these SVs often affect genes related to neurodevelopment and synaptic function. However, most SCZ-related SV studies are still mainly based on short-read sequencing (SRS) data, and the read length limitation of SRS makes it difficult to accurately detect complex SVs, especially variations involving repetitive sequences or genomic rearrangements. For example, Lee et al. conducted a family study, performing LRS sequencing on 10 schizophrenia patients and 5 healthy controls, and identified 88 medium-sized (50 - 2000 bp) SVs, which may explain part of the "missing heritability" of schizophrenia. In addition, current genetic research on SCZ mainly focuses on European populations, and there are fewer SV studies on non-European populations, resulting in limited applicability of existing research and making it difficult to comprehensively analyze the genetic basis of SCZ.

[0004] Existing studies still have multiple deficiencies in the identification and functional evaluation of SCZ-related SVs. First, due to the limitations of SRS, large-scale studies have a high false negative rate in SV identification, especially in the detection of SVs involving repetitive sequences and complex structural regions. Second, even though some studies have identified SCZ-related SVs, existing functional evaluation methods mainly focus on coding region variations, while ignoring the impact of SVs on non-coding regulatory elements (such as promoters, enhancers, transcription factor binding sites), resulting in insufficient understanding of the potential pathogenicity of SVs. Finally, current studies rarely integrate multi-dimensional data (such as single-cell transcriptomics, gene regulatory networks, SCZ drug target data) to deeply analyze the functional effects of SVs, making the pathogenic mechanisms of some high-risk SVs still unclear. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for identifying and functionally evaluating schizophrenia risk structural variations based on third-generation sequencing, so as to solve problems such as incomplete SV detection, insufficient functional evaluation, and lack of population studies in current SCZ genetic research.

[0006] To achieve the above problems, the first aspect of the present invention provides a method for identifying and functionally evaluating schizophrenia risk structural variations based on third-generation sequencing, including:

[0007] Step 1: Extract DNA from peripheral blood samples of schizophrenia patients and perform whole-genome sequencing using third-generation sequencing technology;

[0008] Step 2: Conduct joint detection of SVs using multiple tools, and integrate the SV results of multiple tools and multiple samples to generate a high-confidence SV set;

[0009] Step 3: Cross-cohort comparison to screen patient-specific SVs, combine with SVJudge for prioritization, identify high-risk potential pathogenic SVs, and record the SVJudge score;

[0010] Step 4: Combine transcription factor binding analysis, SCZ drug target data, and tissue cell-specific expression data to evaluate the gene regulatory effects of potential pathogenic SVs and their mechanisms of action in SCZ;

[0011] Step 5: Screen SCZ risk genes based on the SVJudge score and patient carriage status, and verify their genetic risks and pathogenic mechanisms through pathway enrichment, protein-protein interaction networks, and functional module analysis.

[0012] Existing short-read sequencing technologies are difficult to accurately resolve complex SVs, resulting in some potentially pathogenic variants being missed, while traditional functional assessment methods mainly focus on coding regions and insufficiently consider the impact of SVs on regulatory elements. In addition, genetic studies of SCZ mainly focus on European populations, and there is less research on SVs in non-European populations, affecting the wide applicability of research results.

[0013] Preferably, extract DNA from peripheral blood samples of schizophrenia patients and perform whole-genome sequencing using third-generation sequencing technology, including:

[0014] Extract DNA from peripheral blood samples of schizophrenia patients; wherein, schizophrenia patients refer to patients clinically diagnosed with first-episode schizophrenia in the population and having no history of other neuropsychiatric or mental diseases, such as epilepsy or mental retardation.

[0015] Perform whole-genome sequencing using third-generation sequencing technology; wherein, the third-generation sequencing technology is: PacBio Continuous Long Reads (CLR);

[0016] Preferably, carry out multi-tool joint detection of SVs, integrate multi-tool and multi-sample SV results, and generate a high-confidence SV set, including:

[0017] S1: Use the NGMLR tool to align long reads with the GRCh37 human reference genome, and then use cuteSV, pbsv, and Sniffles to detect SVs on the alignment results respectively;

[0018] S2: Use the Jasmine tool to merge SVs detected by multiple tools on the same sample, and retain high-confidence SVs supported by at least two tools;

[0019] S3: Use the Jasmine tool to merge SVs from multiple samples to obtain high-confidence SVs at the population level.

[0020] It should be noted that "long reads" refer to data fragments representing a short sequence of the original DNA or cDNA molecule generated after the sequencing process. Different from the shorter fragments (usually 100-300bp) generated by short-read sequencing (Short-Read Sequencing, SRS), long-read sequencing (Long-Read Sequencing, LRS) can generate continuous sequence fragments of thousands to millions of bases. These long reads retain more complete genomic structural information, giving them significant advantages in detecting complex structural variations (SVs), resolving repetitive sequences, haplotype assembly, epigenetic modification analysis, etc.

[0021] NGMLR is an alignment tool designed specifically for long-read sequencing data. It can handle long-read data with high error rates and effectively identify complex structural variations (SVs). By considering variation patterns such as insertions, deletions, and inversions, NGMLR achieves efficient and accurate sequence alignment across the entire genome, which is one of the preprocessing steps for structural variation analysis.

[0022] cuteSV, pbsv, and Sniffles are three structural variation (SV) detection tools developed specifically for long-read sequencing (PacBio / Nanopore) data. cuteSV adopts a strategy based on segment matching and soft clipping, featuring high efficiency and low false positives, and is suitable for large-scale SV analysis; pbsv is an official PacBio tool optimized for HiFi and CLR data, accurately detecting insertions and deletions; Sniffles identifies various SV types through alignment information, soft clipping, and alignment spacing, supporting single-sample and multi-sample analysis. These tools can be used in combination to improve the coverage and accuracy of SV detection.

[0023] Jasmine is a tool for SV integration and deduplication. It can aggregate SV results from different detection tools and merge them based on information such as coordinates, variation types, and sequence similarity to generate a high-confidence SV dataset.

[0024] Preferably, cross-queue comparison is used to screen for patient-specific SVs, and SVJudge is combined for prioritization to identify high-risk potential pathogenic SVs that are rare in the normal population and located in coding regions, promoters, and enhancers, including:

[0025] Using SVJudge, potential harmful SVs are identified based on genomic annotation results and public population SV datasets, and the SVJudge scores are recorded:

[0026] S1: Analyze the location of the SV in the genome, determine its association with key functional regions, and establish the connection between the SV and the affected gene through genomic annotation or enhancer-promoter interaction information.

[0027] S2: Assign weights to functional regions. Evaluate whether the SV falls into CDS, UTR, promoter, enhancer, TAD boundary, chromatin open region, and assign different weights based on the importance of different regions. For coding region SVs, analyze whether they cause frameshift or premature termination codons;

[0028] S3: Population tolerance analysis. Calculate the enrichment degree of the SV in the normal population, use the gwRVIS framework to evaluate whether the SV is located in an evolutionarily conserved region, and combine the LOEUF value of gnomAD to evaluate the knockout sensitivity of the affected gene. If the SV is located in a low-tolerance region (low LOEUF), it may have high pathogenicity.

[0029] S4: Gene dosage sensitivity assessment, integrating HI (haploinsufficiency) and TS (triploid sensitivity) data to evaluate whether the SV is likely to cause dosage abnormalities;

[0030] S5: Population background correction, using the background incidence of SVs in the normal population to correct the tolerance weights of different functional regions.

[0031] S6: Re-filtering of repetitive regions. If the SV overlaps with the repetitive region by more than 50%, it is screened based on the length of the repeat unit. Even if the overlap threshold is not met, as long as the number of resulting repeat units is similar, it is considered a consistent SV;

[0032] S7: Calculate the final pathogenicity score of the SV, comprehensively considering factors such as the importance of the functional region, the frequency of the SV in the population, the tolerance of the SV in the population, and gene tolerance to comprehensively evaluate the pathogenicity of the SV and give a pathogenicity score.

[0033] S8: Use the ClinVar dataset to verify the accuracy of SVJudge and evaluate its prediction performance through AUC.

[0034] It should be noted that the key genomic functional regions include: CDS, UTR, promoter, enhancer, TAD boundary, and chromatin development region. CDS (Coding Sequence) is the region in a gene that can be transcribed and translated, determining the amino acid sequence of the protein; UTR (Untranslated Region) includes 5'UTR and 3'UTR, regulating the stability and translation efficiency of mRNA; the promoter is located upstream of the gene, containing the RNA polymerase binding site, determining the transcriptional initiation of the gene; the enhancer is a long-distance regulatory element that enhances gene expression by binding to transcription factors; the TAD boundary limits the interaction between genes and regulatory elements within chromatin, maintaining the spatial organizational structure of gene expression; the chromatin open region is the unpacked chromatin region that allows transcription factors and RNA polymerase to bind, regulating gene expression.

[0035] gwRVIS (Genome-wide Residual Variation Intolerance Score) is an indicator used to evaluate the tolerance of genomic regions to structural variations (SVs). This method calculates the variation enrichment degree of each genomic region based on population data and estimates the sensitivity of this region to SVs by adjusting the background mutation rate. Regions with lower gwRVIS values are generally more intolerant to variations, indicating that SVs in this region are more likely to be pathogenic. This indicator is commonly used for SV prioritization to help identify potential high-risk structural variations.

[0036] gnomAD (Genome Aggregation Database) is a large - scale human genome variation database that integrates variant data from multiple large - scale genome and exome sequencing projects, aiming to provide a comprehensive resource to explore the variation frequencies and patterns of the human genome; LOEUF (Loss - of - function Observed / Expected Upper bound Fraction) is an indicator provided by gnomAD to measure the tolerance of genes to loss - of - function mutations. Genes with low LOEUF values are sensitive to LoF variants and may cause diseases; genes with high values are tolerant to LoF variants and have less impact. This indicator is used to evaluate the potential impact of SVs (such as deletions) on gene function.

[0037] HI (Haploinsufficiency): Measures whether a gene is sensitive to copy number reduction (such as deletions). Genes with high HI may lead to loss of function and diseases in the single - copy state.

[0038] TS (Triplosensitivity): Evaluates the tolerance of a gene to copy number increase (such as duplication amplification). Genes with high TS may lead to abnormal expression and pathological changes in the multi - copy state.

[0039] ClinVar is a public clinical variation database maintained by NCBI, which stores and shares the clinical significance of gene variations and their associations with diseases. This database integrates variant data from clinical laboratories, research institutions, and database projects (such as OMIM, dbSNP, gnomAD). SVs are covered in this database and are classified as pathogenic, likely pathogenic, benign, likely benign, variant of uncertain significance (VUS), etc. according to their impacts on diseases. ClinVar is widely used in genetic research, disease diagnosis, and precision medicine and is an important reference database for variant pathogenicity assessment.

[0040] Preferably, by combining transcription factor binding analysis, SCZ drug target data, and tissue - cell - specific expression data, evaluate the gene regulatory effects of potential pathogenic SVs and their mechanisms of action in SCZ, including:

[0041] S1: Analyze the impact of SVs on transcription factor binding. Using the conserved transcription factor binding site data of the UCSC table browser and the PWM matrix of the CIS - BP database, calculate the change in transcription factor binding strength before and after SV. Use TFM - PVALUE to evaluate the disruption or enhancement of binding sites and screen for significantly affected SVs.

[0042] S2: Evaluate the impact of SV in combination with SCZ drug target data. Integrate the SCZ drug target information in DrugBank and CTD databases, screen target genes and their regulatory factors affected by SV, and evaluate the potential impact of SV on SCZ drug targets.

[0043] S3: Evaluate SV regulatory effects based on tissue-cell-specific expression regulation data. Combined with the transcriptional regulatory networks of the hippocampus, adult brain, and fetal brain, analyze the tissue-specific regulatory relationships of SV-affected genes. At the same time, using PsychENCODE single-cell transcriptome data, evaluate the expression changes of these genes in different brain cell types (such as neurons, astrocytes, and oligodendrocytes).

[0044] It should be noted that the UCSC Table Browser is a tool provided by the UCSC genome database that can query and extract genome annotation data, including genes, transcripts, regulatory elements, and variation information.

[0045] The CIS-BP database (Catalog of Inferred Sequence Binding Preferences) is a database that contains transcription factor binding sites (TFBS) and their position weight matrices (PWMs), which are used to predict the binding of transcription factors to DNA sequences.

[0046] TFM-PVALUE is a tool that calculates the effects of mutations on transcription factor binding sites based on the PWM matrix and can be used to assess the potential effects of SVs on transcriptional regulation.

[0047] DrugBank is a comprehensive database containing information on FDA-approved drugs, experimental drugs and their target genes, and is widely used in drug-gene interaction analysis.

[0048] PsychENCODE is a research project that integrates multi-omics data such as transcriptome, epigenetics and single-cell RNA sequencing related to human brain development and mental illness. It is often used to study the molecular mechanisms of neuropsychiatric diseases such as schizophrenia.

[0049] Preferably, SCZ risk genes are screened based on SVJudge scores and patient carrier status, including:

[0050] This method screens high-confidence SCZ risk genes based on the pathogenicity scores of potential pathogenic SVs and the frequency of patients carrying them.

[0051] S1: Calculate the impact score of each gene and add the pathogenicity scores of all potential pathogenic SVs carried by patients in the SCZ patient cohort to ensure that the contributions of all SVs affecting the gene are included;

[0052] S2: Set a screening threshold. Using the median plus the standard deviation of all gene scores as a benchmark, screen out the genes significantly affected by SV to highlight the genes with greater influence in the patient population.

[0053] S3: Only retain the high-scoring genes carried by at least two patients to exclude individual-specific effects and ensure higher reliability of the screened candidate SCZ risk genes.

[0054] Preferably, verify its genetic risk and potential pathogenic mechanisms through pathway enrichment, protein-protein interaction network, and gene function module analysis, including:

[0055] S1: Use the ToppGene and SynGO tools to perform functional enrichment on the candidate SCZ risk genes. If enrichment occurs in known SCZ-related pathways, it indicates that the candidate genes capture the biological mechanisms of SCZ.

[0056] S2: Establish a protein-protein interaction network of SCZ candidate genes based on the Human Gene Connectivity Database; if the newly discovered genes in this network have a more significant connection with known SCZ genes, it indicates that the genes newly discovered through SV pose a potential risk to SCZ.

[0057] S3: Use the Louvain method to detect gene modules in the protein-protein interaction network. If the gene modules are enriched in known SCZ-related pathways, it indicates that the candidate genes capture the biological mechanisms of SCZ.

[0058] It should be noted that ToppGene is a gene function analysis and candidate gene prioritization tool, which can be used for gene enrichment analysis, pathway analysis, protein-protein interaction analysis, and candidate gene screening, and is commonly used for the identification of disease-related genes.

[0059] SynGO is a gene ontology database focusing on synaptic function, which annotates genes related to synapses based on experimentally verified data and is commonly used in the research of nervous system diseases.

[0060] The aforementioned SCZ-related pathways include neurotransmitter pathways (glutamate, dopamine, GABAergic transmission), neurodevelopmental pathways (neuronal migration, synaptic plasticity), immune and inflammatory pathways, and metabolic and mitochondrial function pathways.

[0061] The known SCZ genes in the network refer to the genes that have been statistically associated with SCZ through previous large-scale genome-wide association studies (GWAS), differential expression analysis, differential methylation analysis, etc. New genes refer to the potentially related genes that have not been clearly associated with SCZ but have been identified in this study.

[0062] The Louvain method mentioned above is a community detection algorithm based on modularity optimization, which is used to discover sub-populations (communities) with high-density connections in complex networks. This method iteratively optimizes the modularity. First, it locally assigns nodes to the best community, and then merges communities to construct a new network until the modularity no longer improves significantly. The Louvain method is computationally efficient and applicable to large-scale network analysis, such as the detection of functional modules in gene co-expression networks and protein interaction networks.

[0063] Advantages of the present invention:

[0064] 1. Improve the accuracy of structural variant (SV) detection and overcome the limitations of short-read sequencing. By adopting high-precision sequencing technology and optimized analysis strategies, the present invention can detect complex structural variants more comprehensively, including insertions, deletions, inversions, and repeat expansions, etc. Especially in highly repetitive and structurally complex genomic regions, the accuracy of variant identification is significantly improved. This improvement helps to reduce false negatives and false positives, ensuring that key pathogenic variants will not be missed.

[0065] 2. Optimize the pathogenicity assessment of SVs and enhance the screening ability for high-risk variants. The present invention combines multi-dimensional genetic and functional information to systematically evaluate the potential pathogenicity of SVs, overcoming the limitations of traditional methods that only rely on coding region variants. By comprehensively analyzing the effects of SVs on gene function, regulatory elements, and genomic stability, the ability to prioritize pathogenic variants is improved, enabling high-risk SVs to be identified more accurately.

[0066] 3. Expand the functional analysis of SVs in SCZ and deeply explore their action mechanisms. Existing research mainly focuses on small-scale gene mutations, while the present invention integrates multi-omics data to systematically evaluate the potential impact of SVs on gene regulation, covering not only the coding region but also the functional variants in non-coding regulatory regions. This strategy can more comprehensively reveal the impact of SVs on the key molecular mechanisms of SCZ, providing a deeper perspective for understanding the genetic basis of the disease.

[0067] 4. Construct a systematic SCZ risk gene screening system to improve the biological credibility of candidate genes. By combining gene function, regulatory network, and disease pathway analysis, the present invention can more reliably screen high-confidence candidate genes related to SCZ. Compared with traditional research that only relies on individual variant analysis, this method can more effectively identify pathogenic genes with important biological significance. This screening system helps to promote the genetic risk assessment of SCZ and provides theoretical support for the development of potential biomarkers.

[0068] 5. Improve the applicability of genetic research on SCZ in the population and promote the development of precision medicine. Existing genetic research on SCZ mainly focuses on European populations, while the present invention can fill the gap in SCZ research for different population genetic backgrounds based on genetic data from non-European populations. This progress not only enhances the broad applicability of research results but also provides an important reference for future personalized medicine and precision disease prediction for different populations. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] The present invention will be further described below in conjunction with the drawings.

[0070] Figure 1 It is a schematic diagram of the method steps of an embodiment of the present invention;

[0071] Figure 2 It is the regulation intensity of the screened transcription factor-regulating potential pathogenic SVs in different tissues in an embodiment of the present invention;

[0072] Figure 3 It is the SCZ risk genes enriched in synaptic structures in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0073] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0074] Refer to Figure 1 , an embodiment of the first aspect of the present invention provides a method for identifying and functionally evaluating schizophrenia risk structural variations based on third-generation sequencing, including:

[0075] Step 1: Extract DNA from the peripheral blood samples of schizophrenia patients and perform whole-genome sequencing using third-generation sequencing technology.

[0076] Step 2: Conduct multi-tool joint detection of SVs, and integrate the SV results of multi-tools and multi-samples to generate a high-confidence SV set.

[0077] Step 3: Cross-cohort comparison to screen patient-specific SVs, combine with SVJudge for prioritization, identify high-risk potential pathogenic SVs, and record the SVJudge scores.

[0078] Step 4: Combine transcription factor binding analysis, SCZ drug target data, and tissue cell-specific expression data to evaluate the gene regulatory effects of potential pathogenic SVs and their mechanisms of action in SCZ.

[0079] Step 5: Screen SCZ risk genes based on SVJudge scores and patient carriage, and verify their genetic risks and pathogenic mechanisms through pathway enrichment, protein-protein interaction network, and functional module analysis.

[0080] Extract DNA from peripheral blood samples of schizophrenia patients and perform whole-genome sequencing using a third-generation sequencing platform, including:

[0081] S1: Extract DNA from peripheral blood samples of schizophrenia patients; among them, schizophrenia patients refer to patients in the Chinese population who are clinically diagnosed with first-episode schizophrenia and have no previous history of other neuropsychiatric or mental diseases, such as epilepsy or mental retardation.

[0082] S2: Perform whole-genome sequencing using third-generation sequencing technology; among them, the third-generation sequencing technology is: PacBio Continuous Long Reads (CLR);

[0083] Carry out multi-tool joint detection of SVs, and integrate multi-tool and multi-sample SV results to generate a high-confidence SV set, including:

[0084] S1: Use the NGMLR tool to align long reads with the GRCh37 human reference genome, and then use cuteSV, pbsv, and Sniffles to perform SVs detection on the alignment results respectively;

[0085] S2: Use the Jasmine tool to perform SV merging of multiple detection tools on the same sample, and retain high-confidence SVs supported by at least two tools;

[0086] S3: Use the Jasmine tool to perform multi-sample SV merging to obtain a population-level high-confidence SV set.

[0087] Cross-cohort comparison to screen patient-specific SVs, combine with SVJudge for prioritization, and identify high-risk potential pathogenic SVs, including:

[0088] Use SVJudge to identify potentially harmful SVs based on genomic annotation results and publicly available population SV datasets, and record the SVJudge scores:

[0089] S1: Analyze the location of SVs in the genome, determine their association with key functional regions, and establish the connection between SVs and affected genes through genomic annotation or enhancer-promoter interaction information.

[0090] S2: Functional region weight assignment. Evaluate whether the SV falls into CDS, UTR, promoter, enhancer, TAD boundary, chromatin open region, and assign different weights based on the importance of different regions. For SVs in the coding region, analyze whether they cause frameshift or premature termination codons;

[0091] S3: Population tolerance analysis. Calculate the enrichment degree of SVs in the normal population, use the gwRVIS framework to evaluate whether the SV is located in an evolutionarily conserved region, and combine the LOEUF value of gnomAD to evaluate the knockout sensitivity of affected genes. If the SV is located in a low-tolerance region (low LOEUF), it may have high pathogenicity.

[0092] S4: Gene dosage sensitivity assessment. Integrate HI (haploinsufficiency) and TS (triploid sensitivity) data to evaluate whether the SV may cause dosage abnormalities;

[0093] S5: Population background correction. Use the background incidence of SVs in the normal population to correct the tolerance weights of different functional regions.

[0094] S6: Repeated region re-filtering. If the SV overlaps with the repeated region by more than 50%, then screen based on the length of the repeat unit. Even if the overlap threshold is not met, as long as the number of repeat units caused is similar, it is regarded as a consistent SV;

[0095] S7: Calculate the final pathogenicity score of the SV. Comprehensively consider factors such as the importance of functional regions, the frequency of SVs in the population, the tolerance of SVs in the population, and gene tolerance to comprehensively evaluate the pathogenicity of SVs and give a pathogenicity score.

[0096] S8: Use the ClinVar dataset to verify the accuracy of SVJudge and evaluate its prediction performance through AUC.

[0097] Combined with transcription factor binding analysis, SCZ drug target data, and tissue cell-specific expression data, evaluate the gene regulatory effects of potential pathogenic SVs and their mechanisms of action in SCZ, including:

[0098] S1: Analyze the impact of SVs on transcription factor binding. Use the conserved transcription factor binding site data of the UCSC table browser and the PWM matrix of the CIS-BP database to calculate the change in transcription factor binding strength before and after the SV. Use TFM-PVALUE to evaluate the disruption or enhancement of binding sites and screen for SVs with significant effects.

[0099] S2: Evaluate the impact of SVs in combination with SCZ drug target data. Integrate the SCZ drug target information in the DrugBank and CTD databases, screen the target genes and their regulatory factors affected by the SV, and evaluate the potential impact of the SV on the drug action targets of SCZ.

[0100] S3: Evaluate the SV regulatory effect based on tissue cell-specific expression regulation data. Combine the transcriptional regulatory networks of the hippocampus, adult brain, and fetal brain to analyze the tissue-specific regulatory relationships of genes affected by SV. Meanwhile, use the PsychENCODE single-cell transcriptome data to evaluate the expression changes of these genes in different brain cell types (such as neurons, astrocytes, oligodendrocytes).

[0101] Screen SCZ risk genes based on the SVJudge score and patient carriage status, including:

[0102] This method screens high-confidence SCZ risk genes based on the pathogenicity score of SV and the patient carriage frequency.

[0103] S1: Calculate the impact score for each gene. Accumulate the pathogenicity scores of all SVs carried by patients in the SCZ patient cohort to ensure that all contributions of SVs affecting the gene are included;

[0104] S2: Set a screening threshold. Use the median plus standard deviation of all gene scores as a benchmark to screen out genes significantly affected by SV to highlight genes with greater influence in the patient population;

[0105] S3: Only retain high-scoring genes carried by at least two patients to exclude individual-specific effects and ensure higher reliability of the selected candidate SCZ risk genes.

[0106] Verify its genetic risk and potential pathogenic mechanisms through pathway enrichment, protein-protein interaction network, and gene function module analysis, including:

[0107] S1: Use the ToppGene and SynGO tools to perform functional enrichment on candidate SCZ risk genes. If they are enriched in known SCZ-related pathways, it indicates that the candidate genes capture the biological mechanisms of SCZ;

[0108] S2: Establish a protein-protein interaction network of SCZ candidate genes based on the Human Gene Connectivity Database; if the newly discovered genes in this network have a more significant connection with known SCZ genes, it indicates that the genes newly discovered through SV pose a potential risk to SCZ;

[0109] S3: Use the Louvain method to detect gene modules in the protein-protein interaction network. If the gene modules are enriched in known SCZ-related pathways, it indicates that the candidate genes capture the biological mechanisms of SCZ.

[0110] For example: This embodiment is called the ZIB_SCZ cohort, with a total of 141 schizophrenia patients, all clinically diagnosed as first-episode schizophrenia (or the clinical high-risk syndrome of psychosis). The inclusion criteria are as follows:

[0111] (1) Age, 16 - 40 years old (gender not limited);

[0112] (2) Without other neuropsychiatric disorders (mental retardation, epilepsy or substance dependence);

[0113] (3) Without a history of other neuropsychiatric system diseases;

[0114] (4) Without a history of epilepsy and family history;

[0115] (5) Without mental retardation;

[0116] All patients need to provide blood samples for blood routine and gene sequencing. The blood samples are collected by the hosting center, 5 mL per person.

[0117] Step 1: Extract DNA from the peripheral blood samples of schizophrenia patients and perform whole - genome sequencing using third - generation sequencing technology;

[0118] For example: Extract DNA sequencing samples from whole blood and purify them using a modified CTAB method. The sample quality is evaluated by agarose gel electrophoresis, Nanodrop 2000 spectrophotometer, and Qubit fluorometric dye method. Construct a 20 - kb SMRTbell library using ExpressTemplate Prep Kit 2.0 (PacBio, #100 - 938 - 900), and perform preliminary quantification by Qubit 2.0. Then, verify the fragment distribution using pulsed - field electrophoresis to ensure that the library insert fragment size is approximately 20 kb, the concentration is not less than 70 ng / μL, and the total amount is not less than 7 μg. Sequencing is performed on the PacBio Sequel II platform using the CLR mode and diffusion loading method. The original data is initially Polymerreads. After quality filtering, sequences with a length less than 50 bp, a quality value lower than 0.8, and self - ligated or adapter sequences are removed, and finally Subreads are obtained as the subsequent analysis data.

[0119] Step 2: Conduct a multi - tool joint detection of SVs, integrate the SV results of multiple tools and multiple samples, and generate a high - confidence SV set.

[0120] For example: In this embodiment, a high - confidence SVs detection pipeline is designed to extract high - confidence structural variations.

[0121] First, long-read alignment was performed using NGMLR (version 0.2.7). The PacBio CLR sequencing data was aligned to the GRCh37 reference genome, and SAMtools was used to sort the alignment results to generate a BAM file. The alignment parameters were set according to the recommendations of PacBio CLR to optimize the alignment accuracy of long-read sequences.

[0122] Next, three SV detection tools were used for SV identification. The parameters of cuteSV (version 1.0.11) were set as follows: the minimum number of sequencing reads supporting the variant was 4, the maximum clustering deviation for insertion variants was set to 100 bases, the merging ratio for insertion variants was set to 0.3, the maximum clustering deviation for deletion variants was set to 200 bases, and the merging ratio for deletion variants was set to 0.5. The parameters of pbsv (version 2.4.1) included: the minimum number of sequencing reads supporting the variant in a single sample was 4, the minimum SV length was set to 30 bases, and the minimum number of sequencing reads supporting the breakpoints when calling structural variants across samples was 4. The parameters of Sniffles (version 1.0.12) were set as: the minimum number of sequencing reads supporting the variant was 4, and the minimum sequence length was set to 500 bases. All parameters not explicitly declared used the default settings.

[0123] The SV detection results were integrated by Jasmine (version 1.1.5). SVs of the same type with an overlap of more than 50% and a breakpoint distance within 1000 bases were defined as consensus SVs. The specific parameters included: the maximum linear distance ratio was set to 0.5, the maximum breakpoint distance was set to 1000 bases, the minimum overlap ratio was set to 50%, the plus / minus strand information was ignored, and the genotyping data was output. Finally, only the set of high-confidence SVs supported by at least two tools was retained, with the cuteSV results as the primary and pbsv as the secondary option because it has been verified to have high reliability at 15 - 20 times the PacBio sequencing depth.

[0124] When screening SVs, the following criteria were set:

[0125] (1) Supported by at least 4 sequencing reads to ensure the stability of the variant in the sequencing data.

[0126] (2) The lengths of insertion and deletion variants were less than 1 million bases to avoid detecting overly long variants that could lead to false positives.

[0127] (3) Variants within the mitotic regions or genomic gap regions annotated in the UCSC database were excluded to reduce noise and improve the specificity of the analysis.

[0128] (4) Variants in the GRCh37 that fell within regions of homology were excluded.

[0129] The median number of individual SVs finally identified in this example was 14,392, with the number ranging from 11,658 to 16,876. Deletions and insertions were the most common, accounting for 49.0% and 40.4% respectively. The length distribution of SVs was consistent with the pattern of publicly available third-generation sequencing data, and compared with data from different populations, the overlap rate with Chinese population data was the highest, showing obvious population specificity, further supporting the reliability of SV identification in this study.

[0130] Step 3: Cross-cohort comparison to screen for patient-specific SVs, combined with SVJudge for prioritization to identify high-risk potential pathogenic SVs;

[0131] For example: Use SVJudge to optimize the pathogenicity assessment of SVs and improve the screening ability for high-risk SCZ variants.

[0132] Based on the developed SVJudge, first annotate the SVs, evaluate the influence range in the coding region and regulatory region, and score them in combination with the distribution frequency of variants in the population to reduce errors caused by inaccurate breakpoints. Through an adjustable scoring matrix, preferentially screen for variants that affect low-tolerance exons, promoters, and enhancers, and focus on structural variants that may significantly affect gene expression.

[0133] The effectiveness of SVJudge was verified by the ClinVar dataset, and it performed better than other SV pathogenicity prediction tools (StrVCTVRE and CADD-SV), with an AUC reaching 0.89. Based on the prediction scores, high-risk insertions, deletions, duplications, and inversions were screened out, and different scoring thresholds were set to control the false positive rate not exceeding 0.1, and variants in translocation and genome instability regions were removed to reduce misjudgment.

[0134] Finally, a total of 352 high-confidence potential pathogenic structural variations were identified in this example. Among them, 95 were involved in tandem repeat expansions, and 257 were located in non-repetitive regions, affecting the coding regions, promoters, or enhancers of 502 genes in total. These variations were enriched in schizophrenia-related regulatory regions, indicating that they may affect disease susceptibility through gene regulation. Notably, most of the variations involving tandem repeat expansions were insertions, suggesting that they may affect gene function through the repeat expansion mechanism. For example, 5-HTTLPR VNTR in the promoter of the SLC6A4 gene, and the carrying frequency of its long allele in schizophrenia patients was significantly higher than that in other datasets. In addition, some variations affected the coding regions of known schizophrenia risk genes, such as the inversion in exons 23-25 of the CACNA1H gene, and the voltage-gated calcium channel encoded by this gene has been reported as a significant risk gene for schizophrenia. Other variations were involved in schizophrenia-related pathways such as neurogenesis, cell signal regulation, and cell adhesion, further supporting the role of the screened potential pathogenic SVs in schizophrenia.

[0135] Step 4: Combine transcription factor binding analysis, SCZ drug target data, and tissue cell-specific expression data to evaluate the gene regulatory effects of potential pathogenic SVs and their mechanisms of action in SCZ.

[0136] For example: Since a large number of potential pathogenic SVs fell into the non-coding region, to further explore their functional mechanisms, this example focused on analyzing their interference with transcription factor binding and the regulatory effects on SCZ-related genes. First, based on the transcription factor binding site data of the GRCh37 reference genome, combined with the information of 1,898 transcription factors in the CIS-BP database, the PWM (Position Weight Matrix) method was used to calculate the change in binding affinity of the sequence before and after mutation, and TFM-PVALUE was used to evaluate the significance. 157 SVs that disrupted transcription factor binding were screened out, among which 65 affected known SCZ-related genes or SCZ drug targets. Further analysis showed that 94.9% (149) of the SVs affected SCZ-related transcription factors or gene expression, and 101 of them met both conditions.

[0137] Refer to Figure 2 , in the hippocampus, adult brain, and fetal brain tissues, 51 genes affected by SVs formed 120 regulatory relationships, and the regulatory intensity was significantly higher than that in other tissues. By consulting single-cell expression data, it was found that these genes were significantly enriched in astrocytes of SCZ patients and oligodendrocytes of bipolar disorder patients, suggesting that they may affect synaptic signal transmission and nerve metabolism.

[0138] In addition, during the screening of drug targets for SCZ and related brain diseases, it was found that 7 affected transcription factor genes are approved drug targets for SCZ, 5 are drug targets for other mental diseases (such as personality disorders and bipolar disorder), and 21 genes regulated by SVs are treatment targets for SCZ. These results suggest that SVs may affect SCZ risk by regulating gene expression and provide potential targets for personalized treatment.

[0139] Step 5: Screen SCZ risk genes based on the SVJudge score and patient carriage status, and verify their genetic risks and pathogenic mechanisms through pathway enrichment, protein-protein interaction network, and functional module analysis.

[0140] For example: In this embodiment, SCZ risk genes driven by potentially pathogenic SVs are further screened.

[0141] First, genes affected by SVs that are above the set threshold of the SVJudge score and carried in at least two patients are selected. Ultimately, 130 SCZ risk genes are identified, of which 85 are affected by tandem repeat-related SVs and 45 are only affected by non-tandem repeat-related SVs. Based on pathway analysis, these genes are highly enriched in MHC-related cellular components, synaptic-related functions, and neural connection pathways, indicating their potential role in SCZ.

[0142] See Figure 3 , in the synaptic structure and function annotation of the SynGO database, 22 genes (14 of which are affected by tandem repeat-related structural variations) among the screened SCZ risk genes are enriched in synaptic-related pathways, mainly involving presynaptic processes (q = 1.63×10 -3 ) and synaptic vesicle cycling (p = 3.22×10 -3 ). In addition, there are significant functional associations between these genes and known SCZ susceptibility genes, especially in gene networks related to regulating synaptic function and neural development, further verifying their potential role in the pathogenesis of SCZ.

[0143] Subsequently, use the STRING database to construct the protein-protein interaction network (PPI) of these genes, and identify six significant functional modules, which cover biological processes such as cytoskeleton and neural tissue, stress response and neuronal apoptosis, immune response and cell adhesion, and calcium ion transport, which are known to be closely related to SCZ. In addition, two less studied modules involve cell division regulation, intracellular transport, and mitochondrial membrane structure, suggesting their possible role in SCZ.

[0144] Among these risk genes, some are located at known SCZ susceptibility loci, such as RAC1 and CREB5, which affect the development of neuronal dendritic spines and the regulation of the cAMP signaling pathway, respectively, and are associated with drugs already approved for the treatment of SCZ. Overall, the results of functional enrichment and network analysis of these genes indicate that they play important roles in the genetic mechanism of SCZ and may provide new targets to facilitate precision medicine research.

[0145] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0146] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for identifying structural variation and functional assessment of schizophrenia risk based on third-generation sequencing, characterized in that: include: Step 1: Extract DNA from peripheral blood samples of schizophrenia patients and perform whole genome sequencing using third-generation sequencing technology; Step 2: Conduct multi-tool joint SV detection and integrate multi-tool and multi-sample SV results to generate a high-confidence SV set; Step 3: Screen patient-specific SVs by cross-cohort comparison, prioritize them based on SVJudge, identify high-risk potential pathogenic SVs, and record SVJudge scores; Step 4: Combine transcription factor binding analysis, SCZ drug target data and tissue cell-specific expression data to evaluate the gene regulatory effects of potential pathogenic SVs and their mechanisms of action in SCZ; Step 5: Screen SCZ risk genes based on SVJudge scores and patient carrier status, and verify their genetic risks and pathogenic mechanisms through pathway enrichment, protein interaction network and functional module analysis.

2. The method for identification and functional assessment of schizophrenia risk structural variation based on third-generation sequencing according to claim 1, characterized in that: In step 1, the schizophrenia patients refer to patients who are clinically diagnosed with first-episode schizophrenia and have no history of other neuropsychiatric or mental illnesses; the third-generation sequencing technology is PacBio ContinuousLong Reads.

3. The method for identification and functional assessment of schizophrenia risk structural variation based on third-generation sequencing according to claim 1, characterized in that: The multi-tool joint detection of SV refers to using the NGMLR tool to align the long read segment with the GRCh37 human reference genome, and then using cuteSV, pbsv and Sniffles to compare the results and perform SVs detection.

4. The method for identification and functional assessment of schizophrenia risk structural variation based on third-generation sequencing according to claim 1, characterized in that: Integrating the SV results of multiple tools and multiple samples to generate a high-confidence SV set means using the Jasmine tool to merge the SVs of multiple detection tools for the same sample, and retaining the high-confidence SVs supported by at least two tools; The Jasmine tool was used to merge multi-sample SVs and obtain a high-confidence SV set at the population level.

5. The method for identification and functional assessment of schizophrenia risk structural variation based on third-generation sequencing according to claim 1, characterized in that: Step 3 includes: using SVJudge to identify potentially harmful SVs based on genome annotation results and public population SV datasets, and recording the SVJudge score: S1: Analyze the location of SV in the genome, determine its association with key functional regions, and establish the connection between SV and affected genes through genome annotation or enhancer-promoter interaction information; S2: Weight assignment of functional regions; evaluate whether the SV falls into CDS, UTR, promoter, enhancer, TAD boundary, chromatin open region, and assign different weights based on the importance of different regions; for coding region SV, analyze whether it causes frameshift or premature termination codon; S3: Population tolerance analysis: Calculate the enrichment of SV in the normal population, use the gwRVIS framework to assess whether the SV is located in the evolutionarily conserved region, and combine the LOEUF value of gnomAD to assess the knockout sensitivity of the affected gene; S4: Gene dosage sensitivity assessment: Integrate HI and TS data to assess whether SV causes dosage abnormalities; S5: Population background correction: Use the background incidence of SV in the normal population to correct the tolerance weights of different functional areas; S6: Repeat region re-filtering: SVs in tandem repeat and high repeat regions were annotated using UCSC Table Browser, RepeatMasker, and Tandem RepeatFinder; S7: Calculate the final pathogenicity score of SV, comprehensively consider the importance of the functional region, the frequency of SV in the population, the tolerance of SV in the population, and the genetic tolerance factors to comprehensively evaluate the pathogenicity of SV and give a pathogenicity score; S8: The accuracy of SVJudge was verified using the ClinVar dataset, and its prediction performance was evaluated by AUC.

6. The method for identification and functional assessment of schizophrenia risk structural variation based on third-generation sequencing according to claim 5, characterized in that: In S3, if the SV is located in the low tolerance region, it is pathogenic; In S6, if the SV overlaps with the repeat region by more than 50%, it is screened based on the length of the repeat unit. Even if the overlap threshold is not met, as long as the number of repeat units caused is similar, it is considered a consistent SV.

7. The method for identification and functional assessment of schizophrenia risk structural variation based on third-generation sequencing according to claim 1, characterized in that: The combined transcription factor binding analysis, SCZ drug target data and tissue cell-specific expression data are used to evaluate the gene regulatory effects of potential pathogenic SVs and their mechanisms of action in SCZ, including: S1: Analyze the effect of SV on transcription factor binding; use the conserved transcription factor binding site data of UCSC table browser and the PWM matrix of CIS-BP database to calculate the change of transcription factor binding strength before and after SV; use TFM-PVALUE to evaluate the destruction or enhancement of binding sites and screen SVs with significant effects; S2: Evaluate the impact of SV in combination with SCZ drug target data; integrate the SCZ drug target information in DrugBank and CTD databases, screen target genes and their regulatory factors affected by SV, and evaluate the potential impact of SV on SCZ drug targets; S3: Evaluate SV regulatory effects based on tissue-cell-specific expression regulation data; combine the transcriptional regulatory networks of the hippocampus, adult brain, and fetal brain to analyze the tissue-specific regulatory relationships of SV-affected genes; and use PsychENCODE single-cell transcriptome data to evaluate the expression changes of these genes in different brain cell types.

8. The method for identification and functional assessment of schizophrenia risk structural variation based on third-generation sequencing according to claim 1, characterized in that: The SCZ risk gene screening based on SVJudge score and patient carrier status includes: This method screens high-confidence SCZ risk genes based on the pathogenicity score of SV and the carrier frequency of patients; S1: Calculate the impact score of each gene and add up the pathogenicity scores of all SVs carried by patients in the SCZ patient cohort to ensure that the contributions of all SVs affecting the gene are included; S2: Set a screening threshold, using the median plus standard deviation of all gene scores as a benchmark to screen out genes significantly affected by SV to highlight genes with a large impact in the patient population; S3: Only high-scoring genes carried in at least two patients were retained to exclude individual-specific effects and ensure that the screened candidate SCZ risk genes had higher reliability.

9. The method for identification and functional assessment of schizophrenia risk structural variation based on third-generation sequencing according to claim 1, characterized in that: The genetic risk and potential pathogenic mechanism were verified through pathway enrichment, protein interaction network and gene function module analysis, including: S1: Functional enrichment of candidate SCZ risk genes using ToppGene and SynGO tools; S2: Establishment of protein-protein interaction network of SCZ candidate genes based on human gene connectome database; S3: Detect gene modules in protein interaction networks using the Louvain method.

10. The method for identification and functional assessment of schizophrenia risk structural variation based on third-generation sequencing according to claim 9, characterized in that: In S1, if it is enriched in known SCZ-related pathways, it means that the candidate gene captures the biological mechanism of SCZ; In S2, if the newly discovered genes in the network have more significant connections with known SCZ genes, it means that the newly discovered genes through SV have potential risks for SCZ; In S3, if the gene module is enriched in known SCZ-related pathways, it means that the candidate gene captures the biological mechanism of SCZ.

Citation Information

Patent Citations

  • Construction method and application of schizophrenia abnormal gene-metabolism regulation network

    CN115206420A

  • Screening method of sick sinus syndrome mutant gene

    CN116052780A

  • Rare variation driven Alzheimer disease new gene identification and function evaluation method

    CN119049545A

  • Method for estimating danger of diabetes typ B developed in the human species of Chinese bloodline and composition

    CN1496412A

  • Methods for identification of genes and genetic variants for complex phenotypes using single cell atlases and uses of the genes and variants thereof

    US20210071255A1

Cited By

  • Weighted voting-based cancer-related gene intelligent analysis method and apparatus, and medium

    CN121789766A

  • Screening method of metabolism-related fatty liver disease key transcription factors

    CN121983132A