Disease risk elncRNA and risk gene identification method, equipment, medium and program product

By combining risk SNP, eQTL data and enhancer elements, the risk elncRNAs and target genes related to SCZ are identified and analyzed, and the problem of difficulty in systematic identification and analysis of these RNAs in the prior art is solved, a deeper understanding of SCZ biology and clinical applications is achieved, and a tool to assist in prediction is provided.

CN120048345AActive Publication Date: 2025-05-27THE FIRST AFFILIATED HOSPITAL OF FUJIAN MEDICAL UNIV

Patent Information

Application Number
CN202510525952.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-27
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The prior art is difficult to systematically identify and analyze risk elncRNAs and their roles associated with schizophrenia (SCZ), and fail to effectively utilize the regulatory role of non-coding RNAs in disease biology.

Method used

Provides a process SCZ-elnc that identifies risk elncRNAs and their target genes driven by disease GWAS signals by binding to risk SNPs, eQTL data and enhancer elements in genomic locations. This process includes obtaining disease risk SNPs, identifying lncRNAs associated with these SNPs, screening significantly related eQTL lncRNAs, obtaining elncRNAs in brain tissues and obtaining their intersections, and finally predicting genes regulated by elncRNA through multiple supporting evidence.

Benefits of technology

The risk elncRNA and its target genes related to SCZ are effectively identified and analyzed, enriching the mechanism anatomical and functional characterization of SCZ risk SNPs, improving the understanding of SCZ biology and clinical applications, and providing potential tools to assist SCZ prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048345A_ABST
    Figure CN120048345A_ABST
Patent Text Reader

Abstract

The invention provides a disease risk elncRNA and risk gene identification method, equipment, a medium and a program product, and relates to the field of intelligent medical treatment. The invention provides a process SCZ-elnc. Risk elncRNA driven by a disease GWAS signal and a target gene thereof are identified by integrating the following two steps. The method comprises the following steps: (1) identifying risk elncRNA by combining risk SNP, eQTL data and enhancer elements in a genome position; and (2) predicting the gene regulated by the elncRNA through a plurality of support evidences from transcriptomics and epigenomics data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent medicine, and more specifically, to a method, device, medium and program product for identifying disease risk elncRNA and risk genes. Background Art

[0002] Although the specific etiology of schizophrenia (SCZ) is not fully understood, the high heritability (81%) suggests that it may be a genetic disease. Previous genome-wide association studies (GWAS) have identified many variants associated with SCZ. Therefore, dissecting the functional roles of these genetic variants and translating them into disease biology and clinical applications remains challenging.

[0003] Most disease risk variants are located in non-coding regions rich in non-coding RNA (ncRNA) and cis-regulatory elements (CRE), indicating their regulatory roles. Expression quantitative trait locus (eQTL) analysis provides a bridge between these risk single nucleotide polymorphisms (SNPs) and transcriptional regulation. eQTL data generated in larger-scale projects such as the BrainSeq Consortium have been widely used to elucidate the potential molecular mechanisms underlying genetic associations with SCZ. However, previous studies have only focused on the associations between SCZ risk SNPs and protein-coding genes (PCGs), while ignoring SNPs that regulate non-coding RNA. In recent years, researchers have found that ncRNA is closely related to SCZ. Therefore, identifying lncRNAs associated with SNPs can enrich the mechanistic dissection and functional characterization of SCZ risk SNPs.

[0004] A unique group of lncRNAs is transcribed from enhancer regions in the genome and is called enhancer-related lncRNA (elncRNA), accounting for 30–60% of all lncRNAs. Importantly, elncRNA has been shown to indicate enhancer activity, and its expression level is positively correlated with the abundance of neighboring PCGs, highlighting the role of elncRNA in gene regulation. Mechanistically, elncRNA participates in gene regulation through multiple mechanisms, including recruiting transcription factors and chromatin-modifying enzymes, regulating the activity of RNA polymerase II, and interacting with chromatin structure; in addition, elncRNA is closely related to human diseases. It has been reported that elncRNA DGCR5 is located in the SCZ risk-related copy number variation (CNV) region and contributes to the regulation of the expression of several SCZ-related genes. However, risk elncRNAs have not been systematically identified and their roles in SCZ have not been dissected. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention provides a process SCZ-elnc for identifying risk elncRNAs driven by disease GWAS signals and their target genes by integrating the following two steps: (1) identifying risk elncRNAs by combining risk SNPs, eQTL data, and enhancer elements in genomic positions; (2) predicting genes regulated by elncRNAs through multiple lines of supporting evidence from transcriptomics and epigenomics data.

[0006] The first aspect of the present application discloses a method for identifying disease risk elncRNAs, the method comprising:

[0007] S101, obtaining several disease risk SNPs;

[0008] S102, identifying lncRNAs within the first threshold region upstream and / or downstream centered on a single risk SNP as LBG lncRNAs;

[0009] S103, screening out lncRNAs whose expression is significantly correlated with the risk SNP (FDR < 0.01) from the LBG lncRNAs, and defining them as eQTL lncRNAs;

[0010] S104, obtaining elncRNAs of disease-occurring tissues; the elncRNAs of disease-occurring tissues are obtained by screening lncRNAs that overlap with the active enhancer genomic regions located in disease-occurring tissues (if it is SCZ, the disease-occurring tissues are brain tissues or cells);

[0011] S105, taking the intersection of the eQTL lncRNAs and the elncRNAs of disease-occurring tissues to obtain the risk elncRNAs.

[0012] In some embodiments, the disease-occurring tissues include any one or more of the following: epithelial tissue, connective tissue, muscle tissue, nerve tissue; the eQTL data comes from the corresponding disease-occurring tissue regions.

[0013] The second aspect of the present application discloses a method for identifying disease risk genes, the method being used to predict the disease risk genes of any single risk elncRNA described in the first aspect of the present application; the method comprising:

[0014] S201, obtaining the known target genes of enhancers, and screening out the target genes of the enhancers that can transcribe the single risk elncRNA as the first target genes of the single risk elncRNA;

[0015] S202. Based on a single risk elncRNA, determine the CRE elements contained within the genome of the single risk elncRNA. If the CRE element can form a circular structure with the promoter region of any gene, then consider this gene as the second target gene.

[0016] S203. Obtain transcriptome data related to normal tissues, perform co-expression analysis on the transcriptome data, and screen genes with a correlation coefficient greater than a second threshold (the top 100 according to the largest correlation coefficient) with the single risk elncRNA as the third target genes.

[0017] S204. If a single gene belongs to at least any two of the first target gene, the second target gene, and the third target gene, then this single gene is the disease risk gene.

[0018] In some embodiments, the methods for co-expression analysis include Pearson and Spearman correlation methods.

[0019] The third aspect of the present application discloses a method for constructing an SCZ prediction model, the method comprising:

[0020] S301. Obtain the expression data of risk elncRNAs in the brain regions of the training set samples and the classification labels of the samples according to the method described in the first aspect of the present application.

[0021] S302. Input the expression data of the brain region risk elncRNAs and the classification labels into a machine learning model to obtain a predicted classification result, compare it with the classification label, and optimize the model according to the comparison result to obtain an SCZ prediction model.

[0022] The fourth aspect of the present application discloses a method for predicting SCZ, the method comprising:

[0023] S401. Obtain the expression data of the risk elncRNAs of the subject.

[0024] S402. Input the expression data into the SCZ prediction model constructed by the method described in the third aspect of the present application to obtain an auxiliary prediction result of whether it is SCZ.

[0025] In some embodiments, the risk elncRNAs include any one or more of the following: ENSG00000269293, ENSG00000256028, ENSG00000204387, ENSG00000255571.

[0026] The fifth aspect of the present application discloses a computer device, the device comprising: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the above method.

[0027] The sixth aspect of the present application discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned method are implemented.

[0028] The seventh aspect of the present application discloses a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method are implemented. Description of the Drawings

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0030] Figure 1 It is a schematic flowchart of the method provided in the first aspect of the embodiments of the present invention;

[0031] Figure 2 It is a schematic flowchart of the method provided in the second aspect of the embodiments of the present invention;

[0032] Figure 3 It is a schematic flowchart of the method provided in the third aspect of the embodiments of the present invention;

[0033] Figure 4 It is a schematic flowchart of the method provided in the fourth aspect of the embodiments of the present invention;

[0034] Figure 5 It is a schematic diagram of the computer device provided in the embodiments of the present invention;

[0035] Figure 6 It is a schematic diagram of the architecture of an exemplary computing device provided in the embodiments of the present invention;

[0036] Figure 7 It is a schematic diagram of the storage medium provided in the embodiments of the present invention;

[0037] Figure 8 It is a schematic flowchart of the SCZ-elnc process provided in the embodiments of the present invention; among them, the light blue module above is the first step, which identifies SCZ elncRNAs by combining SCZ GWAS, eQTL, and enhancer data. The light purple module below is the second step. For each SCZ elncRNA, the target genes are predicted by integrating known enhancer target genes, CRE-promoter loops, and correlation analysis between SCZ elncRNAs and target genes;

[0038] Figure 9 It is the verification of the genomic characteristics of SCZ elncRNAs provided in the embodiments of the present invention; among them,Figure 9 Panel a shows that SCZ elncRNA is significantly enriched among the known schizophrenia-related lncRNAs (Gandal’s lncRNA) and has a higher odds ratio (OR) compared to LBG lncRNA, brain-region elncRNA, and eQTL lncRNA. Enrichment analysis was performed using Fisher's exact test, and the number of lncRNAs in the genome-wide background was 19,955. Figure 9 Panel b shows that based on Hi-C data from brain CP, SCZ elncRNA captures more CRE promoter loops compared to WBG lncRNA (candidate lncRNAs in all reference genomes), LBG lncRNA, eQTL lncRNA, brain-region elncRNA, and Gandal’s lncRNA. Figure 9 Panel c shows that SCZ elncRNA is more likely to exhibit differential expression (DE) (SCZ vs. control), and the P values of DE were compared in the hippocampus. Figure 9 Panel d shows that the chromatin regions where SCZ elncRNA is located in the hippocampus are more open. The comparison was performed using the one-sided Wilcoxon rank-sum test. * represents P < 0.05; ** represents P < 0.01; *** represents P < 0.001. ns represents not significant. Box plots show the median and the 25th and 75th percentiles; DE indicates differential expression.

[0039] Figure 10 These are the heritability analysis and tissue expression characteristics of SCZ elncRNA provided by the embodiments of the present invention; among them, Figure 10 Panel a shows the stratified LDSC to evaluate the enrichment of SCZ heritability explained by different groups of lncRNAs. The central value represents enrichment, and the error bars represent standard errors. Figure 10 Panel b shows the tissue specificity of SCZ elncRNA across tissues in GTEx, indicating that SCZ elncRNA is highly expressed in brain-related tissues compared to WBG lncRNA. Figure 10 Panel c shows the analysis of the expression of lncRNAs in different developmental stage groups based on BrainSpan data, indicating that the expression level of SCZ elncRNA in the prenatal stage is higher than that in the postnatal stage, and the expression level of SCZ elncRNA in the brain is higher than that of other groups of lncRNAs, which is consistent with the observation results based on GTEx data in panel b.

[0040] Figure 11 These are the potential plots of using SCZ elncRNA expression to predict SCZ risk provided by the embodiments of the present invention. The results are the 10-fold cross-validation results of different models in the hippocampus; among them, Figure 11 Panel a shows the receiver operating characteristic (ROC) curves of different models. Figure 11b is the accuracy rate of different models. Detailed implementation manners

[0041] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0042] In some processes described in the specification, claims and above-mentioned drawings of the present invention, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0043] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0044] Figure 1 It is a schematic flowchart of a method for identifying disease-risk elncRNAs provided by an embodiment of the present invention. Specifically, the method includes the following steps:

[0045] S101, obtaining several disease-risk SNPs; the risk SNPs are obtained based on GWAS data;

[0046] S102, identifying lncRNAs within the first threshold region upstream and / or downstream centered on a single risk SNP as LBG lncRNAs (local basis lncRNAs);

[0047] S103, screening out lncRNAs whose expression is significantly correlated with the risk SNPs (FDR < 0.01) from the LBG lncRNAs, and defining them as eQTL lncRNAs; wherein, eQTL (expression Quantitative Trait Loci) refers to genetic variant sites that affect gene expression levels.

[0048] S104. Obtain the elncRNAs of the disease-occurring tissue; the elncRNAs of the disease-occurring tissue are obtained by screening the lncRNAs that overlap with the genomic regions of active enhancers located in the disease-occurring tissue (if it is SCZ, the disease-occurring tissue is brain tissue or cells); the active enhancers refer to the active enhancers in brain tissue or cells; in some embodiments, the disease-occurring tissue includes any one or more of the following: epithelial tissue, connective tissue, muscle tissue, nerve tissue; the eQTL data comes from the corresponding disease-occurring tissue region. When the disease-occurring tissue is the nerve tissue of the brain, the eQTL data comes from the hippocampal region, and can also come from the dorsolateral prefrontal cortex and caudate nucleus brain regions.

[0049] S105. Take the intersection of the eQTL lncRNAs and the elncRNAs of the disease-occurring tissue to obtain the risk elncRNAs.

[0050] Figure 2 It is a schematic flow chart of a method for identifying disease risk genes provided by an embodiment of the present invention. Specifically, genes regulated by elncRNAs are predicted from evidence such as transcriptomics + epigenomics data. The method is used to predict the disease risk genes of any single risk elncRNA described in the first aspect of this application; the method includes the following steps:

[0051] S201. Obtain the known target genes of the enhancer (active enhancer), and screen the target genes of the enhancer that can transcribe the single risk elncRNA as the first target gene of the single risk elncRNA; in some embodiments, the screening method is as Figure 8 shown in the screening method of the dark green gene related to elncRNA1 in the first row of the light purple square, specifically including: in the Hippocampus brain region of human normal samples, calculate the Pearson or Spearman correlation coefficient between a SCZ elncRNA and all genes. If the correlation coefficient of a gene ranks among the top 100 in any (Pearson or Spearman) sorted list, then this gene is considered to be related to this SCZ elncRNA.

[0052] S202. Based on the single risk elncRNA, determine the CRE elements contained within the genome of the single risk elncRNA. If the CRE element can form a circular structure with the promoter region of any gene, take this gene as the second target gene.

[0053] S203. Obtain the transcriptome data related to normal tissues, perform co-expression analysis on the transcriptome data, and screen out the genes with a correlation coefficient greater than the second threshold (the top 100 with the largest correlation coefficient) related to a single risk elncRNA as the third target gene. In some embodiments, the methods for co-expression analysis include Pearson and Spearman correlation methods.

[0054] S204. If a single gene belongs to at least any two of the first target gene, the second target gene, and the third target gene, then this single gene is the disease risk gene.

[0055] Figure 3 It is a schematic flowchart of a method for constructing an SCZ prediction model provided by an embodiment of the present invention. Specifically, the method includes the following steps:

[0056] S301. Obtain the expression data of risk elncRNAs in the brain regions of the training set samples and the classification labels of the samples according to the method described in the first aspect of the present application.

[0057] S302. Input the expression data of the brain region risk elncRNAs and the classification labels into a machine learning model to obtain a predicted classification result, compare it with the classification label, and optimize the model according to the comparison result to obtain an SCZ prediction model. Optionally, the machine learning model includes any one or more of the following: support vector machine (SVM), extreme gradient boosting (XGBoost), k-nearest neighbor (KNN), logistic regression (LR), and random forest (RF). In this embodiment, SCZ risk elncRNA and SCZ elncRNA are the same concept.

[0058] Figure 4 It is a schematic flowchart of a method for predicting SCZ provided by an embodiment of the present invention. Specifically, the method includes the following steps:

[0059] S401. Obtain the expression data of the risk elncRNAs of the subject.

[0060] S402. Input the expression data into the SCZ prediction model constructed by the method described in the third aspect of the present application to obtain an auxiliary prediction result of whether it is SCZ.

[0061] In some embodiments, the risk elncRNAs include any one or more of the following: ENSG00000269293, ENSG00000256028, ENSG00000204387, ENSG00000255571.

[0062] In some embodiments, the term "subject" or "test subject" or "test sample" as used herein refers to any animal (e.g., mammal), including but not limited to humans, non-human primates, rodents, etc., which will be the recipient of a particular treatment. Generally, the terms "subject" and "patient" are used interchangeably herein when referring to human subjects. Preferably, the subject is a human. In some embodiments, the test sample is a patient clinically used for prognostic evaluation.

[0063] In some embodiments, the auxiliary prediction results include but are not limited to the form of a paper or electronic report. This result is only obtained by the intelligent machine based on the relevant data of the subject and is only for reference by medical staff and does not serve as the final diagnosis result of the subject.

[0064] In some embodiments, the first threshold and / or the second threshold are obtained by training with training set samples. It can be a specific threshold or an interval range, and the specific form is not specifically limited in this embodiment.

[0065] Figure 5 is a schematic diagram of a computer device provided by an embodiment of the present invention, as Figure 5 shown, the device 2000 may include: one or more processors 2010, and one or more memories 2020; wherein, computer-readable code is stored in the memory, and when the computer-readable code is run by the one or more processors, the methods described above can be executed.

[0066] The processor in this embodiment may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, operations and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc., and may be of the X86 architecture or the ARM architecture.

[0067] Generally speaking, the various example embodiments of the present disclosure may be implemented in hardware or a dedicated circuit, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor or other computing device. When the aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts or using some other graphical representation, it will be understood that the blocks, devices, systems, technologies or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.

[0068] For example, the method or apparatus according to an embodiment of the present disclosure may also be implemented by means of Figure 6 the architecture of the computing device 3000 shown. As Figure 6 shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for the processing and / or communication of the method provided by the present disclosure and the program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 6 the architecture shown is only exemplary, and when implementing different devices, one or more components shown in the Figure 6 computing device may be omitted according to actual needs.

[0069] An embodiment of the present invention also provides a computer-readable storage medium. As Figure 7 shown, it is a schematic diagram of the storage medium 4000 provided by an embodiment of the present invention. A computer-readable instruction 4010 is stored on the computer storage medium 4020. When the computer-readable instruction 4010 runs on a processor, it can execute the method according to an embodiment of the present disclosure described with reference to the above drawings. The computer-readable storage medium in the embodiments of the present disclosure may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memories for the methods described herein are intended to include, but are not limited to, these and any other suitable types of memories. It should be noted that the memories for the methods described herein are intended to include, but are not limited to, these and any other suitable types of memories.

[0070] Embodiments of the present disclosure also provide a computer program product or system, including a computer program, which, when executed by a processor, implements the steps of the above method.

[0071] In some embodiments, this embodiment also discloses a system for identifying risk elncRNAs, the system including:

[0072] An SNP acquisition module, configured to acquire several disease risk SNPs.

[0073] An LBG lncRNA processing module, configured to identify lncRNAs within the first threshold regions upstream and / or downstream centered on a single risk SNP as LBG lncRNAs.

[0074] An eQTL lncRNA processing module, configured to screen out lncRNAs whose expression is significantly correlated (FDR < 0.01) with the risk SNP from the LBG lncRNAs, and define them as eQTL lncRNAs.

[0075] A brain elncRNA acquisition module, configured to acquire brain elncRNAs; the brain elncRNAs are obtained by screening lncRNAs that overlap with the genomic regions of active enhancers located in brain tissues or cells.

[0076] A risk elncRNA processing module, configured to take the intersection of the eQTL lncRNAs and the brain elncRNAs to obtain the risk elncRNAs.

[0077] In some embodiments, this embodiment also discloses a system for identifying disease risk genes, the system for predicting the disease risk genes of any single risk elncRNA described in the first aspect of the present application; the system includes:

[0078] A first target gene determination module, configured to acquire the known target genes of enhancers and screen out the first target genes related to a single risk elncRNA.

[0079] A second target gene determination module, configured to, based on a single risk elncRNA, determine the CRE elements contained within the genome of the single risk elncRNA, and if the CRE elements can form a circular structure with the promoter region of any gene, take this gene as the second target gene.

[0080] A third target gene determination module, configured to acquire disease-related transcriptome data, perform co-expression analysis on the transcriptome data, and screen out genes with a correlation coefficient greater than a second threshold (100) with a single risk elncRNA as the third target genes.

[0081] A disease risk gene determination module, which is used or configured to determine that a single gene is the disease risk gene if the single gene belongs to at least any two of the first target gene, the second target gene, and the third target gene.

[0082] In some embodiments, the present embodiment also discloses a system for constructing an SCZ prediction model, and the system includes:

[0083] A training set data acquisition module, which is used or configured to obtain the SCZ risk elncRNA expression data of the training set samples and the classification labels of the samples according to the method described in the first aspect of the present application;

[0084] An SCZ prediction model training module, which is used or configured to input the SCZ risk elncRNA expression data and the classification labels into a machine learning model, obtain a predicted classification result, compare it with the classification labels, and optimize the model according to the comparison result to obtain an SCZ prediction model.

[0085] In some embodiments, the present embodiment also discloses a prediction system for SCZ, and the system includes:

[0086] A subject data acquisition module, which is used or configured to obtain the risk elncRNA of the subject;

[0087] An auxiliary prediction result prediction module, which is used or configured to input the risk elncRNA of the subject into the SCZ prediction model constructed by the method described in the third aspect of the present application to obtain an auxiliary prediction result of whether it is SCZ. In some embodiments, the risk elncRNA includes any one or more of the following: ENSG00000269293, ENSG00000256028, ENSG00000204387, ENSG00000255571.

[0088] Specific embodiments are described by taking the SCZ disease as an example:

[0089] Method

[0090] SCZ-elnc process description: The SCZ-elnc process aims to identify SCZ elncRNAs and their target genes. First, identify from the SCZ GWAS results those with a P-value less than 5.0×10 -8SCZ-risk SNPs. We collected lncRNAs in the 1-Mb region centered on the risk SNPs as candidates and named them local background lncRNAs (LBG lncRNAs). For LBG lncRNAs, lncRNAs significantly associated with SCZ-risk SNPs (FDR < 0.01) were screened and named eQTL lncRNAs. Then, brain elncRNAs were determined based on data from the EnhancerAtlas 2.0 database, defined as lncRNAs overlapping with genomic regions of active enhancers in brain tissues or cells. We retained the intersection of eQTL lncRNAs and brain-region elncRNAs as SCZ elncRNAs. Next, the SCZ-elnc pipeline predicted the disease-risk genes of SCZ elncRNAs by integrating multiple lines of evidence. For each candidate gene, the evidence came from the following sources: (1) it is a target gene of an enhancer in the EnhancerAtlas 2.0 database, and the enhancer transcribes the relevant SCZ elncRNA; (2) based on 3D genomics data, it has cis-regulatory elements (CREs) in the genomic region of the SCZ elncRNA (further detailed in the "CRE-Promoter Loop Data Collection" section); (3) in co-expression analysis based on transcriptome data, its Pearson or Spearman correlation coefficient (R) with the relevant SCZ elncRNA ranks among the top 100 (more details can be found in the "Gene Expression Analysis" section). If a gene accumulates at least two lines of evidence, it is considered a target gene of the SCZ elncRNA.

[0091] lncRNA and gene set enrichment analysis: lncRNAs and gene sets from different sources with strong evidence of their involvement in SCZ were collected for lncRNA or gene set enrichment analysis. We used Fisher's exact test for enrichment analysis.

[0092] The lncRNA set included the following: The results of eQTL analysis in three brain regions, the hippocampus, dorsolateral prefrontal cortex, and caudate nucleus, were from previous studies. We retained lncRNAs significantly (FDR < 0.01) associated with SCZ-risk SNPs from the previous studies and named them eQTL lncRNAs. Using data from EnhancerAtlas 2.0, brain elncRNAs were identified as lncRNAs overlapping with genomic regions of active brain enhancers. Previous studies have identified lncRNA isoforms differentially expressed in neuropsychiatric diseases including ASD, BD, and SCZ. Here, those lncRNAs specifically differentially expressed in SCZ were retained as known SCZ-related lncRNAs (Gandal’s lncRNAs).

[0093] The gene set includes the following: SCZ GWAS genes from two previous studies, named 2014 GWAS genes and 2022 GWAS genes respectively. SCZ prioritized genes and SCZ mutation (DNM) genes were collected from previous studies. Harmonizome collected SCZ genes extracted from biomedical literature. Postsynaptic genes reported in previous studies include: PSD genes and genes related to postsynaptic proteins (Synaptome DB postsynaptic), calcium channels and signaling (CCS) genes. Considering the common pathophysiology among mental disorders, we also collected two sets of autism spectrum disorder (ASD) genes, including evolutionarily constrained genes (ECGs) and essential genes. We performed gene ontology (GO) biological process and KEGG pathway enrichment analysis through Metascape (https: / / metascape.org), using default parameters for the analysis.

[0094] Gene expression analysis: For differential expression and co-expression analysis, we utilized the expression profiles from previous studies, which included data obtained from RNA sequencing of samples from the hippocampus, DLPFC (dorsolateral prefrontal cortex), and caudate nucleus regions. These samples were from 133, 153, and 154 SCZ patients, and 314, 299, and 266 healthy controls. Differential expression analysis between SCZ and controls was performed using DESeq2. For co-expression analysis, we used Pearson and Spearman correlation methods, calculated in healthy control samples, to evaluate the co-expression of candidate PCGs with SCZ elncRNAs. For each SCZ elncRNA, we selected the top 100 PCGs with the highest correlation coefficient (R) based on Pearson or Spearman correlation as candidate target genes for the SCZ elncRNA.

[0095] For tissue-specific studies, we used data from GTEx version V8. We downloaded the lncRNA RPKM (reads per kilobase per million mapped reads) dataset from the GTEx website (https: / / www.gtexportal.org / home / datasets), covering approximately 50 tissues. We used Jensen Shannon divergence to measure the tissue specificity of each lncRNA.

[0096] To conduct stage-specific research on brain development, RNA sequencing data of the developing human brain was downloaded from BrainSpan (http: / / help.brain-map.org / display / devhumanbrain / Documentation). Then, we calculated the average expression of SCZ elncRNA, eQTLlncRNA, LBG lncRNA, Gandal’s lncRNA, and WBG lncRNA in all brain regions at each developmental stage for each sample using the RPKM value.

[0097] CRE-promoter loop data collection: CRE-promoter loops were collected from multiple 3D genome sources. Previous studies inferred chromosomal contacts by constructing Hi-C libraries in two major regions of the human cerebral cortex: the cortex and subplate, and the germinal zone (referred to as Brain CP and GZ). The predicted CRE-promoter loops in this study were downloaded. As a result, there were 221,069 loops in Brain CP and 228,323 loops in Brain GZ. In addition, we obtained CRE promoter loops from two other studies. One study involved performing capture Hi-C analysis on the GM12878 cell line, and we obtained 1,618,000 predicted CRE-promoter loops from http: / / www.ebi.ac.uk / arrayexpress / experiments / E-MTAB-2323 / . Another dataset came from the FANTOM5 project, which inferred CRE-promoter loops in multiple human tissues using cap analysis of gene expression technology, resulting in 66,899 CRE-promoter loops (http: / / enhancer.binf.ku.dk / presets / ). In addition, we downloaded Hi-C data of the hippocampus and dorsolateral prefrontal cortex (http: / / kobic.kr / 3div / download). We retained 9,186,925 CRE-promoter loops in the hippocampus and 9,669,639 CRE-promoter loops in the dorsolateral prefrontal cortex, where the "dist_foldchange" parameter was ≥2.

[0098] Chromatin accessibility data collection: We collected ATAC-seq from previous studies, including three replicates of the hippocampus and caudate nucleus regions of the brain. For each lncRNA, we calculated the average signal in the region from the TSS (transcription start site) to the TES (transcription termination site) using the UCSC tool "bigwigAverageOverBed".

[0099] Machine learning model construction: A model was developed using SCZ elncRNAs to distinguish SCZ patients from healthy controls. Samples from the hippocampal region were used as training data. The lncRNA expression values were transformed by Z-score and used for model construction. We employed five common machine learning algorithms, including support vector machine (SVM), extreme gradient boosting (XGBoost), k-nearest neighbor (KNN), logistic regression (LR), and random forest (RF) models. To evaluate the performance of these models, we performed 10-fold cross-validation in the dataset and evaluated the performance by comparing the area under the curve (AUC) and accuracy values, as specifically shown in Figure 11 a and Figure 11 b. Specifically, we used the R package e1071 (V1.7-9) to construct the SVM model and xgboost (V1.0.6.1) for the XGBoost model. The KNN algorithm was implemented using the caret package (V6.0-93). Additionally, the glmnet package (V4.1-4) was used to establish the LR model and the randomForest package (V4.6-12) for the RF model.

[0100] Results

[0101] Overview of the SCZ-elnc pipeline: The Psychiatric Genomics Consortium reported thousands of significant SNPs in recent SCZ GWAS studies. First, we identified 22,344 risk SNPs with a GWASP < 5.0×10 -8 . Then, we obtained 2,193 candidate lncRNAs (hereinafter referred to as local background lncRNAs, LBGlncRNAs), which were located within a 1-Mb region centered on the risk SNP. Next, we collected the results of hippocampal eQTL analysis from previous studies. We retained 80 lncRNAs significantly (FDR < 0.01) associated with 4,172 SCZ risk SNPs and named them eQTL lncRNAs. We also collected the positions of active enhancers with tissue-specific information from EnhancerAtlas 2.0 and screened lncRNAs overlapping with the genomic regions of active enhancers in the brain, naming them Brain elncRNAs. Finally, through the intersection of eQTL lncRNAs and Brain elncRNAs, 16 SCZ risk elncRNAs (SCZ elncRNAs) were identified.

[0102] To explore the biological functions of these SCZ elncRNAs, SCZ-elnc predicted the disease risk genes of SCZ elncRNAs through multiple lines of supporting evidence from multi-omics data. Previous studies have reported that the production of elncRNAs is related to the activity of associated enhancers, and their expression levels are positively correlated with the abundance of neighboring PCGs. To determine the disease risk genes of each SCZ elncRNA, we integrated three lines of supporting evidence, including: (1) known enhancer-target genes from the EnhancerAtlas database, (2) CRE-promoter loops between SCZ elncRNAs and disease risk genes from various 3D genomics datasets, and (3) co-expression analysis using transcriptome data. Genes supported by at least two lines of evidence were predicted as target genes of SCZ elncRNAs. Details of the SCZ-elnc pipeline can be found in the methods. The workflow of the SCZ-elnc pipeline is as Figure 8 shown.

[0103] Validation of SCZ elncRNAs: First, we verified the consistency of the identified SCZ elncRNAs when using eQTL datasets from other different brain regions (dorsolateral prefrontal cortex and caudate nucleus). Among the 16 SCZ elncRNAs, 13 were also present in the other two regions, while the remaining three were shared between the hippocampus and the caudate nucleus. Given the current limited understanding of disease-related lncRNAs in SCZ, a comprehensive characterization of the genomic features of SCZ-related lncRNAs has not been clearly established. Gandal et al. previously reported lncRNAs associated with neuropsychiatric diseases, including lncRNAs associated with SCZ, autism spectrum disorder (ASD), and bipolar disorder (BD). Specifically, these lncRNAs were identified based on transcriptome-wide isoform-level dysregulation in disease states, independent of genetic data. We isolated 473 SCZ-related lncRNAs as a standard lncRNA set (Gandal’s lncRNA) for enrichment analysis of SCZ elncRNAs. Although LBG lncRNAs, eQTL lncRNAs, Brain elncRNAs, and SCZ elncRNAs all showed significant (P<0.05) enrichment, SCZ elncRNAs obtained the highest odds ratio (OR) value of 9.13 (as Figure 9 shown in a). These results demonstrate the high confidence and regional consistency of SCZ elncRNAs.

[0104] In addition, our previous studies have shown that SCZ-risk PCGs have some characteristics, such as more incoming CRE connections and more significant differential expression compared to background levels (SCZ vs control). In this study, we further investigated whether SCZ elncRNAs exhibit similar patterns. We collected five Hi-C or Capture Hi-C datasets to study the CRE promoter loops of SCZ elncRNAs, including the cerebral cortex and subplate (Brain CP), the germinal zone of the brain (Brain GZ), GM12878, the hippocampus, and the dorsolateral prefrontal cortex (DLPFC). We also obtained transcriptome data from three brain regions, the hippocampus, DLPFC, and caudate nucleus, and performed differential expression analysis (SCZ vs control). Compared with whole-genome background lncRNAs (WBGlncRNAs), LBG lncRNAs, eQTL lncRNAs, Brain elncRNAs, and Gandal's lncRNAs, we found that SCZ elncRNAs are indeed connected to more CREs (as shown in Figure 9 b), and are more likely to exhibit differential expression (as shown in Figure 9 c). In addition, we collected ATAC-seq data to study the open chromatin levels of SCZ elncRNAs, including three replicates each from the hippocampus and caudate nucleus regions of the brain. Our analysis showed that the chromatin regions where SCZ elncRNAs are located are more open compared to other lncRNA groups (as shown in Figure 9 d). Only one experimental replicate from the hippocampus region is shown in the figure. This result was replicated three times each in the hippocampus and caudate nucleus, and the trend was consistent. In summary, these results demonstrate the effectiveness of the SCZ-elnc pipeline in identifying SCZ disease-risk elncRNAs. Due to space limitations, only the results from the hippocampus region are shown in this example, but the same phenomenon was also found in multiple other datasets.

[0105] SCZ elncRNAs can explain a higher SCZ heritability: Then, we used stratified linkage disequilibrium score regression (LDSC) to evaluate the SCZ heritability explained by SCZ elncRNAs. We included SNPs located within a 20-kb window centered on the transcription start site (TSS) of each lncRNA in the LDSC analysis. We observed that, compared with WBG lncRNAs (enrichment = 1.45, P = 1.1×10 -3 ), LBG lncRNAs (enrichment = 11.40, P = 3.4×10 -29 ), eQTL lncRNAs (enrichment = 66.01, P = 8.2×10 -3 ), Brain elncRNAs (enrichment = 1.76, P = 5.4×10 -6Compared with Gandal’s lncRNA (enrichment = 3.06, P = 0.014), SCZ elncRNA can explain higher disease heritability (enrichment = 104.77, P = 2.6×10 -3 ). (As shown in Figure 10 a). In addition, we evaluated the enrichment of eQTL SNPs associated with SCZ elncRNA in SCZ risk SNPs contained within the CRE regions of these lncRNAs. Significant enrichment was observed in multiple datasets, including Brain CP (OR = 4.3, P < 2.2×10 -16 ), Brain GZ (OR = 3.56, P = 8.46×10 -16 ), and GM12878 (OR = 14.78, P < 2.2×10 -16 ) datasets. The above evaluations demonstrated that SCZ elncRNA can well explain SCZ heritability.

[0106] Tissue-specific and developmental stage-specific expression of SCZ elncRNA: We collected expression data of different tissues from the Genotype-Tissue Expression (GTEx) project and observed that SCZ elncRNA has higher brain tissue specificity than WBG lncRNA, LBG lncRNA, eQTL lncRNA, Brain elncRNA, and Gandal’s lncRNA in the brain (as shown in Figure 10 b). In addition, we observed that the brain tissue expression level of SCZ elncRNA is higher in the prenatal stage than in the postnatal stage (P < 2.2×10 -16 ), as shown in Figure 10 c), as shown by BrainSpan data. At all prenatal stages, the expression level of SCZ elncRNA is consistently higher than that of WBG lncRNA, LBG lncRNA, eQTL lncRNA, brain region elncRNA, and Gandal’s lncRNA (as shown in Figure 10 c). These findings are consistent with the expression patterns of SCZ risk PCGs, and these results emphasize the pattern of SCZ elncRNA in brain development, indicating that it plays a potentially key role in the pathology involved in SCZ.

[0107] SCZ risk prediction potential based on SCZ elncRNA expression: Given that SCZ elncRNAs can explain more SCZ heritability and are more likely to exhibit differential expression (SCZ vs controls), we explored whether they have the potential to distinguish SCZ patients from normal controls. We used a logistic regression model to calculate the Pr(>|z|) value for each SCZ elncRNA. Four SCZ elncRNAs (ENSG00000269293, ENSG00000256028, ENSG00000204387, ENSG00000255571) were finally selected according to the threshold of Pr(>|z|) < 0.1 to construct a logistic regression model to predict the risk of developing SCZ. To evaluate the predictive performance of the model, we performed internal 10-fold cross-validation to compare the area under the curve (AUC) and accuracy values of the model. The results showed that the AUC value was approximately 0.71 and the accuracy was approximately 71.8%( Figure 11 a). We also compared with four other machine learning models, and the results showed that our model had the best AUC and accuracy( Figure 11 b), demonstrating the accuracy of our model. These results highlight the strong potential of the identified SCZ elncRNAs in predicting SCZ risk and assisting in SCZ diagnosis.

[0108] Target gene validation of SCZ elncRNAs: To explore the biological functions of SCZ elncRNAs, we predicted the target genes of each SCZ elncRNA by integrating enhancer-target gene, CRE-promoter loop, and co-expression data. The number of genes regulated by each SCZ elncRNA varied from 1 to 35. SCZ elncRNA SNHG32 regulated the most (35) genes, while LINC00862 and ENSG00000253553 regulated only one gene. Among all target genes, a significant portion was differentially expressed (P<0.05) in the hippocampus (75, 72.1%), DLPFC (39, 37.5%), and caudate nucleus regions (26, 25%) of SCZ patients. To demonstrate that these genes contain true SCZ risk genes, we evaluated the enrichment of target genes in ten gene sets that are widely and repeatedly associated with SCZ (see Methods for details). We observed significant enrichment (P<0.05) in seven gene sets, all of which showed significantly enhanced OR (OR>1), including risk genes identified by SCZ GWAS in 2014 and 2022, text-mined SCZ genes, postsynaptic density (PSD), SCZ-prioritized genes, evolutionarily constrained genes (ECGs), and SynaptomeDB postsynaptic. Then, we performed enrichment analysis of GO biological processes and KEGG pathways and found significant enrichment in some neurological functions and immune pathways, such as the "glial cell differentiation" function and the "antigen processing and presentation" pathway.

[0109] Some of the target genes are established SCZ genes or potential candidate genes that are involved in the major functional categories of SCZ derived from the above gene sets. Individual target genes are involved in neurodevelopment and the regulation of neuronal function, such as MAPT, ULK2, GABBR1, etc. MAPT encodes the microtubule-associated protein tau, and its expression in the nervous system varies depending on the neuronal maturation stage and neuronal type. MAPT gene mutations are associated with various neurodegenerative diseases such as AD and frontotemporal dementia. Recently, several studies have reported that frontotemporal dementia and SCZ occur simultaneously due to MAPT variations; in addition, MAPT is associated with copy number variations in SCZ patients. Moreover, MAPT is one of the neuronal marker genes and plays an important role in the etiology of SCZ. These findings further suggest that MAPT not only plays a crucial role in neurodegenerative diseases but is also closely related to SCZ. It has been reported that copy number variations (CNVs) of the ULK family genes (such as ULK2) are enriched in SCZ patients, and ULK2 is involved in the autophagy process and can affect neuronal health. Reduced expression of the ULK2 gene leads to weakened autophagy, especially in pyramidal neurons in the prefrontal cortex, resulting in an imbalance between excitatory and inhibitory neurotransmission, which may partly explain the emergence of sensorimotor gating deficits and cognitive impairments associated with mental disorders. GABBR1 encodes a GABAB receptor that is widely distributed in the brain and regulates neuronal network activity, neurodevelopment, and synaptic plasticity. Given its prevalence and widespread distribution in the central nervous system, GABA B receptor dysfunction is associated with various central nervous system diseases, including SCZ, major depressive disorder, and BD. In addition, some target genes related to the extracellular matrix are also involved, and abnormalities in these target genes may reflect the core genetic features underlying the pathophysiology of SCZ, including NCAN and matrix metalloproteinases such as MMP16.

[0110] In addition, we found that some target genes coexisted in protein-protein interaction (PPI) modules, further indicating their consistent functions in SCZ. The PPI modules included "transcription and epigenetic regulation", "immune response and inflammatory response", and "cellular stress and protein folding". For example, the PPI module of three genes (including GATAD2A protein, CSNK2B, and EHMT2) showed biological functions of transcription and epigenetic regulation. The GATAD2A protein is a transcriptional repressor involved in methylation-dependent gene silencing. It is preferentially expressed during fetal brain development and is associated with SCZ due to its role in open chromatin regulation. In addition, recent studies have shown that the mouse homolog p66α of the GATAD2A protein contributes to memory preservation by promoting persistent histone modifications in hippocampal neurons. CSNK2B has been identified as a potential SCZ risk gene. It encodes the β subunit of casein kinase II, a ubiquitously expressed protein kinase with significant regulatory functions. Knockdown of CSNK2B enhances neural stem cell proliferation, inhibits differentiation, and alters neuronal morphology and synaptic transmission. These findings suggest a potential role of CSNK2B in the pathophysiology of SCZ, highlighting its importance in neurodevelopment and neuronal function regulation. EHMT2 encodes a methyltransferase that methylates lysine residues on histone H3. Methylation of lysine 9 on histone H3 by EHMT2 facilitates the recruitment of additional epigenetic regulators, leading to transcriptional repression. EHMT2 has been confirmed to be associated with ASD, AD, and PD, but its role in SCZ has not been reported. Given that the target genes EHMT2, GATAD2A protein, and CSNK2B coexist in the same PPI module, it is reasonable to infer that EHMT2 should also be a potential risk gene for SCZ.

[0111] In addition, we conducted an extensive manual review of all target genes. The involvement of approximately 80% of these genes (81 out of 104) in the pathophysiological processes of SCZ was supported to some extent. These results demonstrated the reliability of the target genes predicted by SCZ-elnc.

[0112] Case study of the regulatory axis mediated by SCZ elncRNA SNHG32: We further conducted a case study to demonstrate a new regulatory pattern from GWAS data to SCZ elncRNA and then to the target genes identified for SCZ - elnc. Given that SCZ elncRNA SNHG32 (with Ensemble ID ‘ENSG00000204387’) showed high expression in the hippocampus and DLPFC and was predicted to regulate the largest number of target genes, we conducted in - depth analysis and validation of its regulatory role. The expression of SNHG32 was found to be regulated by SNP rs805825, a known SCZ risk variant consistently identified in previous GWAS studies, as revealed by eQTL analysis in the hippocampus and DLPFC. By studying the rs805825 - SNHG32 and target gene regulatory axis, we used capture Hi - C genomic analysis of GM12878 to reveal complex interaction patterns. The rs660550 locus showed interaction with the SNHG32 region, and SNHG32 interacted with multiple target genes including EHMT2. These findings highlight the complexity of this regulatory axis and suggest the involvement of complex pathological mechanisms in SCZ.

[0113] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware - based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0114] In general, the various example embodiments of the present disclosure may be implemented in hardware or in special-purpose circuits, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, special-purpose circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.

[0115] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0116] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods may be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other may be indirect couplings or communication connections through some interfaces, devices, or units, and may be in electrical, mechanical, or other forms.

[0117] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0118] In addition, the functional units in the various embodiments of the present invention may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.

[0119] The example embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art should understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.

Claims

1. A method for identifying disease risk elncRNA, characterized in that: The method comprises: S101, obtain several disease risk SNPs; S102, identifying lncRNAs in the upstream and / or downstream first threshold regions centered on a single risk SNP as LBG lncRNAs; S103, screening out lncRNAs whose expression is significantly correlated with the risk SNP from the LBG lncRNAs, defining them as eQTL lncRNAs; S104, obtaining elncRNA of the disease-occurring tissue; the elncRNA of the disease-occurring tissue is obtained by screening lncRNAs that overlap with active enhancer genomic regions located in the disease-occurring tissue; S105, taking the intersection of the eQTL lncRNA and the elncRNA of the disease-occurring tissue to obtain the risk elncRNA.

2. The method for identifying disease risk elncRNA according to claim 1, characterized in that: Disease-occurring tissues include any one or more of the following: epithelial tissue, connective tissue, muscle tissue, and neural tissue; The eQTL data come from the tissue region where the disease occurs.

3. A method for identifying disease risk genes, characterized in that: The method is used to predict the disease risk gene of any single risk elncRNA in claim 1 or 2; the method comprises: S201, obtaining known target genes of enhancers, and screening target genes of enhancers that can transcribe the single risk elncRNA as the first target genes of the single risk elncRNA; S202, based on a single risk elncRNA, determining a CRE element contained in the genome of the single risk elncRNA, and if the CRE element can form a loop structure with the promoter region of any gene, taking the gene as the second target gene; S203, obtaining transcriptome data related to normal tissues, performing co-expression analysis on the transcriptome data, and screening genes having a correlation coefficient with a single risk elncRNA greater than a second threshold as third target genes; S204: If the single gene belongs to at least any two of the first target gene, the second target gene and the third target gene, the single gene is the disease risk gene.

4. The method for identifying disease risk genes according to claim 3, characterized in that: The co-expression analysis methods include Pearson and Spearman correlation methods.

5. A method for constructing a SCZ prediction model, characterized in that: The method comprises: S301, obtaining the risk elncRNA expression data of brain regions of training set samples and classification labels of the samples according to the method of claim 1 or 2; S302, inputting the brain region risk elncRNA expression data and classification labels into a machine learning model to obtain a predicted classification result, comparing it with the classification label, optimizing the model according to the comparison result, and obtaining a SCZ prediction model.

6. A method for predicting SCZ, characterized in that: The method comprises: S401, obtaining the expression data of the tested risk elncRNA; S402, inputting the expression data into the SCZ prediction model constructed by the method described in claim 5 to obtain an auxiliary prediction result of whether it is SCZ.

7. The method for predicting SCZ according to claim 6, characterized in that: The risk elncRNA includes any one or more of the following: ENSG00000269293, ENSG00000256028, ENSG00000204387, and ENSG00000255571.

8. A computer device, characterized in that: The device comprises: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method and system for identifying non-coding RNA (lnRNA) regulated disease risk target pathways

    CN111899788A

  • Application of SNP rs62065444 site as target in preparation of product for detecting and / or treating ovarian cancer

    CN116676394A

  • Screening and identification method of mammalian enhancer lncRNA

    CN117542411A

  • Methods and compositions for altering function and structure of chromatin loops and / or domains

    US20180245079A1

Cited By

  • Tumor-driven IncRNA screening method and system based on interaction and co-expression characteristics of three-dimensional genome

    CN121075422A