A method, device, medium and program product for identifying disease risk elncRNA and risk genes
By integrating genomic and epigenomic data, identifying and predicting SCZ-related elncRNAs and their target genes, the insufficient recognition of the genetic mechanism of SCZ in existing technologies was solved, and auxiliary prediction of SCZ was achieved.
Patent Information
- Application Number
- CN202510525952.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-04-25
AI Technical Summary
In the existing technology, single nucleotide polymorphisms (SNPs) that identify the regulatory role of non-coding RNA (ncRNA) associated with schizophrenia (SCZ) have not been systematically identified and analyzed, leading to challenges in disease biology and clinical applications.
By integrating risk SNPs, eQTL data, and enhancer elements in genomic locations, combined with transcriptomics and epigenomic data, disease risk elncRNAs and their target genes were identified, and a SCZ prediction model was constructed.
The SCZ-related elncRNAs and their target genes were effectively identified and predicted, which improved our understanding of the genetic mechanism of SCZ and provided tools and methods to assist in the prediction of SCZ.
Smart Images

Figure CN120048345B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medicine, and more specifically, to a method, device, medium, and program product for identifying disease risk elncRNA and risk genes. Background Art
[0002] Although the specific etiology of schizophrenia (SCZ) remains incompletely understood, its high heritability (81%) suggests it may be a genetic disorder. Previous genome-wide association studies (GWAS) have identified numerous variants associated with SCZ. However, dissecting the functional roles of these genetic variants and translating them into disease biology and clinical applications remains challenging.
[0003] Most disease risk variants are located in noncoding regions rich in noncoding RNAs (ncRNAs) and cis-regulatory elements (CREs), suggesting their regulatory roles. Expression quantitative trait loci (eQTL) analysis provides a bridge between these risk single nucleotide polymorphisms (SNPs) and transcriptional regulation. eQTL data generated in larger-scale projects such as the BrainSeq Consortium have been widely used to elucidate the molecular mechanisms underlying genetic associations associated with SCZ. However, previous studies have focused solely on the association of SCZ-risk SNPs with protein-coding genes (PCGs), while overlooking SNPs with regulatory effects on noncoding RNAs. In recent years, researchers have discovered that ncRNAs are closely associated with SCZ. Therefore, identifying SNP-associated lncRNAs can enrich the mechanistic dissection and functional characterization of SCZ-risk SNPs.
[0004] A unique group of lncRNAs, termed enhancer-associated lncRNAs (elncRNAs), are transcribed from enhancer regions within the genome and represent 30–60% of all lncRNAs. Importantly, elncRNAs have been shown to indicate enhancer activity, with their expression levels positively correlated with the abundance of adjacent PCGs, highlighting the role of elncRNAs in gene regulation. Mechanistically, elncRNAs participate in gene regulation through multiple mechanisms, including recruiting transcription factors and chromatin-modifying enzymes, regulating RNA polymerase II activity, and interacting with chromatin structure. Furthermore, elncRNAs are closely associated with human diseases. The elncRNA DGCR5 has been reported to reside in regions of copy number variation (CNV) associated with SCZ risk and to help regulate the expression of several SCZ-associated genes. However, systematic identification of risk elncRNAs and dissection of their roles in SCZ have yet to be undertaken. Summary of the Invention
[0005] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention provides a pipeline, SCZ-elnc, to identify risk elncRNAs and their target genes driven by disease GWAS signals by integrating the following two steps: (1) combining risk SNPs, eQTL data, and enhancer elements in genomic locations to identify risk elncRNAs; and (2) predicting genes regulated by elncRNAs using multiple lines of supporting evidence from transcriptomics and epigenomic data.
[0006] In a first aspect, the present application discloses a method for identifying disease risk elncRNA, the method comprising:
[0007] S101, obtain several disease risk SNPs;
[0008] S102, identifying lncRNAs within a first threshold region upstream and / or downstream of a single risk SNP as LBG lncRNAs;
[0009] S103, screening lncRNAs whose expression is significantly correlated with the risk SNP (FDR<0.01) from the LBG lncRNAs, and defining them as eQTL lncRNAs;
[0010] S104, obtaining elncRNAs from disease-affecting tissues; elncRNAs from disease-affecting tissues are obtained by screening lncRNAs that overlap with active enhancer genomic regions located in disease-affecting tissues (in the case of SCZ, the disease-affecting tissue is brain tissue or cells);
[0011] S105, taking the intersection of the eQTL lncRNA and the elncRNA of the disease-occurring tissue to obtain the risk elncRNA.
[0012] In some embodiments, the disease-occurring tissue includes any one or more of the following: epithelial tissue, connective tissue, muscle tissue, and neural tissue; and the eQTL data comes from the corresponding disease-occurring tissue region.
[0013] The second aspect of the present application discloses a method for identifying disease risk genes, which is used to predict the disease risk gene of any single risk elncRNA described in the first aspect of the present application; the method comprises:
[0014] S201, obtaining known target genes of enhancers, and screening target genes of enhancers that can transcribe the single risk elncRNA as the first target genes of the single risk elncRNA;
[0015] S202, based on a single risk elncRNA, determining a CRE element contained in the genome of the single risk elncRNA; if the CRE element can form a loop structure with the promoter region of any gene, the gene is used as the second target gene;
[0016] S203, obtaining transcriptome data related to normal tissues, performing co-expression analysis on the transcriptome data, and screening genes whose correlation coefficient with a single risk elncRNA is greater than a second threshold (according to the top 100 with the largest correlation coefficient) as third target genes;
[0017] S204, if the single gene belongs to at least any two of the first target gene, the second target gene, and the third target gene, then the single gene is the disease risk gene;
[0018] In some embodiments, the co-expression analysis method includes Pearson and Spearman correlation methods.
[0019] The third aspect of the present application discloses a method for constructing an SCZ prediction model, the method comprising:
[0020] S301, obtaining the risk elncRNA expression data of brain regions of training set samples and the classification labels of the samples according to the method described in the first aspect of the present application;
[0021] S302: Input the brain region risk elncRNA expression data and classification labels into a machine learning model to obtain a predicted classification result, compare it with the classification label, and optimize the model according to the comparison result to obtain an SCZ prediction model.
[0022] A fourth aspect of the present application discloses a method for predicting SCZ, the method comprising:
[0023] S401, obtaining the expression data of the tested risk elncRNA;
[0024] S402 , inputting the expression data into the SCZ prediction model constructed by the method described in the third aspect of the present application to obtain an auxiliary prediction result of whether it is SCZ.
[0025] In some embodiments, the risk elncRNA includes any one or more of the following: ENSG00000269293, ENSG00000256028, ENSG00000204387, ENSG00000255571.
[0026] In a fifth aspect, the present application discloses a computer device, comprising: a memory and a processor; the memory is used to store a computer program; and the processor executes the computer program to implement the steps of the above method.
[0027] In a sixth aspect, the present application discloses a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above-mentioned method when the computer program is executed by a processor.
[0028] In a seventh aspect, the present application discloses a computer program product, comprising a computer program, which implements the steps of the above method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0030] Figure 1 This is a schematic diagram of the method flow provided by the first aspect of the embodiment of the present invention;
[0031] Figure 2 is a schematic flow chart of the method provided by the second aspect of an embodiment of the present invention;
[0032] Figure 3 is a schematic flow chart of a method provided in the third aspect of an embodiment of the present invention;
[0033] Figure 4 is a schematic flow chart of a method provided in the fourth aspect of an embodiment of the present invention;
[0034] Figure 5 is a schematic diagram of a computer device provided by an embodiment of the present invention;
[0035] Figure 6 is a schematic diagram of the architecture of an exemplary computing device provided by an embodiment of the present invention;
[0036] Figure 7 is a schematic diagram of a storage medium provided by an embodiment of the present invention;
[0037] Figure 8 Figure 1 is a schematic diagram of the SCZ-elnc process provided by an embodiment of the present invention; the upper light blue module is the first step, which identifies SCZ elncRNAs by combining SCZ GWAS, eQTL, and enhancer data. The lower light purple module is the second step, which predicts the target gene for each SCZ elncRNA by integrating known enhancer target genes, CRE-promoter loops, and correlation analysis between SCZ elncRNAs and target genes.
[0038] Figure 9 This is the verification of the SCZ elncRNA genomic characteristics provided by the embodiment of the present invention; wherein, Figure 9 a shows that SCZelncRNA is significantly enriched in known schizophrenia-related lncRNAs (Gandal's lncRNAs) and has a higher odds ratio (OR) compared with LBG lncRNAs, brain region elncRNAs, and eQTL lncRNAs. Enrichment analysis was performed using Fisher's exact test. The number of lncRNAs in the whole genome background is 19,955. Figure 9 b Based on Hi-C data from brain CP, SCZ elncRNA captured more CRE promoter loops compared with WBG lncRNA (candidate lncRNA in all reference genomes), LBG lncRNA, eQTL lncRNA, brain elncRNA and Gandal's lncRNA. Figure 9 c SCZel ncRNA was more likely to show DE (SCZ vs. control), and the P value of DE was used for comparison in the hippocampus. Figure 9 d shows that the chromatin region where the SCZ elncRNA is located is more open in the hippocampus. Comparisons were made using a one-sided Wilcoxon rank sum test. * indicates a P value < 0.05; ** indicates a P value < 0.01; *** indicates a P value < 0.001. ns indicates not significant. Boxplots show the median and the 25th and 75th percentiles; DE indicates differential expression.
[0039] Figure 10 The heritability analysis and tissue expression characteristics of SCZ elncRNA provided by the embodiment of the present invention; wherein, Figure 10 a Stratified LDSC to evaluate the enrichment of SCZ heritability explained by different groups of lncRNAs, the central value represents the enrichment, and the error bars represent the standard error. Figure 10 b Tissue specificity of SCZ elncRNAs across tissues in GTEx shows that SCZ elncRNAs are highly expressed in brain-related tissues compared with WBG lncRNAs. Figure 10 c is an analysis of the expression of lncRNAs in different developmental stages based on BrainSpan data, showing that the expression level of SCZ elncRNA in the prenatal stage is higher than that in the postnatal stage, and the expression level of SCZ elncRNA in the brain is higher than that of lncRNAs in other groups, which is consistent with the observation results based on GTEx data in b.
[0040] Figure 11 This is a potential diagram of SCZ elncRNA expression for predicting SCZ risk provided by the embodiment of the present invention. The result is a 10-fold cross-validation result of different models in the hippocampus; wherein, Figure 11 a is the receiver operating characteristic (ROC) curve of different models, Figure 11b is the accuracy of different models. DETAILED DESCRIPTION
[0041] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0042] In some of the processes described in the specification and claims of the present invention and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to be different types.
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0044] Figure 1 1 is a flow chart of a method for identifying disease risk elncRNAs provided by an embodiment of the present invention. Specifically, the method comprises the following steps:
[0045] S101, obtaining several disease risk SNPs; the risk SNPs are obtained based on GWAS data;
[0046] S102, identifying lncRNAs within the first threshold region upstream and / or downstream of a single risk SNP as LBG lncRNAs (local basic lncRNAs);
[0047] S103, screening out lncRNAs whose expression is significantly correlated with the risk SNP (FDR<0.01) from the LBG lncRNAs, defining them as eQTL lncRNAs; wherein eQTL (expression Quantitative Trait Loci) refers to genetic variation sites that affect gene expression levels.
[0048] S104, obtaining elncRNAs from the disease-affecting tissue; elncRNAs from the disease-affecting tissue are obtained by screening lncRNAs that overlap with active enhancer genomic regions located in the disease-affecting tissue (in the case of SCZ, the disease-affecting tissue is brain tissue or cells); active enhancers refer to enhancers that are active in brain tissue or cells; in some embodiments, the disease-affecting tissue includes any one or more of the following: epithelial tissue, connective tissue, muscle tissue, and neural tissue; and eQTL data is obtained from the corresponding disease-affecting tissue region. When the disease-affecting tissue is neural tissue of the brain, eQTL data is obtained from the hippocampus, and may also be obtained from the dorsolateral prefrontal cortex and caudate nucleus.
[0049] S105, taking the intersection of the eQTL lncRNA and the elncRNA of the disease-occurring tissue to obtain the risk elncRNA.
[0050] Figure 2 This is a flow chart of a method for identifying disease risk genes provided by an embodiment of the present invention. Specifically, genes regulated by elncRNA are predicted from evidence such as transcriptomics + epigenomic data. The method is used to predict the disease risk gene of any single risk elncRNA described in the first aspect of this application; the method comprises the following steps:
[0051] S201, obtaining known target genes of enhancers (active enhancers), and screening target genes of enhancers that can transcribe the single risk elncRNA as the first target genes of the single risk elncRNA; in some embodiments, the screening method is as follows Figure 8 The screening method for the dark green genes associated with elncRNA1 in the first row of the light purple blocks shown here specifically involves calculating the Pearson or Spearman correlation coefficient between a SCZ elncRNA and all genes in the Hippocampus brain region of normal human samples. If a gene's correlation coefficient ranks in the top 100 in either Pearson or Spearman ranking list, then the gene is considered associated with the SCZ elncRNA.
[0052] S202: Based on the single risk elncRNA, determine the CRE element contained in the genome of the single risk elncRNA. If the CRE element can form a loop structure with the promoter region of any gene, use the gene as the second target gene.
[0053] S203, obtaining transcriptome data related to normal tissue, performing co-expression analysis on the transcriptome data, and screening genes whose correlation coefficient with a single risk elncRNA is greater than a second threshold (according to the top 100 with the largest correlation coefficient) as third target genes; in some embodiments, the co-expression analysis method includes Pearson and Spearman correlation methods.
[0054] S204: If the single gene belongs to at least any two of the first target gene, the second target gene, and the third target gene, then the single gene is the disease risk gene.
[0055] Figure 3 1 is a flow chart of a method for constructing an SCZ prediction model provided by an embodiment of the present invention. Specifically, the method includes the following steps:
[0056] S301, obtaining the risk elncRNA expression data of brain regions of training set samples and the classification labels of the samples according to the method described in the first aspect of the present application;
[0057] S302, input the brain region risk elncRNA expression data and classification labels into a machine learning model to obtain a predicted classification result, compare it with the classification label, optimize the model according to the comparison result, and obtain an SCZ prediction model; optionally, the machine learning model includes any one or more of the following: support vector machine (SVM), extreme gradient boosting (XGBoost), k-nearest neighbor (KNN), logistic regression (LR), and random forest (RF). In this embodiment, SCZ risk elncRNA and SCZ elncRNA are the same concept.
[0058] Figure 4 : is a flow chart of a method for predicting SCZ provided by an embodiment of the present invention. Specifically, the method includes the following steps:
[0059] S401, obtaining the expression data of the tested risk elncRNA;
[0060] S402 , inputting the expression data into the SCZ prediction model constructed by the method described in the third aspect of the present application to obtain an auxiliary prediction result of whether it is SCZ.
[0061] In some embodiments, the risk elncRNA includes any one or more of the following: ENSG00000269293, ENSG00000256028, ENSG00000204387, ENSG00000255571.
[0062] In some embodiments, the terms "subject," "person to be tested," or "test sample" as used herein refer to any animal (e.g., mammal), including but not limited to humans, non-human primates, rodents, etc., that will be the recipient of a particular treatment. Generally, the terms "subject" and "patient" are used interchangeably herein when referring to a human subject. Preferably, the subject is a human. In some embodiments, the test sample is a patient undergoing a prognostic assessment in a clinical setting.
[0063] In some embodiments, the auxiliary prediction results include but are not limited to paper or electronic report forms. The results are only obtained by the intelligent machine based on the analysis of relevant data of the subject, and are only used as a reference for medical staff, and are not the final diagnosis results of the subject.
[0064] In some embodiments, the first threshold and / or the second threshold are obtained through training of training set samples, which may be specific thresholds or interval ranges. The specific form is not specifically limited in this embodiment.
[0065] Figure 5 is a schematic diagram of a computer device provided by an embodiment of the present invention, such as Figure 5 As shown, the device 2000 may include: one or more processors 2010, and one or more memories 2020; wherein the memories store computer-readable codes, and when the computer-readable codes are run by the one or more processors, they may execute the method described above.
[0066] The processor in this embodiment can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. It can implement or execute the various methods, operations, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor can be a microprocessor or any conventional processor, etc., and can be an X86 architecture or an ARM architecture.
[0067] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0068] For example, the method or apparatus according to the embodiment of the present disclosure may also be implemented by Figure 6 The architecture of the computing device 3000 shown in FIG. Figure 6 As shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to a network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for processing and / or communication of the method provided in the present disclosure, as well as program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 6 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Figure 6 One or more components of a computing device are shown.
[0069] The embodiment of the present invention further provides a computer-readable storage medium, such as Figure 7 As shown, it is a schematic diagram of a storage medium 4000 provided in an embodiment of the present invention, and computer-readable instructions 4010 are stored on the computer storage medium 4020. When the computer-readable instructions 4010 are executed by the processor, the method according to the embodiment of the present disclosure described with reference to the above figures can be executed. The computer-readable storage medium in the embodiment of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory. It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0070] The embodiments of the present disclosure further provide a computer program product or system, including a computer program, which implements the steps of the above method when executed by a processor.
[0071] In some embodiments, this embodiment further discloses a system for identifying risky elncRNAs, the system comprising:
[0072] A SNP acquisition module, used for or configured to acquire a plurality of disease risk SNPs;
[0073] an LBG lncRNA processing module, configured to or configured to identify lncRNAs within a first threshold region upstream and / or downstream of a single risk SNP as LBG lncRNAs;
[0074] an eQTL lncRNA processing module, used for or configured to screen out, from the LBG lncRNAs, lncRNAs whose expression is significantly correlated with the risk SNP (FDR<0.01), and define them as eQTL lncRNAs;
[0075] A brain elncRNA acquisition module, used or configured to acquire brain elncRNA; the brain elncRNA is obtained by screening lncRNAs that overlap with active enhancer genomic regions located in brain tissues or cells;
[0076] The risk elncRNA processing module is used for or configured to obtain the risk elncRNA by taking the intersection of the eQTL lncRNA and the brain elncRNA.
[0077] In some embodiments, this embodiment further discloses a disease risk gene identification system, which is used to predict the disease risk gene of any single risk elncRNA described in the first aspect of this application; the system comprises:
[0078] A first target gene determination module is used or configured to obtain known target genes of enhancers and screen first target genes associated with a single risk elncRNA;
[0079] The second target gene determination module is used or configured to determine, based on a single risk elncRNA, a CRE element contained in the genome of a single risk elncRNA; if the CRE element can form a loop structure with the promoter region of any gene, the gene is used as the second target gene;
[0080] A third target gene determination module is used or configured to obtain disease-related transcriptome data, perform co-expression analysis on the transcriptome data, and screen genes whose correlation coefficient with a single risk elncRNA is greater than a second threshold (100) as third target genes;
[0081] The disease risk gene determination module is used for or configured to determine that if a single gene belongs to at least any two of the first target gene, the second target gene, and the third target gene, the single gene is the disease risk gene.
[0082] In some embodiments, this embodiment further discloses a system for constructing an SCZ prediction model, the system comprising:
[0083] A training set data acquisition module, configured or configured to obtain the SCZ risk elncRNA expression data and classification labels of the training set samples according to the method described in the first aspect of the present application;
[0084] The SCZ prediction model training module is used or configured to input the SCZ risk elncRNA expression data and classification labels into a machine learning model to obtain a predicted classification result, compare it with the classification label, optimize the model according to the comparison result, and obtain an SCZ prediction model.
[0085] In some embodiments, this embodiment further discloses a SCZ prediction system, the system comprising:
[0086] A test data acquisition module, used for or configured to acquire the risk elncRNA of the test;
[0087] The auxiliary prediction result prediction module is used or configured to input the risk elncRNA of the subject into the SCZ prediction model constructed by the method described in the third aspect of the present application to obtain an auxiliary prediction result of whether it is SCZ. In some embodiments, the risk elncRNA includes any one or more of the following: ENSG00000269293, ENSG00000256028, ENSG00000204387, ENSG00000255571.
[0088] A specific embodiment is described using SCZ disease as an example:
[0089] method
[0090] SCZ-elnc pipeline description: The SCZ-elnc pipeline aims to identify SCZ elncRNAs and their target genes. First, we identify SCZ elncRNAs with a P value less than 5.0×10 from the SCZ GWAS results. -8SCZ risk SNPs. We collected lncRNAs in a 1Mb region centered on the risk SNP as candidates and named them local background lncRNAs (LBG lncRNAs). For LBG lncRNAs, we screened for lncRNAs that were significantly associated with SCZ risk SNPs (FDR < 0.01) and named them eQTL lncRNAs. Then, brain elncRNAs were identified based on data from the EnhancerAtlas 2.0 database, defined as lncRNAs that overlap with active enhancer genomic regions in brain tissue or cells. We retained the intersection of eQTL lncRNAs and brain elncRNAs as SCZ elncRNAs. Next, the SCZ-elnc pipeline predicted disease risk genes for SCZ elncRNAs by combining multiple lines of evidence. For each candidate gene, evidence was collected from the following sources: (1) it was a target gene of an enhancer that transcribed a related SCZ elncRNA in the EnhancerAtlas2.0 database; (2) it had a cis-regulatory element (CRE) in the genomic region of the SCZ elncRNA based on 3D genomics data (further detailed in the “CRE-promoter loop data collection” section); and (3) it ranked in the top 100 in the Pearson or Spearman correlation coefficient (R) with the related SCZ elncRNA in the co-expression analysis based on transcriptome data (see the “Gene expression analysis” section for more details). If a gene accumulated at least two pieces of evidence, it was considered a target gene of the SCZ elncRNA.
[0091] LncRNA and gene set enrichment analysis: We collected lncRNAs and gene sets from different sources with strong evidence of involvement in SCZ for lncRNA or gene set enrichment analysis. We performed enrichment analysis using Fisher's exact test.
[0092] The lncRNA set included the following: eQTL analysis results for three brain regions, the hippocampus, dorsolateral prefrontal cortex, and caudate nucleus, were derived from previous studies. We retained lncRNAs significantly associated with SCZ risk SNPs (FDR < 0.01) from these previous studies and designated them as eQTL lncRNAs. Using data from EnhancerAtlas 2.0, brain elncRNAs were identified as lncRNAs overlapping genomic regions of active brain enhancers. Previous studies have identified differentially expressed lncRNA isoforms in neuropsychiatric disorders, including ASD, BD, and SCZ. Here, we retained those lncRNAs specifically differentially expressed in SCZ as known SCZ-associated lncRNAs (Gandal's lncRNAs).
[0093] The gene sets included the following: SCZ GWAS genes from two previous studies, named 2014 GWAS genes and 2022 GWAS genes. SCZ priority genes and SCZ mutation (DNM) genes were collected from previous studies. Harmonizome collected SCZ genes extracted from the biomedical literature. Postsynaptic genes reported in previous studies included: PSD genes and genes related to postsynaptic proteins (Synaptome DB Postsynaptic), calcium channels and signaling (CCS) genes. Considering the shared pathophysiology between psychiatric disorders, we also collected two sets of autism spectrum disorder (ASD) genes, including evolutionarily constrained genes (ECGs) and essential genes. We performed Gene Ontology (GO) biological process and KEGG pathway enrichment analysis using Metascape (https: / / metascape.org) using default parameters.
[0094] Gene Expression Analysis: For differential expression and coexpression analysis, we utilized expression profiles from previous studies, including data obtained from RNA sequencing of samples from the hippocampus, DLPFC (dorsolateral prefrontal cortex), and caudate nucleus regions. These samples were from 133, 153, and 154 SCZ patients, as well as 314, 299, and 266 healthy controls. Differential expression analysis between SCZ and controls was performed using DESeq2. For coexpression analysis, we used Pearson and Spearman correlation methods, calculated in healthy control samples, to assess the coexpression of candidate PCGs with SCZ elncRNAs. For each SCZ elncRNA, we selected the top 100 PCGs with the highest correlation coefficient (R) based on Pearson or Spearman correlation as candidate target genes of the SCZ elncRNA.
[0095] For tissue-specific studies, we used GTEx version V8 data. We downloaded the lncRNA RPKM (reads per kilobase of transcript per million mapped reads) dataset from the GTEx website (https: / / www.gtexportal.org / home / datasets), which covers approximately 50 tissues. We used the Jensen-Shannon divergence to measure the tissue specificity of each lncRNA.
[0096] To conduct stage-specific studies of brain development, we downloaded RNA sequencing data of the developing human brain from BrainSpan (http: / / help.brain-map.org / display / devhumanbrain / Documentation). We then calculated the average expression of SCZ elncRNAs, eQTL lncRNAs, LBG lncRNAs, Gandal's lncRNAs, and WBG lncRNAs in all brain regions for each sample at each developmental stage using RPKM values.
[0097] CRE-promoter loop data collection: We collected CRE-promoter loops from multiple 3D genomic sources. A previous study inferred chromosome contacts by constructing Hi-C libraries from two major regions of the human cerebral cortex: the cortical and subcortical plates, and the germinal zone (referred to as the Brain CP and GZ). We downloaded the predicted CRE-promoter loops from that study, which yielded 221,069 loops in the Brain CP and 228,323 loops in the Brain GZ. Additionally, we obtained CRE-promoter loops from two other studies, one involving capture Hi-C analysis of the GM12878 cell line, and 1,618,000 predicted CRE-promoter loops from http: / / www.ebi.ac.uk / arrayexpress / experiments / E-MTAB-2323 / . Another dataset, from the FANTOM5 project, uses gene expression cap analysis to infer CRE-promoter loops in various human tissues, yielding 66,899 CRE-promoter loops (http: / / enhancer.binf.ku.dk / presets / ). Additionally, we downloaded Hi-C data from the hippocampus and dorsolateral prefrontal cortex (http: / / kobic.kr / 3div / download). We retained 9,186,925 CRE-promoter loops from the hippocampus and 9,669,639 CRE-promoter loops from the dorsolateral prefrontal cortex, with a "dist_foldchange" parameter of ≥ 2.
[0098] Chromatin accessibility data collection: We collected ATAC-seq data from a previous study, including three replicates of the hippocampus and caudate nucleus regions of the brain. For each lncRNA, we used the UCSC tool "bigwigAverageOverBed" to calculate the average signal from the TSS (transcription start site) to the TES (transcription end site) region.
[0099] Machine learning model construction: A model was developed using SCZ elncRNA to distinguish SCZ patients from healthy controls. Samples from the hippocampus were used as training data. The lncRNA expression values were Z-score transformed and used for model construction. We used five common machine learning algorithms, including support vector machine (SVM), extreme gradient boosting (XGBoost), k-nearest neighbor (KNN), logistic regression (LR), and random forest (RF) models. To evaluate the performance of these models, we performed 10-fold cross validation in the dataset and evaluated the performance by comparing the area under the curve (AUC) and accuracy values, as shown in the following figure. Figure 11 a and Figure 11 b. Specifically, we used the R package e1071 (V1.7-9) for the SVM model and xgboost (V1.0.6.1) for the XGBoost model. The KNN algorithm was implemented using the caret package (V6.0-93). Furthermore, the glmnet package (V4.1-4) was used to build the LR model, and the randomForest package (V4.6-12) was used for the RF model.
[0100] result
[0101] SCZ-elnc pipeline overview: The Psychiatric Genomics Consortium reported thousands of significant SNPs in the recent SCZ GWAS study. First, we identified 22,344 SNPs with a GWASP < 5.0 × 10 -8 We then identified 2,193 candidate lncRNAs (hereafter referred to as local background lncRNAs, LBG lncRNAs) located within a 1-Mb region centered on the risk SNP. Next, we collected hippocampal eQTL analysis results from previous studies. We retained 80 lncRNAs that were significantly associated with 4,172 SCZ risk SNPs (FDR < 0.01) and named them eQTL lncRNAs. We also collected active enhancer locations with tissue-specific information from EnhancerAtlas 2.0 and screened for lncRNAs that overlapped with active enhancer genomic regions in the brain, naming them Brain elncRNAs. Finally, through the intersection of eQTL lncRNAs and Brain elncRNAs, we identified 16 SCZ risk elncRNAs (SCZ elncRNAs).
[0102] To explore the biological functions of these SCZ elncRNAs, SCZ-elnc predicted the disease risk genes of SCZ elncRNAs using multiple supporting evidence from multi-omics data. Previous studies have reported that the production of elncRNAs is associated with the activity of related enhancers, and their expression levels are positively correlated with the abundance of neighboring PCGs. To identify the disease risk genes of each SCZ elncRNA, we integrated three supporting evidences, including: (1) known enhancer-target genes from the EnhancerAtlas database, (2) CRE-promoter loops between SCZ elncRNAs and disease risk genes from various 3D genomics datasets, and (3) co-expression analysis using transcriptome data. Genes supported by at least two pieces of evidence were predicted to be target genes of SCZ elncRNAs. The detailed information of the SCZ-elnc pipeline can be found in the Methods. The workflow of the SCZ-elnc pipeline is shown in Figure 2. Figure 8 shown.
[0103] Validation of SCZ elncRNAs: First, we verified the consistency of the SCZ elncRNAs identified using eQTL datasets from other brain regions (dorsolateral prefrontal cortex and caudate nucleus). Of the 16 SCZ elncRNAs, 13 were also present in the other two regions, while the remaining three were shared between the hippocampus and caudate nucleus. Given the current limited understanding of disease-associated lncRNAs in SCZ, a comprehensive genomic characterization of SCZ-associated lncRNAs has not yet been established. Gandal et al. previously reported lncRNAs associated with neuropsychiatric disorders, including those associated with SCZ, autism spectrum disorder (ASD), and bipolar disorder (BD). Specifically, these lncRNAs were identified based on transcriptome-wide isoform dysregulation in disease states, independent of genetic data. We isolated 473 SCZ-associated lncRNAs as a standard lncRNA set (Gandal's lncRNAs) for enrichment analysis of SCZ elncRNAs. Although LBG lncRNA, eQTL lncRNA, Brain elncRNA, and SCZ elncRNA all showed significant (P < 0.05) enrichment, SCZ elncRNA obtained the highest odds ratio (OR) value of 9.13 (e.g. Figure 9 a). These results show high confidence and regional consistency of SCZ elncRNAs.
[0104] In addition, our previous studies have shown that SCZ-risk PCGs have some characteristics, such as being connected to more incoming CREs and showing more significant differential expression compared with background levels (SCZ vs. control). In this study, we further investigated whether SCZ elncRNAs show similar patterns. We collected five Hi-C or capture Hi-C datasets to study the CRE promoter loops of SCZ elncRNAs, including the cerebral cortex and subcortical plate (Brain CP), brain germinal zone (BrainGZ), GM12878, hippocampus, and dorsolateral prefrontal cortex (DLPFC). We also obtained transcriptome data from three brain regions, hippocampus, DLPFC, and caudate nucleus, and performed differential expression analysis (SCZ vs. control). Compared with whole-genome background lncRNAs (WBGlncRNAs), LBG lncRNAs, eQTL lncRNAs, Brain elncRNAs, and Gandal's lncRNAs, we found that SCZelncRNAs were indeed connected to more CREs (such as Figure 9 b) and are more likely to show differential expression (e.g. Figure 9 c). In addition, we collected ATAC-seq data to investigate the open chromatin levels of SCZ elncRNAs, including three replicates each in the hippocampus and caudate nucleus regions of the brain. Our analysis showed that the chromatin regions where SCZ elncRNAs were located were more open compared to other lncRNA groups (e.g. Figure 9 d) The figure shows only one experimental replicate in the hippocampus. The results were replicated three times in the hippocampus and caudate nucleus, and the results showed consistent trends. In summary, these results demonstrate the effectiveness of the SCZ-elnc pipeline in identifying elncRNAs that confer SCZ disease risk. Due to space limitations, only the results for the hippocampus are shown in this example, but similar phenomena were observed in multiple other datasets.
[0105] SCZ elncRNAs explain higher SCZ heritability: We then used stratified linkage disequilibrium score regression (LDSC) to assess the heritability of SCZ explained by SCZ elncRNAs. We included SNPs located within a 20 kb window centered on the transcription start site (TSS) of each lncRNA in the LDSC analysis. We observed that SNPs located within a 20 kb window centered on the transcription start site (TSS) of each lncRNA were significantly associated with WBG lncRNA (enrichment = 1.45, P = 1.1 × 10 -3 ), LBG lncRNA (enrichment = 11.40, P = 3.4 × 10 -29 ), eQTL lncRNA (enrichment = 66.01, P = 8.2 × 10 -3 ), Brain elncRNA (enrichment = 1.76, P = 5.4 × 10 -6) and Gandal's lncRNA (enrichment = 3.06, P = 0.014), SCZ elncRNAs explained a higher disease heritability (enrichment = 104.77, P = 2.6 × 10 -3 )(like Figure 10 a). In addition, we evaluated the enrichment of eQTL SNPs associated with SCZ elncRNAs in SCZ risk SNPs contained within the CRE regions of these lncRNAs. Significant enrichment was observed in multiple datasets, including Brain CP (OR = 4.3, P < 2.2 × 10 -16 ), Brain GZ (OR=3.56, P=8.46×10 -16 ) and GM12878 (OR=14.78, P<2.2×10 -16 The above evaluations demonstrated that SCZ elncRNAs can better explain the heritability of SCZ.
[0106] Tissue-specific and developmental stage-specific expression of SCZ elncRNAs: We collected expression data from different tissues from the Genotype-Tissue Expression (GTEx) project and observed that SCZ elncRNAs had higher brain tissue specificity in the brain than WBG lncRNAs, LBG lncRNAs, eQTL lncRNAs, Brain elncRNAs, and Gandal's lncRNAs (e.g., Figure 10 b). In addition, we observed that the expression level of SCZ elncRNA in brain tissue was higher in the prenatal stage than in the postnatal stage (P<2.2×10 -16 ,like Figure 10 c), as shown in the BrainSpan data. In all prenatal stages, the expression levels of SCZ elncRNAs were consistently higher than those of WBG lncRNAs, LBG lncRNAs, eQTL lncRNAs, brain region elncRNAs, and Gandal's lncRNAs (e.g. Figure 10 c). These findings are consistent with the expression patterns of SCZ-risk PCGs, and these results highlight the patterning of SCZ elncRNAs in brain development, suggesting their potential key roles in the pathology involved in SCZ.
[0107] Potential for SCZ risk prediction based on SCZ elncRNA expression: Considering that SCZ elncRNAs can explain more SCZ heritability and are more likely to show differential expression (SCZ vs. controls), we explored whether they have the potential to distinguish SCZ patients from normal controls. We used a logistic regression model to calculate the Pr(>|z|) value of each SCZ elncRNA. Based on the threshold of Pr(>|z|)<0.1, four SCZ elncRNAs (ENSG00000269293, ENSG00000256028, ENSG00000204387, ENSG00000255571) were finally screened to construct a logistic regression model to predict the risk of SCZ. To evaluate the predictive performance of the model, we performed an internal 10-fold cross-validation to compare the area under the curve (AUC) and accuracy values of the model. The results showed that the AUC value was approximately 0.71 and the accuracy was approximately 71.8% ( Figure 11 a). We also compared with four other machine learning models, and the results showed that our model had the best AUC and accuracy ( Figure 11 b), illustrating the accuracy of our model. These results highlight the strong potential of the identified SCZ elncRNAs in predicting SCZ risk and assisting SCZ diagnosis.
[0108] Target gene validation of SCZ elncRNAs: To explore the biological functions of SCZ elncRNAs, we predicted the target genes of each SCZ elncRNA by combining enhancer-target gene, CRE-promoter loop and co-expression data. The number of genes regulated by each SCZ elncRNA ranged from 1 to 35. SCZ elncRNA SNHG32 regulated the most genes (35), while LINC00862 and ENSG00000253553 regulated only one gene. Among all the target genes, a considerable number were significantly (P < 0.05) differentially expressed in the hippocampus (75, 72.1%), DLPFC (39, 37.5%) and caudate nucleus (26, 25%) of SCZ patients. To prove that these genes contain true SCZ risk genes, we evaluated the enrichment of target genes in ten gene sets that have been widely and repeatedly associated with SCZ (see Methods for details). We observed significant enrichment of seven gene sets (P < 0.05), all of which showed significantly enhanced ORs (OR > 1), including risk genes identified by the 2014 and 2022 SCZ GWAS, text mining SCZ genes, postsynaptic density (PSD), SCZ priority genes, evolutionarily constrained genes (ECG), and SynaptomeDB postsynaptic. We then performed enrichment analysis of GO biological processes and KEGG pathways and found that several neurological and immune pathways were significantly enriched, such as the "glial cell differentiation" function and the "antigen processing and presentation" pathway.
[0109] Some of these target genes are established SCZ genes or potential candidate genes involved in the major functional categories of SCZ derived from the aforementioned gene set. Some target genes, such as MAPT, ULK2, and GABBR1, are involved in neurodevelopment and neuronal function regulation. MAPT encodes the microtubule-associated protein tau, whose expression in the nervous system varies depending on the stage of neuronal maturation and neuronal type. Mutations in the MAPT gene are associated with various neurodegenerative diseases, including AD and frontotemporal dementia. Recently, several studies have reported that frontotemporal dementia and SCZ co-occur due to MAPT variants. Furthermore, MAPT is associated with copy number variations in SCZ patients. Furthermore, MAPT is a neuronal marker gene and plays an important role in the etiology of SCZ. These findings further suggest that MAPT not only plays a crucial role in neurodegenerative diseases but is also closely associated with SCZ. Copy number variations (CNVs) in ULK family genes, such as ULK2, have been reported to be enriched in SCZ patients. ULK2 is involved in autophagy and can affect neuronal health. Reduced ULK2 gene expression leads to decreased autophagy, particularly in pyramidal neurons of the prefrontal cortex, resulting in an imbalance between excitatory and inhibitory neurotransmission, which may partly explain the sensorimotor gating defects and cognitive impairment associated with psychiatric disorders. GABBR1 encodes a GABAB receptor that is widely distributed in the brain and regulates neuronal network activity, neurodevelopment, and synaptic plasticity. Given its ubiquity and widespread distribution in the central nervous system, GABA B Receptor dysfunction has been implicated in various central nervous system disorders, including SCZ, major depressive disorder, and BD. Furthermore, several target genes associated with the extracellular matrix have been implicated, and abnormalities in these target genes may reflect core genetic features underlying the pathophysiology of SCZ, including NCAN and matrix metalloproteinases such as MMP16.
[0110] Furthermore, we found that some target genes coexisted within protein-protein interaction (PPI) modules, further suggesting their consistent functions in SCZ. PPI modules include "transcriptional and epigenetic regulation," "immune response and inflammatory response," and "cellular stress and protein folding." For example, the PPI module for three genes (GATAD2A, CSNK2B, and EHMT2) displayed biological functions in transcriptional and epigenetic regulation. GATAD2A is a transcriptional repressor involved in methylation-dependent gene silencing. It is preferentially expressed during fetal brain development and has been associated with SCZ due to its role in regulating open chromatin. Furthermore, recent studies have shown that p66α, the mouse homolog of GATAD2A, contributes to memory preservation by promoting persistent histone modifications in hippocampal neurons. CSNK2B has been identified as a potential SCZ risk gene. It encodes the β subunit of casein kinase II, a ubiquitous protein kinase with significant regulatory functions. Knockdown of CSNK2B enhances neural stem cell proliferation, inhibits differentiation, and alters neuronal morphology and synaptic transmission. These findings suggest a potential role for CSNK2B in the pathophysiology of SCZ and highlight its importance in regulating neurodevelopment and neuronal function. EHMT2 encodes a methyltransferase that methylates lysine residues on histone H3. Methylation of histone H3 at lysine 9 by EHMT2 facilitates the recruitment of additional epigenetic regulators, leading to transcriptional repression. EHMT2 has been implicated in ASD, AD, and PD, but its role in SCZ has not been reported. Given that the target gene EHMT2, GATAD2A protein, and CSNK2B coexist in the same PPI module, it is reasonable to infer that EHMT2 should also be a potential risk gene for SCZ.
[0111] In addition, we performed extensive manual review of all target genes. Approximately 80% of these genes (81 out of 104) were supported to be involved in the pathophysiology of SCZ to some extent. These results demonstrate the reliability of the target genes predicted by SCZ-elnc.
[0112] Case Study of the Regulatory Axis Mediated by the SCZ ElncRNA SNHG32: We further conducted a case study demonstrating novel regulatory patterns from GWAS data to SCZ elncRNAs and subsequently to target genes identified in SCZ-elncs. Given that the SCZ elncRNA SNHG32 (with Ensemble ID 'ENSG00000204387') is highly expressed in the hippocampus and DLPFC and has the largest number of predicted regulated target genes, we conducted in-depth analysis and validation of its regulatory role. SNHG32 expression was found to be regulated by SNP rs805825, a known SCZ risk variant consistently identified in previous GWAS studies and revealed by eQTL analysis in the hippocampus and DLPFC. By investigating the regulatory axis between rs805825 and SNHG32 and its target genes, we revealed complex interaction patterns using capture Hi-C genomic analysis of GM12878. The rs660550 locus showed interaction with the SNHG32 region, which interacts with multiple target genes including EHMT2. These findings highlight the complexity of this regulatory axis and suggest that complex pathological mechanisms are involved in SCZ.
[0113] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or operations, or can be implemented using a combination of dedicated hardware and computer instructions.
[0114] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the disclosed embodiments are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0115] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0116] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0117] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0118] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0119] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art will appreciate that various modifications and combinations may be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A method for identifying disease risk elncRNA, characterized in that: The method comprises: S101, obtain several disease risk SNPs; S102, identifying lncRNAs within a first threshold region upstream and / or downstream of a single risk SNP as LBG lncRNAs; S103, screening lncRNAs whose expression is significantly correlated with the risk SNP from the LBG lncRNAs, defining them as eQTL lncRNAs, wherein the eQTL data is obtained from the corresponding disease-occurring tissue region; S104, obtaining elncRNAs from disease-affected tissues; the elncRNAs from disease-affected tissues are obtained by screening lncRNAs that overlap with active enhancer genomic regions located in the disease-affected tissues; S105, taking the intersection of the eQTL lncRNA and the elncRNA of the disease-occurring tissue to obtain the risk elncRNA; the disease is schizophrenia.
2. A method for constructing a schizophrenia prediction model, characterized in that: The method comprises: S301, obtaining risk elncRNA expression data of brain regions of training set samples and classification labels of the samples according to the method of claim 1; S302: Input the brain region risk elncRNA expression data and classification labels into a machine learning model to obtain a predicted classification result, compare it with the classification label, optimize the model according to the comparison result, and obtain a schizophrenia prediction model.
3. A computer device, characterized in that: The device comprises: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 2.
4. A computer device, characterized in that: The device comprises: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to implement the following method steps: Obtain expression data of the tested risk elncRNAs; the risk elncRNAs are: ENSG00000269293, ENSG00000256028, ENSG00000204387, and ENSG00000255571; The expression data is input into the schizophrenia prediction model constructed by the method according to claim 2 to obtain an auxiliary prediction result of whether schizophrenia is present.
5. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 2 are implemented.
6. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the following method steps are implemented: Obtain expression data of the tested risk elncRNAs; the risk elncRNAs are: ENSG00000269293, ENSG00000256028, ENSG00000204387, and ENSG00000255571; The expression data is input into the schizophrenia prediction model constructed by the method according to claim 2 to obtain an auxiliary prediction result of whether schizophrenia is present.
7. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 2 are implemented.
8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the following method steps are implemented: Obtain expression data of the tested risk elncRNAs; the risk elncRNAs are: ENSG00000269293, ENSG00000256028, ENSG00000204387, and ENSG00000255571; The expression data is input into the schizophrenia prediction model constructed by the method according to claim 2 to obtain an auxiliary prediction result of whether schizophrenia is present.
Citation Information
Patent Citations
Application of SNP rs62065444 site as target in preparation of product for detecting and / or treating ovarian cancer
CN116676394A
Screening and identification method of mammalian enhancer lncRNA
CN117542411A