Alzheimer's disease gene prediction method and device based on alternative splicing

By obtaining the expression levels of multiple transcripts in the gene signals of Alzheimer's disease samples, performing differential analysis, and constructing functional association networks, the problem of insufficient accuracy and reliability in gene prediction in existing technologies has been solved, achieving higher-precision Alzheimer's disease gene prediction.

CN122392640APending Publication Date: 2026-07-14CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610529005.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-21
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing methods for predicting Alzheimer's disease genes neglect the heterogeneity of features at different levels and cannot specifically handle different types of data distributions, resulting in insufficient accuracy and reliability.

Method used

By obtaining the expression levels of multiple transcripts in gene signals from Alzheimer's disease samples, differential analysis was performed, a functional association network was constructed, the hierarchical relationship between genes, transcripts, and alternative splicing events was explored, and a probabilistic model of genes and Alzheimer's disease was constructed.

Benefits of technology

It improves the accuracy and reliability of Alzheimer's disease gene prediction by using information-rich feature data to predict disease genes and enhance the precision of prediction probability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122392640A_ABST
    Figure CN122392640A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of gene analysis, and provides an Alzheimer disease gene prediction method and equipment based on alternative splicing, which comprises the following steps: performing difference analysis on the expression amount of all genes and the expression amount of transcripts to obtain expression amount difference characteristics, expression amount difference characteristics and expression proportion difference characteristics of the genes, and performing difference analysis on alternative splicing events of all sample gene signals to obtain alternative splicing difference characteristics of each gene; splicing the expression amount difference characteristics, the transcript expression amount difference characteristics, the expression proportion difference characteristics and the alternative splicing difference characteristics of each gene to obtain final characteristics of each gene; constructing a function correlation network according to the final characteristics of all genes; and performing gene disease prediction on the function correlation network to obtain the probability that each gene is related to Alzheimer disease. The method can improve the accuracy and reliability of Alzheimer disease gene prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of gene analysis technology, and in particular to a method and device for predicting Alzheimer's disease genes based on alternative splicing. Background Technology

[0002] The development of next-generation sequencing technology has generated massive amounts of RNA sequencing (RNA-seq) data, enabling researchers to quantitatively characterize gene expression levels. Gene expression information contains rich functional cues and has therefore been widely used in predicting pathogenic genes. Meanwhile, transcriptomics is also a crucial research area for elucidating gene regulatory mechanisms. Alternative splicing refers to the phenomenon during transcription where pre-mRNA generates multiple mature mRNAs through different splicing mechanisms. Existing research has shown that alternative splicing events are closely related to the occurrence and development of various diseases. Based on RNA-seq data, in addition to expression levels, transcriptomic information and alternative splicing profiling features such as the percentage of splitting in indices (PSI) can be further obtained. These indicators can characterize gene expression at a more refined level. A key prerequisite for predicting disease-causing genes within a supervised learning framework is constructing effective gene feature vectors. Existing methods typically utilize a combination of information, including protein-protein interaction networks, gene phenotypic information, protein sequence characteristics, and gene expression data, to characterize genes. However, research on the joint modeling and effective integration of gene expression features with multi-level features from alternative splicing omics remains relatively lacking. Traditional methods often simply concatenate multiple types of features before inputting them into the model. This approach has the following limitations: First, it ignores the heterogeneity of features at different levels and cannot specifically handle different types of data distributions; second, it applies the same implicit weights to all genes, failing to reflect the fact that different genes may depend on information from different levels. Therefore, existing methods still have shortcomings in terms of accuracy and reliability in predicting Alzheimer's disease genes. Summary of the Invention

[0003] This application provides a method and device for predicting Alzheimer's disease genes based on alternative splicing, which can solve the problems of low accuracy and reliability in predicting Alzheimer's disease genes.

[0004] In a first aspect, embodiments of this application provide an Alzheimer's disease gene prediction method based on alternative splicing, which includes: The expression levels of multiple transcripts in the gene signals of each sample of Alzheimer's disease were obtained, and the expression level of the gene corresponding to each transcript was obtained. Differential analysis was performed on the expression levels of all genes to obtain the expression difference characteristics of each gene. Differential analysis was also performed on the expression levels of all transcripts to obtain the expression difference characteristics of each transcript. Differential analysis was also performed on the expression ratios of all transcripts to obtain the expression ratio difference characteristics of each transcript. Finally, differential analysis was performed on the alternative splicing events of all sample gene signals to obtain the alternative splicing difference characteristics of each gene. The expression level difference features of each gene, the expression level difference features of the corresponding transcript, the expression level ratio difference features of the corresponding transcript, and the alternative splicing difference features are spliced ​​together to obtain the final features of each gene. A functional association network is constructed based on the final characteristics of all genes; in the functional association network, multiple nodes correspond one-to-one with multiple genes, and the edges between nodes represent the interaction relationships between the corresponding two genes. Genetic disease prediction is performed on functional association networks to obtain the probability that each gene is associated with Alzheimer's disease.

[0005] Optionally, differential analysis can be performed based on the expression levels of all genes to obtain the expression level difference characteristics of each gene; differential analysis can be performed on the expression levels of all transcripts to obtain the expression level difference characteristics of each transcript; and differential analysis can be performed on the expression ratios of all transcripts to obtain the expression ratio difference characteristics of each transcript, including: The expression levels of all transcripts are integrated into a matrix to obtain the transcript expression matrix; The expression level of each gene in each sample gene signal is calculated based on the expression levels of all transcripts, and the expression levels of all genes are integrated into a matrix to obtain the gene expression matrix. The transcript proportion expression matrix is ​​calculated based on the transcript expression matrix; the elements in the transcript proportion expression matrix are the proportion of transcripts expressed in the sample gene signals. Differential analysis was performed based on transcript expression matrix and gene expression matrix to obtain the differential expression characteristics of transcripts contained in each gene and the differential expression characteristics of the gene itself. Differential analysis was performed based on the transcript proportion expression matrix to obtain the differential expression characteristics of the transcripts contained in each gene.

[0006] Optionally, the expression level of each gene in each sample gene signal is calculated based on the expression levels of all transcripts, including: Through the formula:

[0007] Computational genes In sample gene signals Expression level in ; in, Transcript In sample gene signals The amount of expression in Indicates gene The corresponding set of transcripts; The transcript proportion expression matrix is ​​calculated based on the transcript expression matrix, including: Through the formula:

[0008] Calculate transcripts In sample gene signals The proportion of values ​​expressed in ; in, Indicates low expression of filtered transcripts In sample gene signals The amount of expression in Indicates gene The sum of the expression levels of the corresponding transcripts, ; All proportional expression values ​​are integrated into a matrix to obtain the transcript proportional expression matrix.

[0009] Optionally, the gene signals in the Alzheimer's disease samples can be categorized as disease-related or disease-unrelated. Differential analysis was performed based on the transcript expression matrix and gene expression matrix to obtain the differential expression characteristics of transcripts contained in each gene and the differential expression characteristics of the gene itself, including: Based on the gene expression matrix of each gene, calculate the expression level difference characteristics of each gene; Using transcript expression matrices and gene expression matrices, the differential expression values ​​for each transcript were calculated; For each gene, the differential expression values ​​of all transcripts corresponding to the gene are sorted from largest to smallest absolute value, and the differential expression values ​​of the top few transcripts in the sorting results are selected to constitute the gene expression level differential feature. Differential analysis was performed based on the transcript proportion expression matrix to obtain the differential expression proportion characteristics of transcripts contained in each gene, including: Based on the transcript proportion expression matrix, the mean proportion expression value of each transcript in disease-related sample gene signals is calculated, and the difference between the mean proportion expression value of each transcript in disease-unrelated sample gene signals is calculated, and the difference proportion value of transcripts is calculated. For each gene, the differential proportion values ​​of all transcripts corresponding to the gene are sorted from largest to smallest in absolute value, and the differential proportion values ​​of the top few transcripts in the sorting results are selected to constitute the expression proportion difference feature of the transcripts contained in the gene.

[0010] Optionally, differential analysis of alternative splicing events is performed on all sample gene signals to obtain the differential features of alternative splicing for each gene, including: Alternative splicing event analysis was performed on all sample gene signals to obtain the degree of splicing use of each alternative splicing event in the sample gene signals. For each variable splice event, calculate the splice usage difference value of the variable splice event based on all splice usage values ​​of the variable splice event; For each gene, the splicing use difference values ​​of all alternative splicing events contained in the gene are sorted from largest to smallest absolute value, and the splicing use difference values ​​of the top few alternative splicing events in the sorting results are selected to constitute the alternative splicing difference features of the gene.

[0011] Optionally, the expression level difference features of each gene, the expression level difference features of the corresponding transcript, the expression ratio difference features of the corresponding transcript, and the alternative splicing difference features are spliced ​​together to obtain the final features of each gene, including: Through the formula:

[0012] Genes were spliced ​​together The final feature ; in, Indicates gene The expression differences Indicates gene The differential expression characteristics of transcripts. Indicates gene The characteristics of differences in the expression ratio of transcripts, Indicates gene The variable shear differential characteristics.

[0013] Optionally, a functional association network is constructed based on the final characteristics of all genes, including: The final characteristics of each gene are encoded to obtain the gene representation of each gene; A corresponding node is generated for each gene, and the corresponding gene is used as the node attribute. When there is an interaction relationship between genes, an edge is generated between the two corresponding nodes to obtain the functional association network.

[0014] Optionally, gene disease prediction can be performed on the functional association network to obtain the probability that each gene is associated with Alzheimer's disease, including: Graph attention aggregation is performed on the functional association network to obtain the final node representation of each node; Based on the final node representation of each node, the probability that the corresponding gene is associated with Alzheimer's disease is calculated.

[0015] Secondly, embodiments of this application provide an Alzheimer's disease gene prediction device based on alternative splicing, comprising: The acquisition module is used to acquire the expression levels of multiple transcripts in the gene signals of each sample of Alzheimer's disease, and to acquire the expression level of the gene corresponding to each transcript; The differential analysis module is used to perform differential analysis based on the expression levels of all genes to obtain the differential expression characteristics of each gene, to perform differential analysis on the expression levels of all transcripts to obtain the differential expression characteristics of each transcript, to perform differential analysis on the expression ratios of all transcripts to obtain the differential expression ratios of each transcript, and to perform differential analysis on alternative splicing events of all sample gene signals to obtain the alternative splicing differential characteristics of each gene. The splicing module is used to splice the expression level difference features of each gene, the expression level difference features of the corresponding transcripts, the expression level ratio difference features of the corresponding transcripts, and the alternative splicing difference features to obtain the final features of each gene. The building module is used to construct a functional association network based on the final features of all genes; in the functional association network, multiple nodes correspond one-to-one with multiple genes, and the edges between nodes represent the interaction relationships between the corresponding two genes. The prediction module is used to predict genetic diseases in the functional association network and obtain the probability that each gene is associated with Alzheimer's disease.

[0016] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the aforementioned Alzheimer's disease gene prediction method based on alternative splicing.

[0017] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned Alzheimer's disease gene prediction method based on alternative splicing.

[0018] The above-mentioned solution in this application has the following beneficial effects: In the embodiments of this application, the expression levels of multiple transcripts in each sample gene signal of Alzheimer's disease are obtained, and the expression level corresponding to each transcript is obtained. Then, differential analysis is performed based on all expression levels to obtain the expression level difference characteristics of each gene. Differential analysis is performed on the expression levels of all transcripts to obtain the expression level difference characteristics of each transcript. Differential analysis is performed on the expression ratio of all transcripts to obtain the expression ratio difference characteristics of each transcript. Differential analysis of alternative splicing events is performed on all sample gene signals to obtain the alternative splicing difference characteristics of each gene. Then, the expression level difference characteristics of each gene, the expression level difference characteristics of the transcripts contained in the gene, the transcript expression ratio difference characteristics, and the alternative splicing difference characteristics are spliced ​​together to obtain the final characteristics of each gene. Then, a functional association network is constructed based on the final characteristics of all genes. Finally, gene disease prediction is performed on the functional association network to obtain the probability that each gene is associated with Alzheimer's disease. Among these methods, mining sample gene signals, transcripts, and hierarchical relationships between genes, and analyzing alternative splicing events, can fully study the characteristics of gene information at different levels, improve the richness and comprehensiveness of information in Alzheimer's disease gene prediction, and make disease gene prediction based on information-rich feature data, thereby increasing the accuracy of the predicted probability and improving the accuracy and reliability of Alzheimer's disease gene prediction.

[0019] Other beneficial effects of this application will be described in detail in the following detailed description section. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 A flowchart illustrating an Alzheimer's disease gene prediction method based on alternative splicing, provided as an embodiment of this application; Figure 2 A schematic diagram of the structure of an Alzheimer's disease gene prediction device based on alternative splicing provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0022] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0023] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0024] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0025] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0026] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0027] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0028] To address the low accuracy and reliability of existing Alzheimer's disease gene prediction methods, this application provides an Alzheimer's disease gene prediction method based on alternative splicing. This method obtains the expression levels of multiple transcripts in each sample gene signal for Alzheimer's disease, and acquires the expression level corresponding to each transcript. Then, it performs differential analysis on all expression levels to obtain the expression level difference characteristics of each gene. It further performs differential analysis on the expression levels of all transcripts to obtain the expression level difference characteristics of each transcript, and performs differential analysis on the expression ratios of all transcripts to obtain the expression ratio difference characteristics of each transcript. Finally, it performs differential analysis on alternative splicing events in all sample gene signals to obtain the alternative splicing difference characteristics of each gene. The expression level difference characteristics of each gene, the expression level difference characteristics of the transcripts contained in the gene, the transcript expression ratio difference characteristics, and the alternative splicing difference characteristics are then concatenated to obtain the final characteristics of each gene. A functional association network is then constructed based on the final characteristics of all genes. Finally, the functional association network is used to predict gene diseases, obtaining the probability that each gene is associated with Alzheimer's disease. Among these methods, mining sample gene signals, transcripts, and hierarchical relationships between genes, and analyzing alternative splicing events, can fully study the characteristics of gene information at different levels, improve the richness and comprehensiveness of information in Alzheimer's disease gene prediction, and make disease gene prediction based on information-rich feature data, thereby increasing the accuracy of the predicted probability and improving the accuracy and reliability of Alzheimer's disease gene prediction.

[0029] The Alzheimer's disease gene prediction method based on alternative splicing provided in this application will be illustrated by example below.

[0030] like Figure 1 As shown, the Alzheimer's disease gene prediction method based on alternative splicing provided in this application includes the following steps: Step 11: Obtain the expression levels of multiple transcripts in the gene signals of each sample of Alzheimer's disease, and obtain the expression level of the gene corresponding to each transcript.

[0031] The gene signals mentioned above are Alzheimer's disease-related RNA-seq samples. The transcripts mentioned above are mature mRNA molecules formed by gene transcription. The expression level of the transcript refers to the abundance of the transcript in the gene signal of the sample. The gene signals of the Alzheimer's disease sample are categorized as disease-related or disease-unrelated. If disease-related, the gene signals of the sample come from individuals with Alzheimer's disease; if disease-unrelated, the gene signals of the sample come from healthy individuals.

[0032] In some embodiments of this application, multiple sample gene signals of Alzheimer's disease, the expression levels of multiple transcripts in the sample gene signals, and the genes corresponding to the transcripts can be obtained by accessing public datasets.

[0033] For example, relevant datasets include: MayoRNAseq-CER, MayoRNAseq-TCX, MSBB-BM10, MSBB-BM22, MSBB-BM36, MSBB-BM44, GSE95587-FG, GSE125583-FG, GSE159699-LTL, GSE231341-iPSC, ROSMAP-B1-DLPFC, and ROSMAP-B3-DLPFC. The MayoRNAseq-CER and MayoRNAseq-TCX datasets are from the Mayo Clinic, containing 149 samples from the temporal cortex (TCX) and 94 samples from the cerebellum (CER), respectively. The MSBB-BM10, MSBB-BM22, MSBB-BM36, and MSBB-BM44 datasets are from the Mount Sinai NIH Brain Bank (MSBB) cohort, containing 112 samples from the frontal pole (BM10, Brodmann Area 10), 106 samples from the superior temporal gyrus (BM22, Brodmann Area 22), 96 samples from the parahippocampal gyrus (BM36, Brodmann Area 36), and 88 samples from the inferior frontal gyrus (BM44, Brodmann Area 44), respectively. The datasets GSE95587-FG, GSE125583-FG, GSE159699-LTL, and GSE231341-iPSC are from the Gene Expression Omnibus database. They contain 117 samples from the fusiform gyrus (FG), 288 samples from the FG, 30 samples from the lateral temporal lobe (LTL), and 24 samples from induced pluripotent stem cell (iPSC) neurons, respectively. The ROSMAP-B1-DLPFC and ROSMAP-B3-DLPFC datasets are from the Religious Orders Study (ROS) and the Memory and Aging Project (MAP), respectively, containing 359 samples from the dorsolateral prefrontal cortex (DLPFC) and 183 samples from the DLPFC.

[0034] Download the annotation file (Gene Transfer Format (GTF)) of release 35 of the Human Genome Build 38 (GRCh38) from the GENCODE Database. This file details the mapping information between genes and their corresponding transcripts.

[0035] It should be noted that after obtaining the expression levels of transcripts, low expression levels can be screened. To remove transcripts with low expression levels and limited information content in most samples, the transcript expression ratio (ER) is introduced. The specific definition is as follows: First, a transcript expression threshold τ=1 is set. For transcripts... Count the number of samples whose expression level exceeds the threshold:

[0036] in, This is an indicator function that takes the value 1 when the condition is true and 0 otherwise. Subsequently, the transcript is defined. Expression ratio index:

[0037] when (in When a transcript is considered to have sufficient expression stability in the dataset, it is retained. This yields the set of transcripts in the dataset filtered for low expression:

[0038] Considering the differences in sample sources and sequencing conditions across different datasets, a unified transcript filtering standard was adopted to ensure the comparability of subsequent analysis results across multiple datasets. Specifically, the union of the transcript sets filtered for low expression was taken from all datasets.

[0039] in, This represents the total number of data sets. This represents the final unified set of transcripts.

[0040] Step 12: Perform differential analysis based on the expression levels of all genes to obtain the differential expression characteristics of each gene; perform differential analysis on the expression levels of all transcripts to obtain the differential expression characteristics of each transcript; perform differential analysis on the expression ratios of all transcripts to obtain the differential expression ratios of each transcript; and perform differential analysis on alternative splicing events of all sample gene signals to obtain the alternative splicing differential characteristics of each gene.

[0041] The above-mentioned gene expression level difference features are used to describe the difference in gene expression levels in different sample gene signals; the above-mentioned transcript expression level difference features are used to describe the difference in transcript expression levels in different sample gene signals; the above-mentioned expression ratio difference features are used to describe the difference in the proportion of transcripts in different sample gene signals; and the above-mentioned alternative splicing difference features are used to describe the difference in alternative splicing events corresponding to genes in different sample gene signals.

[0042] In some embodiments of this application, the steps of performing differential analysis based on the expression levels of all genes to obtain the expression level difference characteristics of each gene, performing differential analysis on the expression levels of all transcripts to obtain the expression level difference characteristics of each transcript, performing differential analysis on the expression ratios of all transcripts to obtain the expression ratio difference characteristics of each transcript, and performing differential analysis on alternative splicing events of all sample gene signals to obtain the alternative splicing difference characteristics of each gene include: The first step is to integrate the expression levels of all transcripts into a matrix to obtain the transcript expression matrix.

[0043] The second step is to calculate the expression level of each gene in each sample gene signal based on the expression levels of all transcripts, and then integrate the expression levels of all genes into a matrix to obtain the gene expression matrix.

[0044] Specifically, through the formula:

[0045] Computational genes In sample gene signals Expression level in .

[0046] in, Transcript In sample gene signals The amount of expression in Indicates gene The corresponding set of transcripts.

[0047] The third step is to calculate the transcript proportion expression matrix based on the transcript expression matrix.

[0048] The elements in the transcript proportion expression matrix above represent the proportion of transcript expression in the sample gene signal.

[0049] Specifically, through the formula:

[0050] Calculate transcripts In sample gene signals The proportion of values ​​expressed in .

[0051] in, Indicates low expression of filtered transcripts In sample gene signals The amount of expression in Indicates gene The sum of the expression levels of the corresponding transcripts, .

[0052] All proportional expression values ​​are integrated into a matrix to obtain the transcript proportional expression matrix.

[0053] The fourth step involves performing differential analysis based on the transcript expression matrix and gene expression matrix to obtain the differential expression characteristics of transcripts contained in each gene and the differential expression characteristics of the gene itself.

[0054] Specifically, based on the gene expression matrix of each gene, the expression level difference characteristics of each gene are calculated.

[0055] The differential expression value of each transcript was calculated using the transcript expression matrix and gene expression matrix.

[0056] For each gene, the differential expression values ​​of all transcripts corresponding to the gene are sorted from largest to smallest absolute value, and the differential expression values ​​of the top few transcripts in the sorting results are selected to constitute the gene expression level differential feature.

[0057] For example, the differential expression value is defined as the logarithm of the fold change in gene expression between the experimental group (i.e., the gene signaling set of disease-related samples) and the control group (i.e., the gene signaling set of disease-unrelated samples):

[0058] in, These are differential expression values. To express the multiple of change.

[0059] This step takes the transcript expression matrix and gene expression matrix as input data and inputs them into a negative binomial distribution (NB) model for calculation to obtain differential expression values.

[0060] Specifically, the edgeR software package (Bioconductor) was used for implementation. Based on the Negative Binomial (NB) distribution model, it effectively characterizes the mean-variance relationship in RNA-seq counting data. For each dataset, the constructed isoform expression matrix and gene expression matrix were used as inputs. To correct for systematic bias introduced by differences in sequencing depth and RNA composition between different samples, the TMM (Trimmed Mean of M-values) method in edgeR was used to normalize the counting data. In differential expression analysis, two statistical modeling strategies were employed depending on whether each dataset had sample covariate information: generalized linear model (GLM) analysis with covariates and exact test without covariates. When the dataset provided covariate information (such as age, sex, RIN value, etc.), a design matrix was constructed:

[0061]

[0062] in, Representation of features The counting vector, Group This represents the grouping variable between the experimental group and the control group. Indicates the first One covariate. When covariate information is not provided, use... edgeR The exact test method based on the negative binomial distribution was used to examine the expression differences between the experimental and control groups. For each analytical feature, edgeR Output the corresponding logarithmic multiple change (log2FC), raw p-value, and other statistics. The Benjamini–Hochberg (BH) method is used to perform multiple test corrections on the p-values ​​of all features to control the false discovery rate (FDR).

[0063] It should be noted that the differential expression values ​​of the first few transcripts in the sorting results are selected, and the number selected is a preset number. If the number of differential expression values ​​in the sorting results is less than the preset number, the insufficient part is padded with 0.

[0064] The fifth step involves differential analysis based on the transcript expression ratio matrix to obtain the differential expression ratio characteristics of the transcripts contained in each gene.

[0065] Specifically, based on the transcript proportion expression matrix, the mean proportion expression value of each transcript in disease-related sample gene signals is calculated, and the difference between the mean proportion expression value of each transcript in disease-unrelated sample gene signals is calculated, and the difference proportion value of the transcripts is then recorded.

[0066] For each gene, the differential proportion values ​​of all transcripts corresponding to the gene are sorted from largest to smallest in absolute value, and the differential proportion values ​​of the top few transcripts in the sorting results are selected to constitute the expression proportion difference feature of the transcripts contained in the gene.

[0067] It should be noted that the differential proportion values ​​of the first few transcripts in the sorting results are selected, and the selected number is a preset number. If the number of differential expression values ​​in the sorting results is less than the preset number, the insufficient part is padded with 0.

[0068] For example, nonparametric statistical methods are used to test differences between groups. For each dataset, the constructed transcript proportion expression matrix is ​​used as input. To avoid interference from uninformative features in the statistical test results, the transcript proportion expression matrix undergoes a multi-step filtering process: First, features with identical transcript proportion expression values ​​across all samples are removed, i.e., transcript proportion features that do not change between samples. Second, for genes containing only a single transcript, their transcript proportion is theoretically always 1, and has no biological significance regarding alternative splicing or proportion changes. Furthermore, when such features exhibit a binary distribution (0 or 1) across all samples, they are removed from the analysis. After the above filtering steps, transcript proportion features that are biologically interpretable and exhibit variation between samples are retained for subsequent statistical analysis.

[0069] For each retained transcript proportion feature, the Mann–Whitney U test (also known as the Wilcoxon rank-sum test) was used to compare the distribution differences between the experimental and control groups. In addition to the significance test, to quantitatively describe the direction and magnitude of transcript proportion changes between groups, ΔPSI (deltaPSI) was further calculated as an effect size indicator, defined as the difference between the mean isoform proportions of the experimental and control groups:

[0070] in, and They represent transcripts. Mean transcript proportion expression values ​​in the experimental and control groups. The raw p-values ​​for all features were corrected using the Benjamini–Hochberg (BH) method with multiple tests to control for the false discovery rate (FDR).

[0071] The sixth step is to perform alternative splicing event analysis on all sample gene signals to obtain the degree of splicing use of each alternative splicing event in the sample gene signals.

[0072] For example, the Replicate Multivariate Analysis of Transcript Splicing (rMATS) software can be used to perform alternative splicing event analysis on all sample gene signals to obtain the degree of splicing use of each alternative splicing event in the sample gene signals.

[0073] rMATS can identify and quantify the following five classic alternative splicing event types: Skipped Exon (SE), Mutually Exclusive Exons (MXE), Alternative 5′ Splice Site (A5SS), Alternative 3′ Splice Site (A3SS), and Retained Intron (RI). The core output metric of rMATS is the Percent Spliced ​​In (PSI), which measures the frequency of a specific alternative splicing event in a given sample. Let N be the set of all variable splicing events, and N be the total number of samples. For any variable splicing event... In the sample middle, This indicates the number of sequencing reads that support the inclusion of this alternative splicing event. The number of reads supporting the exclusion of this variable splicing event is represented by the PSI value of that splicing event in sample n, which is then defined as:

[0074] in: This indicator directly reflects the event. In the sample The degree of splicing used in [the context]. Taking skipping exon (SE) events as an example: A value of approximately 1 indicates that the exon is included in most transcripts. A value of approximately 0 indicates that the exon is skipped in most transcripts.

[0075] Step 7: For each variable splice event, calculate the splice usage difference value of the variable splice event based on all splice usage values ​​of the variable splice event.

[0076] For example, splicing uses difference values. The expression is: ,in, This represents the mean value of the degree of splicing use of variable splicing events in the experimental group. This represents the mean value of the degree of splicing use of the variable splicing event in the control group.

[0077] rMATS assumes that inclusion reads and exclusion reads follow a binomial or beta-binomial distribution, and estimates the PSI distribution for both the case group and the control group. For each splicing event, the p-value and the difference in splice inclusion rate (ΔPSI) are calculated using the likelihood ratio test (LRT). .

[0078] Because it examines tens of thousands of alternative splicing events simultaneously, rMATS applies the Benjamini-Hochberg (BH) correction method to the P-value to control the false discovery rate (FDR). When performing differential analysis of alternative splicing events, the binary alignment map (BAM) file and corresponding sample grouping information for each dataset are used as input. rMATS systematically identifies and classifies alternative splicing events by traversing all potential splicing junctions on the genome and calculates the corresponding PSI value. , ΔΨ , p Values ​​and FDR. When defining the Alternative Splicing Event Differential Feature, gene values ​​were extracted based on the analysis results from rMATS software. The ΔPSI value of all variable splice events (including five types: SE, MXE, A5SS, A3SS, and RI) in analysis group m.

[0079] Step 8: For each gene, sort the splicing use difference values ​​of all alternative splicing events contained in the gene in descending order of absolute value, and select the splicing use difference values ​​of the top few alternative splicing events in the sorting results to form the alternative splicing difference features of the gene.

[0080] For example, the splicing difference values ​​of the first few variable splicing events in the sorting results are selected, where the selected number is a preset number. If the number of difference values ​​in the sorting results is less than the preset number, the insufficient part is padded with 0.

[0081] Step 13: The expression level difference features of each gene, the expression level difference features of the corresponding transcript, the expression level ratio difference features of the corresponding transcript, and the alternative splicing difference features are spliced ​​together to obtain the final features of each gene.

[0082] Specifically, through the formula:

[0083] Genes were spliced ​​together The final feature .

[0084] in, Indicates gene The expression difference characteristic (this characteristic is the logarithmic value of the fold change in gene expression between the experimental group and the control group) ), Indicates gene The differential expression characteristics of transcripts. Indicates gene The characteristics of differences in the expression ratio of transcripts, Indicates gene The variable shear differential characteristics.

[0085] Step 14: Construct a functional association network based on the final characteristics of all genes.

[0086] In the above functional association network, multiple nodes correspond one-to-one with multiple genes, and the edges between nodes represent the interaction relationships between the corresponding two genes.

[0087] In some embodiments of this application, the steps of constructing a functional association network based on the final characteristics of all genes include: The first step is to encode the final characteristics of each gene to obtain the gene representation of each gene.

[0088] For example, a multi-layer perceptron (MLP) can be used to encode the final features of each gene, obtaining the gene representation for each gene. The encoding process can be represented as follows:

[0089] in Indicates the first Each level corresponds to a nonlinear mapping function. The encoding network consists of two fully connected layers, with layer normalization and nonlinear activation functions introduced between layers to ensure stable representations of features from different levels within a unified space. After encoding, features from each level are converted into vectors of the same dimension. To adaptively learn the contribution of information from different levels to the final prediction result, the four level encoding vectors are concatenated and input into the attention weight calculation network to obtain the corresponding weight coefficients for each level of information. The weights are normalized using the Softmax function, satisfying:

[0090] in, Indicates the first These characteristics at various levels are in genes Weighting on.

[0091] Based on the learned layer information weights, the encoded representations of each layer are weighted and summed to obtain the fused gene representation:

[0092] Finally, the output representation of this module is obtained through a linear mapping layer, which is used as the node input features of subsequent graph neural networks.

[0093] The second step is to generate a corresponding node for each gene, and use the corresponding gene representation as the node attribute. When there is an interaction relationship between genes, an edge is generated between the corresponding two nodes to obtain the functional association network.

[0094] If there is no interaction between genes, no edge is generated between the corresponding two nodes.

[0095] For example, publicly available databases can be used to analyze gene interactions. Two genes that interact are considered to have an interaction relationship; otherwise, no interaction relationship exists. For instance, the STRING (Search Tool for the Retrieval of Interacting Genes / Proteins) database can be used as a source of protein interaction information. First, gene names are mapped to their corresponding STRING protein IDs. Genes that cannot be successfully matched with proteins in the STRING database are recorded as unmapped genes and removed in subsequent network construction. Then, the interaction relationships between target genes and their corresponding proteins are extracted, retaining only records where both interacting proteins belong to the target gene set. Each protein interaction in the STRING database is accompanied by a confidence score. This study uses this confidence score as the weight of edges in the gene network, constructing a weighted functional association network in this way.

[0096] Step 15: Perform gene disease prediction on the functional association network to obtain the probability that each gene is associated with Alzheimer's disease.

[0097] Specifically, graph attention aggregation is performed on the functional association network to obtain the final node representation of each node. Then, based on the final node representation of each node, the probability of the corresponding gene being associated with Alzheimer's disease is calculated.

[0098] For example, the expression for graph attention aggregation is:

[0099] in, The final node representation of the node. This represents the gene representation corresponding to the node. Indicates gene The set of adjacent nodes of the corresponding node. Indicates the attention coefficient. Indicates gene Gene representation, This represents the updated node representation.

[0100] The Sigmoid function and similar methods can be used to calculate the probability that a corresponding gene is associated with Alzheimer's disease. The expression is:

[0101] in, Indicates gene Probability associated with Alzheimer's disease Indicates weight, This is a bias term.

[0102] It should be noted that after obtaining the probability that a gene is associated with Alzheimer's disease, the genes associated with Alzheimer's disease can be identified based on probability analysis. For example, a probability threshold can be set, and genes with a probability greater than the probability threshold can be considered as genes associated with Alzheimer's disease.

[0103] In some embodiments of this application, before performing this step, the parameters in the graph attention aggregation formula and the sigmoid function can be trained under supervised supervision. For example, the above process can be performed using multiple sample gene signals as training data to obtain the probability that the corresponding gene is associated with Alzheimer's disease, and the binary cross-entropy loss function can be calculated. The parameters in the graph attention aggregation formula and the sigmoid function can then be iteratively optimized using a backpropagation algorithm until the loss function converges or reaches a preset number of training rounds. The binary cross-entropy loss function is as follows:

[0104] in, This represents the value of the loss function. This represents the number of genes in the sample gene signal used as training data. This represents the actual label corresponding to the gene. , This indicates that genes are related to diseases. This indicates that the gene is not related to the disease.

[0105] For example, in order to verify the performance of the Alzheimer's disease gene prediction method proposed in this application, the method was compared with other benchmark methods for predicting pathogenic genes, and the experimental results are shown in Table 1.

[0106] Table 1

[0107] In summary, the disease gene prediction method used in this paper has achieved certain advantages, with both metrics outperforming other methods. For the Area Under the Receiver Operating Characteristic curve (AUROC) metric, our method achieves 0.912, which is an improvement of 0.08, 0.13, 0.03, and 0.02 compared to the following methods: Random Walk with Restart-Meta Heuristic (RWR-MH) (0.827), Gene Semantic Inference (GSI) (0.784), Disease Hypergraph (DISHyper) (0.889), and Latent Interaction Multi-Objective Graph Convolutional Network (LIMO-GCN) (0.898). For the Area Under the Precision-Recall Curve (AUPRC) metric, our method achieves a score of 0.713, which is an improvement of 0.36, 0.43, 0.14, and 0.04 compared to RWR-MH's 0.351, GSI's 0.281, DISHyper's 0.568, and LIMO-GCN's 0.683.

[0108] It is worth mentioning that mining the hierarchical relationships between sample gene signals, transcripts, and genes, and analyzing alternative splicing events, can fully study the characteristics of gene information at different levels, improve the richness and comprehensiveness of Alzheimer's disease gene prediction, and make disease gene prediction based on information-rich feature data, thereby increasing the accuracy of the predicted probability and improving the accuracy and reliability of Alzheimer's disease gene prediction.

[0109] Furthermore, this application adaptively integrates multi-level information from gene differential expression, transcript differential expression, transcript ratio differential expression, and differential analysis of alternative splicing events by designing a multi-layered alternative splicing omics feature fusion mechanism and a residual graph attention network model. This integration is combined with gene network topology for joint modeling, enabling gene prediction for Alzheimer's disease. Compared to existing single-feature or traditional graph model methods, this application improves feature utilization efficiency and model stability, alleviates the oversmoothing problem of deep networks, and enhances prediction accuracy and generalization ability, demonstrating high application value and promising prospects for wider adoption.

[0110] The Alzheimer's disease gene prediction device based on alternative splicing provided in this application is described below as an example.

[0111] like Figure 2 As shown, this application provides an Alzheimer's disease gene prediction device based on alternative splicing. The Alzheimer's disease gene prediction device 200 based on alternative splicing includes: The acquisition module 201 is used to acquire the expression levels of multiple transcripts in the gene signals of each sample of Alzheimer's disease, and to acquire the expression level of the gene corresponding to each transcript; The differential analysis module 202 is used to perform differential analysis based on the expression levels of all genes to obtain the differential expression characteristics of each gene, to perform differential analysis on the expression levels of all transcripts to obtain the differential expression characteristics of each transcript, to perform differential analysis on the expression ratios of all transcripts to obtain the differential expression ratios of each transcript, and to perform differential analysis on alternative splicing events of all sample gene signals to obtain the alternative splicing differential characteristics of each gene. The splicing module 203 is used to splice the expression level difference features of each gene, the expression level difference features of the corresponding transcripts, the expression level ratio difference features of the corresponding transcripts, and the alternative splicing difference features to obtain the final features of each gene. Module 204 is used to construct a functional association network based on the final features of all genes; in the functional association network, multiple nodes correspond one-to-one with multiple genes, and the edges between nodes represent the interaction relationships between the corresponding two genes. The prediction module 205 is used to predict gene diseases in the functional association network and obtain the probability that each gene is associated with Alzheimer's disease.

[0112] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0113] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0114] like Figure 3 As shown, an embodiment of this application provides a terminal device, wherein the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 3 The diagram shows only one processor, a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 executes the computer program D102 to implement the steps in any of the above method embodiments.

[0115] Specifically, when the processor D100 executes the computer program D102, it acquires the expression levels of multiple transcripts in each sample gene signal of Alzheimer's disease and obtains the expression level corresponding to each transcript. Then, it performs differential analysis based on all expression levels to obtain the expression level difference characteristics of each gene. It performs differential analysis on the expression levels of all transcripts to obtain the expression level difference characteristics of each transcript. It performs differential analysis on the expression ratio of all transcripts to obtain the expression ratio difference characteristics of each transcript. It performs differential analysis on alternative splicing events of all sample gene signals to obtain the alternative splicing difference characteristics of each gene. Then, it splices together the expression level difference characteristics of each gene, the expression level difference characteristics of the transcripts contained in the gene, the transcript expression ratio difference characteristics, and the alternative splicing difference characteristics to obtain the final characteristics of each gene. Then, it constructs a functional association network based on the final characteristics of all genes. Finally, it performs gene disease prediction on the functional association network to obtain the probability that each gene is associated with Alzheimer's disease. Among these methods, mining sample gene signals, transcripts, and hierarchical relationships between genes, and analyzing alternative splicing events, can fully study the characteristics of gene information at different levels, improve the richness and comprehensiveness of information in Alzheimer's disease gene prediction, and make disease gene prediction based on information-rich feature data, thereby increasing the accuracy of the predicted probability and improving the accuracy and reliability of Alzheimer's disease gene prediction.

[0116] The processor D100 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0117] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may be an external storage device of the terminal device D10, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device D10. Furthermore, the memory D101 may include both internal and external storage units of the terminal device D10. The memory D101 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory D101 can also be used to temporarily store data that has been output or will be output.

[0118] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0119] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.

[0120] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above-described embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to an Alzheimer's disease gene prediction method device / terminal device based on alternative splicing, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0121] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0122] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0123] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention.

Claims

1. A method for predicting Alzheimer's disease genes based on alternative splicing, characterized in that, include: The expression levels of multiple transcripts in the gene signals of each sample of Alzheimer's disease were obtained, and the expression level of the gene corresponding to each transcript was obtained. Differential analysis was performed on the expression levels of all genes to obtain the expression difference characteristics of each gene. Differential analysis was also performed on the expression levels of all transcripts to obtain the expression difference characteristics of each transcript. Differential analysis was also performed on the expression ratios of all transcripts to obtain the expression ratio difference characteristics of each transcript. Finally, differential analysis was performed on the alternative splicing events of all sample gene signals to obtain the alternative splicing difference characteristics of each gene. The expression level difference features of each gene, the expression level difference features of the corresponding transcript, the expression level ratio difference features of the corresponding transcript, and the alternative splicing difference features are spliced ​​together to obtain the final features of each gene. A functional association network is constructed based on the final characteristics of all genes; in the functional association network, multiple nodes correspond one-to-one with multiple genes, and the edges between nodes represent the interaction relationships between the corresponding two genes. Gene disease prediction is performed on the functional association network to obtain the probability that each gene is associated with Alzheimer's disease.

2. The Alzheimer's disease gene prediction method according to claim 1, characterized in that, The differential expression analysis based on the expression levels of all genes yields the expression level difference characteristics of each gene; the differential expression analysis based on the expression levels of all transcripts yields the expression level difference characteristics of each transcript; and the differential expression analysis based on the expression ratio of all transcripts yields the expression ratio difference characteristics of each transcript. This includes: The expression levels of all transcripts are integrated into a matrix to obtain the transcript expression matrix; The expression level of each gene in each sample gene signal is calculated based on the expression levels of all transcripts, and the expression levels of all genes are integrated into a matrix to obtain the gene expression matrix. A transcript proportion expression matrix is ​​calculated based on the transcript expression matrix; the elements in the transcript proportion expression matrix are the proportion of transcript expression values ​​in the sample gene signals. Based on the transcript expression matrix and gene expression matrix, differential analysis was performed to obtain the differential expression characteristics of transcripts contained in each gene and the differential expression characteristics of the gene itself. Differential analysis was performed based on the transcript proportion expression matrix to obtain the expression proportion difference characteristics of the transcripts contained in each gene.

3. The Alzheimer's disease gene prediction method according to claim 2, characterized in that, The calculation of the expression level of each gene in each sample gene signal based on the expression levels of all transcripts includes: Through the formula: Computational genes In sample gene signals Expression level in ; in, Transcript In sample gene signals The amount of expression in Indicates gene The corresponding set of transcripts; The step of calculating the transcript proportion expression matrix based on the transcript expression matrix includes: Through the formula: Calculate transcripts In sample gene signals The proportion of values ​​expressed in ; in, Indicates low expression of filtered transcripts In sample gene signals The amount of expression in Indicates gene The sum of the expression levels of the corresponding transcripts, ; All proportional expression values ​​are integrated into a matrix to obtain the transcript proportional expression matrix.

4. The Alzheimer's disease gene prediction method according to claim 3, characterized in that, The categories of the gene signals in the Alzheimer's disease samples are disease-related or disease-unrelated; The differential analysis based on the transcript expression matrix and gene expression matrix yields the differential expression characteristics of transcripts contained in each gene and the differential expression characteristics of the gene itself, including: Based on the gene expression matrix of each gene, calculate the expression level difference characteristics of each gene; Using the transcript expression matrix and gene expression matrix, the differential expression value of each transcript was calculated; For each gene, the differential expression values ​​of all transcripts corresponding to the gene are sorted from largest to smallest absolute value, and the differential expression values ​​of the top few transcripts in the sorting results are selected to constitute the differential expression features of the gene. The differential analysis based on the transcript proportion expression matrix yields the differential expression proportion characteristics of the transcripts contained in each gene, including: Based on the transcript proportion expression matrix, the difference between the mean proportion expression value of each transcript in disease-related sample gene signals and the mean proportion expression value in disease-unrelated sample gene signals is calculated, and the difference is expressed as the difference proportion value of the transcript. For each gene, the differential proportion values ​​of all transcripts corresponding to the gene are sorted from largest to smallest in absolute value, and the differential proportion values ​​of the top few transcripts in the sorting results are selected to constitute the expression proportion differential feature of the transcripts contained in the gene.

5. The Alzheimer's disease gene prediction method according to claim 1, characterized in that, The differential analysis of alternative splicing events on all sample gene signals yields the differential features of alternative splicing for each gene, including: Alternative splicing event analysis was performed on all sample gene signals to obtain the degree of splicing use of each alternative splicing event in the sample gene signals. For each variable splicing event, calculate the splicing usage difference value of the variable splicing event based on all splicing usage values ​​of the variable splicing event; For each gene, the splicing use difference values ​​of all alternative splicing events contained in the gene are sorted from largest to smallest absolute value, and the splicing use difference values ​​of the top multiple alternative splicing events in the sorting results are selected to constitute the alternative splicing difference features of the gene.

6. The Alzheimer's disease gene prediction method according to claim 1, characterized in that, The expression level difference features of each gene, the expression level difference features of the corresponding transcript, the expression ratio difference features of the corresponding transcript, and the alternative splicing difference features are spliced ​​together to obtain the final features of each gene, including: Through the formula: Genes were spliced ​​together The final feature ; in, Indicates gene The expression differences Indicates gene The differential expression characteristics of transcripts. Indicates gene The characteristics of differences in the expression ratio of transcripts, Indicates gene The variable shear differential characteristics.

7. The Alzheimer's disease gene prediction method according to claim 6, characterized in that, The construction of a functional association network based on the final characteristics of all genes includes: The final characteristics of each gene are encoded to obtain the gene representation of each gene; A corresponding node is generated for each gene, and the corresponding gene is used as the node attribute. When there is an interaction relationship between genes, an edge is generated between the two corresponding nodes to obtain the functional association network.

8. The Alzheimer's disease gene prediction method according to claim 7, characterized in that, The process of predicting gene diseases using the functional association network to obtain the probability that each gene is associated with Alzheimer's disease includes: Graph attention aggregation is performed on the functional association network to obtain the final node representation of each node; Based on the final node representation of each node, the probability that the corresponding gene is associated with Alzheimer's disease is calculated.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the Alzheimer's disease gene prediction method based on alternative splicing as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the Alzheimer's disease gene prediction method based on alternative splicing as described in any one of claims 1 to 8.