Pathogenic gene prediction method based on whole transcriptome correlation research and related equipment

By constructing a splicing percentage and genotype matrix model based on RNA-seq samples, and combining genotype and splicing characteristics, the problem of low accuracy in predicting pathogenic genes in existing technologies has been solved, and more accurate pathogenicity prediction has been achieved.

CN120108497BActive Publication Date: 2025-12-05CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510285374.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-12-05
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

Existing whole transcriptome association studies do not adequately account for alternative splicing events, resulting in low accuracy in predicting pathogenic genes.

Method used

By obtaining the splicing percentage matrix and genotype matrix of RNA-seq samples, first and second prediction models were constructed. Pathogenicity analysis was performed by combining genotype and splicing percentage characteristics to reveal the regulatory mechanism between genotype and splicing events.

Benefits of technology

It improves the accuracy and information richness of pathogenic gene prediction, captures the potential association between splicing events and pathogenic traits, and provides a new perspective for pathogenicity prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108497B_ABST
    Figure CN120108497B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of pathogenic gene prediction, and provides a pathogenic gene prediction method based on whole-transcriptome correlation research and related equipment. The method provided by the application comprises the following steps: calculating a splicing percentage matrix of each gene according to an RNA-seq sample; obtaining a genotype matrix of each gene and a genotype matrix of each splicing factor; constructing a first prediction model based on the genotype matrix of all splicing factors and all splicing percentage matrices, and obtaining a predicted splicing percentage matrix corresponding to all splicing factors by using the first prediction model; constructing a second prediction model based on the predicted splicing percentage matrix and the genotype matrix of all genes, and obtaining a final predicted splicing percentage matrix of all genes by using the second prediction model; and performing pathogenicity analysis according to the final predicted splicing percentage matrix of all genes to obtain a gene pathogenicity prediction result. The method provided by the application can improve the accuracy of pathogenic gene prediction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of pathogenic gene prediction, and in particular relates to a pathogenic gene prediction method based on transcriptome-wide association study and related equipment. BACKGROUND

[0002] Transcriptome-wide association study (TWAS) aims to explore the potential relationship between genes and traits through genotype and gene expression association analysis, and provides important clues for the genetic mechanism research of complex traits. Existing TWAS methods are usually based on multiple biological information such as gene expression level, regulatory element, genetic variation, etc., combined with machine learning or statistical model to predict trait-related genes. For example, by predicting gene expression level using genotype, and analyzing its association with traits; combining exon expression and genotype data, and evaluating trait correlation through a Bayesian framework. However, these methods mostly focus on steady-state gene expression level, and do not fully consider the dynamic change characteristics at the transcriptome level, especially the alternative splicing events. Alternative splicing is a key step of post-transcriptional regulation, which generates multiple messenger RNA (mRNA) transcripts from the same pre-messenger RNA (pre-mRNA) through different splicing ways (such as exon skipping, alternative splicing site selection, etc.), greatly increasing the complexity of the transcriptome. Percent spliced in (PSI) is a commonly used indicator to quantify splicing events, which can accurately describe the proportion of splicing events in transcripts. Each gene usually contains multiple alternative splicing events, and these events may contain trait-related regulatory signals. However, in current TWAS research, little consideration is given to PSI, resulting in inaccurate representation of splicing event characteristics, which cannot describe the relationship between splicing event and pathogenicity of gene, and further leads to low accuracy of pathogenic gene prediction. SUMMARY

[0003] The present application provides a pathogenic gene prediction method based on transcriptome-wide association study and related equipment, which can solve the problem of low accuracy of pathogenic gene prediction.

[0004] In a first aspect, the embodiments of the present application provide a pathogenic gene prediction method based on transcriptome-wide association study, which comprises:

[0005] An RNA-seq sample containing multiple genes is obtained, and a splicing percentage matrix of each gene is calculated according to the RNA-seq sample; there are multiple splicing factors in all genes, and the splicing percentage matrix is used to describe the proportion of splicing events in the gene;

[0006] Obtain the genotype matrix for each gene and the genotype matrix for each splicing factor. The genotype matrix of a gene is used to describe the single nucleotide polymorphisms within a certain range upstream and downstream of the gene, and the genotype matrix of a splicing factor is used to describe the single nucleotide polymorphisms within a certain range upstream and downstream of the location of the splicing factor.

[0007] The first prediction model is constructed based on the genotype matrix of all splicing factors and the splicing percentage matrix of all splicing factors, and the predicted splicing percentage matrix corresponding to all splicing factors is obtained using the first prediction model.

[0008] A second prediction model is constructed based on the predicted splice percentage matrix and the genotype matrix of all genes, and the final predicted splice percentage matrix of all genes is obtained using the second prediction model.

[0009] Pathogenicity analysis was performed based on the final predicted splicing percentage matrix of all genes to obtain gene pathogenicity prediction results. The gene pathogenicity prediction results are used to describe the association between splicing factors, corresponding single nucleotide polymorphism sites and pathogenicity in each gene.

[0010] Optionally, the first prediction model is:

[0011] Y = X SF ω+ε

[0012] Where Y represents the matrix of all splice percentages, X SF This represents the genotype matrix of all splicing factors, where ω represents the effect size and ε represents the error.

[0013] Optionally, a second prediction model is constructed based on the predicted splicing percentage matrix and the genotype matrix of all genes, including:

[0014] Multiple initial secondary prediction models were constructed based on the predicted splicing percentage matrix and the genotype matrix of all genes;

[0015] For each initial second prediction model, predictions are made using the initial second prediction model to obtain prediction results including multiple splice percentage prediction values, and the fitted value of the initial second prediction model is calculated based on the prediction results;

[0016] The initial second prediction model with a fitted value greater than or equal to the fitted threshold is used as the second prediction model.

[0017] Optionally, the initial second prediction model is:

[0018] Y = (X, Y) SF )β+ε

[0019] Where Y represents the splicing percentage matrix, and X represents the genotype matrix of all genes, Y SFrepresents a predicted splicing percentage matrix, β represents a weight, and ε represents an error.

[0020] Optionally, the fitting value of the initial second prediction model is calculated according to the prediction result, and the fitting value comprises:

[0021] The fitting value R is calculated by the formula:

[0022]

[0023] The fitting value R is calculated by the formula: 2 ;

[0024] wherein y i represents the ith real splicing percentage, represents the ith splicing percentage prediction value, represents the mean of y i in all prediction results, and n represents the number of splicing percentage prediction values.

[0025] Optionally, pathogenicity analysis is performed according to the final predicted splicing percentage matrix of all genes to obtain a gene pathogenicity prediction result, and the pathogenicity analysis comprises:

[0026] Obtaining pathogenic label information of each gene; the pathogenic label information is used to describe whether the gene has pathogenicity;

[0027] Performing pathogenicity analysis based on the pathogenic label information of all genes and the final predicted splicing percentage matrix to obtain a gene pathogenicity prediction result.

[0028] In a second aspect, an embodiment of the present application provides a pathogenic gene prediction device based on whole transcriptome correlation research, comprising:

[0029] A calculation module is configured to obtain an RNA-seq sample containing a plurality of genes, and calculate a splicing percentage matrix of each gene according to the RNA-seq sample; a plurality of splicing factors exist in all genes, and the splicing percentage matrix is used to describe the proportion of splicing events in the gene;

[0030] An obtaining module is configured to obtain a genotype matrix of each gene and obtain a genotype matrix of each splicing factor; the genotype matrix of the gene is used to describe single nucleotide polymorphisms within a certain range upstream and downstream of the gene, and the genotype matrix of the splicing factor is used to describe single nucleotide polymorphisms within a certain range upstream and downstream of the position of the splicing factor;

[0031] A first construction module is configured to construct a first prediction model based on the genotype matrix of all splicing factors and the splicing percentage matrix of all genes, and obtain a predicted splicing percentage matrix corresponding to all splicing factors by using the first prediction model;

[0032] a second constructing module, configured to construct a second prediction model based on the predicted splicing percentage matrix and the genotype matrix of all genes, and obtain a final predicted splicing percentage matrix of all genes by using the second prediction model;

[0033] a pathogenicity analysis module, configured to perform pathogenicity analysis according to the final predicted splicing percentage matrix of all genes, and obtain a gene pathogenicity prediction result; the gene pathogenicity prediction result is used to describe the association relationship between a splicing factor, a corresponding single nucleotide polymorphism site and pathogenicity in each gene.

[0034] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor implements the pathogenic gene prediction method based on whole transcriptome association study when executing the computer program.

[0035] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the pathogenic gene prediction method based on whole transcriptome association study.

[0036] The above-mentioned scheme of the present application has the following beneficial effects:

[0037] In the embodiments of the present application, the RNA-seq sample containing multiple genes is obtained, the splicing percentage matrix of each gene is calculated according to the RNA-seq sample, then the genotype matrix of each gene is obtained, the genotype matrix of each splicing factor is obtained, the first prediction model is constructed based on the genotype matrix of all splicing factors and all splicing percentage matrices, the predicted splicing percentage matrix of each gene is obtained by using the first prediction model, then the second prediction model is constructed based on all predicted splicing percentage matrices and the genotype matrix of all genes, the final predicted splicing percentage matrix of all genes is obtained by using the second prediction model, and finally the pathogenicity analysis is performed according to the final predicted splicing percentage matrix of all genes to obtain the gene pathogenicity prediction result. The prediction model is constructed based on the genotype matrix and the splicing percentage matrix, which can combine the information of the genotype and the splicing percentage of the gene, improve the accuracy and information richness of the result obtained by the prediction model, reveal the regulation mechanism between the genotype and the splicing event, and perform the pathogenicity prediction according to the accurate final predicted splicing percentage matrix, so as to capture the potential association of the splicing event with the pathogenicity, provide a new perspective for the pathogenicity prediction research, and effectively improve the accuracy of the gene pathogenicity prediction.

[0038] Other beneficial effects of the present application will be described in detail in the following specific embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only some of the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0040] Figure 1 The flow chart of the pathogenic gene prediction method based on whole transcriptome correlation study provided by an embodiment of the present application;

[0041] Figure 2 The structural schematic diagram of the pathogenic gene prediction device based on whole transcriptome correlation study provided by an embodiment of the present application;

[0042] Figure 3 The structural schematic diagram of the terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0043] In the following description, for the purpose of explanation and not limitation, specific details are set forth, such as particular system configurations, techniques, etc., in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present application with unnecessary detail.

[0044] It should be understood that the term "comprising" as used in the specification and the appended claims indicates the presence of the recited features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0045] It should also be understood that the term "and / or" as used herein refers to any combination of one or more of the associated listed items, and all possible combinations, and includes these combinations.

[0046] As used in the specification and the appended claims, the term "if' can be interpreted as meaning "when" or "once" or "in response to a determination" or "in response to a detection" depending on the context. Similarly, the phrase "if determined" or "if detected [the described condition or event]" can be interpreted to mean "once determined" or "in response to a determination" or "once detected [the described condition or event]" or "in response to a detection [the described condition or event]" depending on the context.

[0047] In addition, in the description of the present application and the appended claims, the terms "first", "second", "third", etc. are used only to distinguish descriptions, and cannot be understood as indicating or implying relative importance.

[0048] In the present application, the reference to "one embodiment" or "some embodiments" means that the specific features, structures or characteristics described in connection with the embodiment are included in one or more embodiments of the present application. Therefore, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in further some embodiments" and the like appearing in the present description are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "include", "contain", "have" and their variants mean "including but not limited to", unless otherwise specifically emphasized.

[0049] In view of the low accuracy of existing pathogenic gene prediction, the present application provides a pathogenic gene prediction method based on whole transcriptome association study. The pathogenic gene prediction method comprises the following steps: obtaining an RNA-seq sample containing a plurality of genes, calculating a splicing percentage matrix of each gene according to the RNA-seq sample, obtaining a genotype matrix of each gene, obtaining a genotype matrix of each splice factor, constructing a first prediction model based on the genotype matrix of all splice factors and the splicing percentage matrix of all genes, obtaining a predicted splicing percentage matrix of each gene by using the first prediction model, constructing a second prediction model based on all predicted splicing percentage matrices and the genotype matrix of all genes, obtaining a final predicted splicing percentage matrix of all genes by using the second prediction model, and finally performing pathogenicity analysis according to the final predicted splicing percentage matrix of all genes to obtain a gene pathogenicity prediction result. The construction of the prediction model based on the genotype matrix and the splicing percentage matrix can combine the information of the genotype and the splicing percentage of the gene, improve the accuracy and information richness of the result obtained by the prediction model, reveal the regulation mechanism between the genotype and the splicing event, and perform pathogenicity prediction according to the accurate final predicted splicing percentage matrix, which can capture the potential association of the splicing event with the pathogenicity, provide a new perspective for the research of pathogenicity prediction, and effectively improve the accuracy of gene pathogenicity prediction.

[0050] Some terms in the present application are described below.

[0051] Splice factors refer to proteins and small RNA molecules involved in the process of RNA splicing. They play an important role in the process of converting pre-mRNA (pre-mRNA) from the initial transcript to mature mRNA.

[0052] A splicing event refers to a splicing reaction between an exon and an intron in a pre-mRNA during post-transcriptional modification. Through a splicing event, the original pre-mRNA will remove the intron and connect the exons, thereby forming a mature mRNA.

[0053] An RNA-seq sample refers to a biological sample used for performing an RNA sequencing (RNA-seq) experiment.

[0054] A Cis-SNP (Cis-acting Single Nucleotide Polymorphism) site refers to a single nucleotide polymorphism (SNP) located near a certain gene or its regulatory region in the genome (usually upstream or downstream of the gene).

[0055] A single nucleotide polymorphism refers to a variation of a single base (such as A, T, C, or G) at a certain position in the genome, which usually has a certain variation frequency in the population.

[0056] Next, the pathogenic gene prediction method based on whole transcriptome association study provided by the present application is exemplarily described.

[0057] As shown in the following, Figure 1 The pathogenic gene prediction method based on whole transcriptome association study provided by the present application comprises the following steps:

[0058] Step 11, obtaining an RNA-seq sample containing a plurality of genes, and calculating a splicing percentage matrix of each gene according to the RNA-seq sample.

[0059] There are a plurality of splicing factors in all genes, and the splicing percentage matrix is used to describe the proportion of splicing events in the gene, and is specifically used to describe the relative proportion of exons spliced in the splicing event in the gene.

[0060] It should be noted that for a splicing factor with a specific function, it can exist in different genes.

[0061] For example, 80 brain RNA-seq samples of healthy humans and 84 brain RNA-seq samples of Alzheimer's disease patients were collected from the Mayo Clinic Brain Bank database; the human gene annotation file GRCh38 release77 was downloaded from the comprehensive database providing genome data (Ensembl database) as the input for calculating the splicing percentage matrix.

[0062] The splicing percentage PSI can be calculated by methods such as replicate Multivariate Analysis of Transcript Splicing (rMATS) using the RNA-seq samples collected above and the gene annotation file as input, and a PSI value is obtained for each gene a matrix (splicing percentage matrix), wherein is 164, representing the number of RNA-seq samples, represents the number of splicing events corresponding to the gene. The method for calculating PSI in this step is not limited to rMATS, nor is it limited to a single calculation method. For example, the proportion of Isoform expression in a gene can be used to calculate the PSI value, and the RNA-seq samples can be brain samples from other databases.

[0063] Step 12, obtain the genotype matrix of each gene and obtain the genotype matrix of each splicing factor.

[0064] The genotype matrix of the above gene is used to describe the single nucleotide polymorphism within a certain range upstream and downstream of the gene, and the genotype matrix of the splicing factor is used to describe the single nucleotide polymorphism within a certain range upstream and downstream of the position of the splicing factor.

[0065] For example, the SNP annotation file and the information of the gene in the gene annotation can be used to obtain the genotype matrix within a certain range (such as 5kb) upstream and downstream of the gene. 147 genes with splicing factors have been collected from the existing database, and the existing gene information is obtained from the gene annotation file GRCh38 release77. By locating the coordinates of these genes, the expression of single nucleotide polymorphisms (SNPs) within 5KB upstream and downstream of the genes in the samples can be found in the same version of SNP annotation information, and the genotype matrix can be obtained.

[0066] For the splicing factors obtained above, further screening can be performed by other tools to select splicing factors that have a strong correlation with the PSI value of the corresponding gene, such as the splicing factor participating in the splicing of the pre-RNA of the gene. In this way, the splicing factors selected according to the screening and the PSI value are modeled, which has stronger biological interpretability.

[0067] Step 13, based on the genotype matrix of all splicing factors and the splicing percentage matrix, a first prediction model is constructed, and the prediction splicing percentage matrix corresponding to all splicing factors is obtained by using the first prediction model.

[0068] Specifically, the first prediction model is:

[0069] Y = XSF ω+ε

[0070] wherein Y represents all splicing percentage matrices, X SF represents genotype matrices of all splicing factors, ω represents effect size, and ε represents error.

[0071] It should be noted that the step of obtaining the predicted splicing percentage matrix of each gene by using the first prediction model is specifically: taking all splicing percentage matrices as Y in the first prediction model, and taking the genotype matrix of each splicing factor as X SF For k splicing factors fitted with splicing percentage, n linear models (fitted first prediction model) and predicted values of k splicing percentages (obtained by inputting the preset genotype matrix into the linear model to calculate the splicing percentage Y) are obtained, that is, a Y n×k matrix (predicted splicing percentage matrix).

[0072] For example, in the above linear model, a least absolute shrinkage and selection operator (LASSO) regression is selected for fitting. The LASSO regression is a statistical method for regression modeling in data analysis, which realizes feature selection and regularization at the same time. The LASSO regression is suitable for the case where the number of features is greater than the number of samples and for the case where the feature correlation is high, and is suitable for the scene of PSI value prediction, and has the following basic characteristics:

[0073] The LASSO controls the complexity of the model by introducing an L1 regularization constraint (i.e., a penalty term of the sum of the absolute values of the coefficients), and the objective function is as follows:

[0074]

[0075] where y i is the target variable, x ij is the jth feature of the ith sample, β j represents the regression coefficient of the feature, β0 represents the coefficient, n represents the number of samples, p represents the number of features, and λ is the regularization strength, which controls the trade-off between model complexity and fitting accuracy.

[0076] Step 14, constructing a second prediction model based on the predicted splicing percentage matrix and the genotype matrix of all genes, and obtaining the final predicted splicing percentage matrix of all genes by using the second prediction model.

[0077] In some embodiments of the present application, the step of constructing a second prediction model based on the predicted splicing percentage matrix and the genotype matrix of all genes, and obtaining a final predicted splicing percentage matrix of all genes by using the second prediction model comprises:

[0078] Firstly, a plurality of initial second prediction models are constructed based on the predicted splicing percentage matrix and the genotype matrix of all genes.

[0079] Specifically, the initial second prediction model is:

[0080] Y = (X, Y SF ) β + ε

[0081] wherein Y represents the splicing percentage matrix, X represents the genotype matrix of all genes, Y SF represents the predicted splicing percentage matrix, β represents the weight, and ε represents the error.

[0082] Secondly, for each initial second prediction model, a prediction result comprising a plurality of splicing percentage prediction values is obtained by using the initial second prediction model for prediction, and a fitting value of the initial second prediction model is calculated according to the prediction result.

[0083] For example, the second prediction model is fitted by using LASSO regression, and then the preset genotype matrix and the predicted splicing percentage matrix are substituted into the fitted second prediction model for calculation to obtain the prediction result.

[0084] Specifically, the fitting value R

[0085]

[0086] is calculated by the formula: 2 .

[0087] wherein y i represents the i-th real splicing percentage, represents the i-th splicing percentage prediction value, represents the mean value of y i in all prediction results, and n represents the number of splicing percentage prediction values.

[0088] Thirdly, the initial second prediction model with a fitting value greater than or equal to a fitting threshold value is taken as the second prediction model.

[0089] It should be noted that if the initial second prediction models with the fitting value greater than or equal to the fitting threshold value are multiple, the initial second prediction model with the largest fitting value is selected as the second prediction model.

[0090] Fourthly, a final predicted splicing percentage matrix of all genes is obtained by using the second prediction model.

[0091] Specifically, for each gene, all splice percentages in the splice percentage matrix of the gene are clustered to obtain a plurality of clustering clusters. For each clustering cluster, if the number of splice percentages in the clustering cluster is greater than 1, multivariate linear regression is performed on the clustering cluster to obtain a plurality of final predicted splice percentages corresponding to the clustering cluster; if the number of splice percentages in the clustering cluster is equal to 1, the second prediction model is used to predict the splice percentage to obtain a final predicted splice percentage. Finally, all final predicted splice percentages are integrated to obtain a final predicted splice percentage matrix of the gene.

[0092] For example, clustering can be performed using a hierarchical clustering method. When predicting using the second prediction model, a genotype matrix previously set according to the actual condition of the gene sample is substituted into the second prediction model to obtain a final predicted splice percentage.

[0093] Step 15: Perform pathogenicity analysis according to the final predicted splice percentage matrix of all genes to obtain a gene pathogenicity prediction result.

[0094] The gene pathogenicity prediction result is used to describe the association between the splice factor in each gene, the corresponding single nucleotide polymorphism site and pathogenicity. For example, the gene pathogenicity prediction result can be a label 1 or 0, and the label 1 represents that the splice factor or single nucleotide polymorphism site is associated with the pathogenicity of the gene, and the label 0 represents that the splice factor or single nucleotide polymorphism site is not associated with the pathogenicity of the gene.

[0095] Specifically, the pathogenic label information of each gene is obtained, and then pathogenicity analysis is performed based on the pathogenic label information of all genes and the final predicted splice percentage matrix to obtain a gene pathogenicity prediction result.

[0096] The pathogenic label information is used to describe whether the gene has pathogenicity. For example, for the gene of the brain RNA-seq sample from the Alzheimer's disease patient in step 11, the pathogenic label information is: pathogenic; and for the gene of the brain RNA-seq sample from the healthy human in step 11, the pathogenic label information is: not pathogenic.

[0097] It should be noted that based on the pathogenic label information of all genes and the final predicted splice percentage matrix, the Bayes correction and significance test algorithm can be used to perform pathogenicity analysis to obtain a gene pathogenicity prediction result.

[0098] Exemplarily, the significant difference features are found by the Bayesian correction and the significance test, and the results of the significance test including the p value and the correction information are saved. The variable splicing events with significant expression difference between the Alzheimer's disease patients and the healthy people, the positions on the specific genes, and the correlations between the SNPs and the variable splicing events can be obtained.

[0099] After obtaining the gene pathogenicity prediction result, for the gene whose pathogenicity is not known, the correlation between the splicing factors in the gene and the pathogenicity described in the gene pathogenicity prediction result can be used to predict whether the gene has pathogenicity. If there is a correlation between the splicing factors of the gene and the pathogenicity, it is considered that the gene has pathogenicity.

[0100] It is worth mentioning that the prediction model is constructed based on the genotype matrix and the splicing percentage matrix, which can combine the information of the genotype and the splicing percentage of the gene, improve the accuracy and information richness of the results obtained by the prediction model, reveal the regulation mechanism between the genotype and the splicing event, and make pathogenicity prediction according to the accurate final prediction splicing percentage matrix. The potential correlation of the splicing event to the pathogenicity can be captured, a new perspective for the research of pathogenicity prediction is provided, and the accuracy of the gene pathogenicity prediction is effectively improved.

[0101] The device for predicting pathogenic genes based on whole transcriptome correlation research provided in the present application will be exemplarily described below.

[0102] As shown in Figure 2 , the device for predicting pathogenic genes based on whole transcriptome correlation research provided in the embodiments of the present application comprises:

[0103] The computing module 201 is configured to obtain an RNA-seq sample containing a plurality of genes, and calculate a splicing percentage matrix of each gene according to the RNA-seq sample. There are a plurality of splicing factors in all genes, and the splicing percentage matrix is used to describe the proportion of the splicing event in the gene.

[0104] The obtaining module 202 is configured to obtain a genotype matrix of each gene, and obtain a genotype matrix of each splicing factor. The genotype matrix of the gene is used to describe the single nucleotide polymorphism within a certain range upstream and downstream of the gene, and the genotype matrix of the splicing factor is used to describe the single nucleotide polymorphism within a certain range upstream and downstream of the position of the splicing factor.

[0105] The first constructing module 203 is configured to construct a first prediction model based on the genotype matrices of all splicing factors and all splicing percentage matrices, and obtain a prediction splicing percentage matrix corresponding to each splicing factor by using the first prediction model.

[0106] The second building module 204 is used to build a second prediction model based on the predicted splice percentage matrix and the genotype matrix of all genes, and to use the second prediction model to obtain the final predicted splice percentage matrix of all genes.

[0107] The pathogenicity analysis module 205 is used to perform pathogenicity analysis based on the final predicted splicing percentage matrix of all genes to obtain gene pathogenicity prediction results. The gene pathogenicity prediction results are used to describe the association between splicing factors, corresponding single nucleotide polymorphism sites and pathogenicity in each gene.

[0108] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0109] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0110] like Figure 3 As shown, an embodiment of this application provides a terminal device, wherein the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 3 The diagram shows only one processor, a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 executes the computer program D102 to implement the steps in any of the above method embodiments.

[0111] Specifically, the processor D100 executes the computer program D102 to obtain RNA-seq samples containing multiple genes, calculate the splicing percentage matrix of each gene according to the RNA-seq samples, then obtain the genotype matrix of each gene, obtain the genotype matrix of each splicing factor, construct a first prediction model based on the genotype matrix of all splicing factors and all splicing percentage matrices, obtain the predicted splicing percentage matrix of each gene by using the first prediction model, then construct a second prediction model based on all predicted splicing percentage matrices and the genotype matrix of all genes, obtain the final predicted splicing percentage matrix of all genes by using the second prediction model, and finally perform pathogenicity analysis according to the final predicted splicing percentage matrix of all genes to obtain the gene pathogenicity prediction result. The construction of the prediction model based on the genotype matrix and the splicing percentage matrix can combine the information of the genotype and the splicing percentage of the gene, improve the accuracy and information richness of the result obtained by the prediction model, reveal the regulation mechanism between the genotype and the splicing event, and perform pathogenicity prediction according to the accurate final predicted splicing percentage matrix, which can capture the potential correlation of the splicing event to the pathogenicity and provide a new perspective for the research on pathogenicity prediction, thereby effectively improving the accuracy of gene pathogenicity prediction.

[0112] The processor D100 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or can also be any conventional processor.

[0113] The storage D101 can be an internal storage unit of the terminal device D10 in some embodiments, such as a hard disk or a memory of the terminal device D10. The storage D101 can also be an external storage device of the terminal device D10 in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, and the like. Further, the storage D101 can include both an internal storage unit and an external storage device of the terminal device D10. The storage D101 is used to store an operating system, an application program, a boot loader, data, and other programs, such as program codes of the computer program, and the like. The storage D101 can also be used to temporarily store data that has been output or is to be output.

[0114] The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps in each of the above method embodiments.

[0115] The computer program product, when executed on a terminal device, causes the terminal device to implement the steps in each of the above method embodiments.

[0116] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the present application implements all or part of the processes in the above embodiments, which can be completed by instructing related hardware through a computer program. The computer program can be stored in a computer readable storage medium, and the computer program, when executed by a processor, can implement the steps in each of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms. The computer readable medium at least includes any entity or device capable of carrying the computer program code based on the whole transcriptome correlation study, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a U disk, a mobile hard disk, a magnetic disk or an optical disk, and the like. In some jurisdictions, according to legislation and patent practice, the computer readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0117] In the above embodiments, the description of each embodiment is focused on, and the part not described or recorded in a certain embodiment can be referred to the relevant description of other embodiments.

[0118] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software manner depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0119] The above is the preferred embodiment of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the principles described in the present application, a number of improvements and refinements can be made, which should be considered as the protection scope of the present application.

Claims

1. A method for predicting pathogenic genes based on whole transcriptome association studies, characterized in that, The method comprises the following steps: obtaining an RNA-seq sample containing a plurality of genes, and calculating a splicing percentage matrix of each gene according to the RNA-seq sample; a plurality of splicing factors exist in all genes, and the splicing percentage matrix is used to describe the proportion of splicing events in the genes; obtaining a genotype matrix of each gene, and obtaining a genotype matrix of each splicing factor; the genotype matrix of the gene is used to describe single nucleotide polymorphisms within a certain range upstream and downstream of the gene, and the genotype matrix of the splicing factor is used to describe single nucleotide polymorphisms within a certain range upstream and downstream of the position of the splicing factor; constructing a first prediction model based on the genotype matrix of all splicing factors and all splicing percentage matrices, and obtaining a predicted splicing percentage matrix corresponding to all splicing factors by using the first prediction model; constructing a second prediction model based on the predicted splicing percentage matrix and the genotype matrix of all genes, and obtaining a final predicted splicing percentage matrix of all genes by using the second prediction model; performing pathogenicity analysis according to the final predicted splicing percentage matrix of all genes to obtain a gene pathogenicity prediction result; the gene pathogenicity prediction result is used to describe the correlation between the splicing factor, the corresponding single nucleotide polymorphism site and the pathogenicity in each gene.

2. The disease causing gene prediction method of claim 1, wherein, The first prediction model is: Y = X SF ω + ε where Y represents all splicing percentage matrices, X SF represents the genotype matrix of all splicing factors, ω represents the effect size, and ε represents the error.

3. The disease causing gene prediction method of claim 1, wherein, The second prediction model is constructed based on the predicted splicing percentage matrix and the genotype matrix of all genes, which comprises: constructing a plurality of initial second prediction models based on the predicted splicing percentage matrix and the genotype matrix of all genes; for each initial second prediction model, performing prediction by using the initial second prediction model to obtain a prediction result comprising a plurality of splicing percentage prediction values, and calculating a fitting value of the initial second prediction model according to the prediction result; the initial second prediction model with a fitting value greater than or equal to a fitting threshold is used as the second prediction model.

4. The disease causing gene prediction method of claim 3, wherein, The initial second prediction model is: Y = (X, Y SF )β + ε where Y represents all splicing percentage matrices, X represents the genotype matrix of all genes, Y SF represents the predicted splicing percentage matrix, β represents the weight, and ε represents the error.

5. The disease causing gene prediction method of claim 4, wherein, The fitting value of the initial second prediction model is calculated according to the prediction result, which comprises: by the formula: Compute the fit value R 2 ; where y i represents the ith real splicing percentage, represents the ith splicing percentage prediction value, represents the mean of y i in all prediction results, and n represents the number of splicing percentage prediction values.

6. The disease causing gene prediction method of claim 1, wherein, The pathogenicity analysis according to the final predicted splicing percentage matrix of all genes to obtain a gene pathogenicity prediction result comprises: obtaining pathogenic label information of each gene; the pathogenic label information is used to describe whether the gene has pathogenicity; performing pathogenicity analysis based on the pathogenic label information of all genes and the final predicted splicing percentage matrix to obtain a gene pathogenicity prediction result.

7. A pathogenic gene prediction device based on whole transcriptome association studies, characterized in that, The method comprises the following steps: a calculation module is configured to obtain an RNA-seq sample containing a plurality of genes, and calculate a splicing percentage matrix of each gene according to the RNA-seq sample; a plurality of splicing factors exist in all genes, and the splicing percentage matrix is used to describe the proportion of splicing events in the genes; The acquisition module is configured to acquire a genotype matrix of each gene and a genotype matrix of each splicing factor, wherein the genotype matrix of the gene is used to describe single nucleotide polymorphisms within a certain range upstream and downstream of the gene, and the genotype matrix of the splicing factor is used to describe single nucleotide polymorphisms within a certain range upstream and downstream of the splicing factor; The first construction module is configured to construct a first prediction model based on the genotype matrices of all splicing factors and the splicing percentage matrices of all genes, and obtain a predicted splicing percentage matrix corresponding to each splicing factor by using the first prediction model; The second construction module is configured to construct a second prediction model based on the predicted splicing percentage matrix and the genotype matrices of all genes, and obtain a final predicted splicing percentage matrix of all genes by using the second prediction model; The pathogenicity analysis module is configured to perform pathogenicity analysis according to the final predicted splicing percentage matrix of all genes, and obtain a gene pathogenicity prediction result, wherein the gene pathogenicity prediction result is used to describe the association relationship between a splicing factor, a corresponding single nucleotide polymorphism site and pathogenicity in each gene.

8. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The computer program is executed by the processor to implement the pathogenic gene prediction method based on the whole transcriptome correlation study according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 8. The computer program is executed by the processor to implement the pathogenic gene prediction method based on the whole transcriptome correlation study according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Disease gene prediction method and device based on graph neural network

    CN118351945A

  • Predicting lung cancer survival using gene expression

    US20100267574A1