A method for processing gene expression data and related devices

By constructing a classification model, combining gene expression data and signaling pathways, the problem of inaccurate analysis of gene expression data is solved, and the accurate interpretation of biological information related to the target disease is achieved.

CN114678072BActive Publication Date: 2025-07-08NEUSOFT CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210243208.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-11
Publication Date
2025-07-08
Estimated Expiration
2042-03-11

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively analyze the gene expression data, resulting in inaccurate analysis results.

Method used

A classification model to be used is constructed, and the information analysis results of gene expression data under the target disease are determined using gene expression data, signal pathways and labeling information, and the contribution parameters of the signal pathway and gene characterization values.

Benefits of technology

It improves the accuracy of gene expression data analysis and the reliability of information analysis results, and can better explain the biological information related to the target disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114678072B_ABST
    Figure CN114678072B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a method for processing gene expression data and related devices. The method includes: after obtaining a large number of signal pathways, a large amount of gene expression data, and their annotation information under a target disease, a classification model to be used can be constructed first by using these signal pathways, this gene expression data, and their annotation information under the target disease, so that the classification model to be used has good classification performance under the target disease; then, based on the classification model to be used and these signal pathways, information analysis and processing are performed on these gene expression data to obtain an information analysis result of these gene expression data under the target disease, so that the information analysis result can accurately represent biological information related to the target disease. In this way, it is possible to mine biological information related to the target disease from a large amount of gene expression data, and thus it is possible to perform biological information analysis on a large amount of gene expression data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of medical data processing, and in particular to a gene expression data processing method and related equipment. Background Art

[0002] A large amount of gene expression data (for example, transcriptome sequencing (RNA-seq) data, etc.) has been accumulated in the databases of medical institutions. These data contain a large amount of biological information and have extremely high research value.

[0003] Among them, RNA-seq is a technology that combines experimental methods and computer means to determine the characteristic and abundance of ribonucleic acid (RNA) sequences in biological samples, so that RNA-seq data can be used for gene expression. In other words, the composition order of adenine, cytosine, guanine and uracil ribonucleic acid residues present in each single-stranded RNA molecule can be identified through RNA sequencing.

[0004] However, how to perform bioinformatics analysis on these gene expression data is a technical problem that needs to be solved urgently. Summary of the invention

[0005] In view of this, the embodiments of the present application provide a gene expression data processing method and related equipment, which can realize bioinformatics analysis of a large amount of gene expression data.

[0006] To solve the above problems, the technical solutions provided in the embodiments of the present application are as follows:

[0007] The present application provides a method for processing gene expression data, the method comprising:

[0008] Acquire at least one gene expression data, annotation information of the at least one gene expression data under a target disease, and at least one signal pathway;

[0009] Constructing a classification model to be used by using the at least one gene expression data, the annotation information of the at least one gene expression data under the target disease, and the at least one signal pathway;

[0010] According to the classification model to be used and the at least one signal pathway, an information analysis result of the at least one gene expression data under the target disease is determined.

[0011] In a possible implementation, determining the information analysis result of the at least one gene expression data under the target disease according to the classification model to be used and the at least one signal pathway includes:

[0012] Determine the contribution parameters of the at least one signaling pathway from the classification model to be used;

[0013] Using the contribution parameters of the at least one signaling pathway, determine the information analysis result of the at least one gene expression data under the target disease.

[0014] In a possible implementation manner, the classification model to be used includes a pathway layer; the pathway layer is used to determine the pathway contribution value of the gene expression data under each signaling pathway;

[0015] The determining the contribution parameters of the at least one signaling pathway from the classification model to be used includes:

[0016] Determine the contribution parameters of the at least one signaling pathway from the layer parameters of the pathway layer.

[0017] In a possible implementation manner, the pathway layer belongs to the shallow network of the classification model to be used.

[0018] In a possible implementation manner, the contribution parameters include a weighting parameter and a bias parameter;

[0019] The using the contribution parameters of the at least one signaling pathway to determine the information analysis result of the at least one gene expression data under the target disease includes:

[0020] According to the contribution parameters of each signaling pathway, the gene expression samples corresponding to the at least one gene expression data, and the annotation information of the at least one gene expression data under the target disease, determine the contribution analysis result of each signaling pathway to the target disease;

[0021] According to the weighting parameters of the at least one signaling pathway and the at least one gene expression data, determine the contribution analysis result of at least one gene to the target disease;

[0022] According to the contribution analysis result of the at least one signaling pathway to the target disease and the contribution analysis result of the at least one gene to the target disease, determine the information analysis result of the at least one gene expression data under the target disease.

[0023] In a possible implementation manner, the number of signaling pathways is M; the number of gene expression data is N; where M is a positive integer and N is a positive integer;

[0024] The determination process of the contribution analysis result of the m-th signaling pathway to the target disease includes:

[0025] Determine the pathway contribution value of the nth gene expression data under the mth signal pathway according to the contribution parameter of the mth signal pathway and the gene expression sample corresponding to the nth gene expression data; where m is a positive integer, m ≤ M;

[0026] Determine the contribution analysis result of the mth signal pathway to the target disease according to the pathway contribution values of N gene expression data under the mth signal pathway and the annotation information of the N gene expression data for the target disease.

[0027] In a possible implementation manner, the method further includes:

[0028] Determine at least one ith data sample from the N gene expression data according to the annotation information of the N gene expression data for the target disease; where i is a positive integer, i ≤ I, I is a positive integer, and I represents the number of types of the data samples;

[0029] The step of determining the contribution analysis result of the mth signal pathway to the target disease according to the pathway contribution values of N gene expression data under the mth signal pathway and the annotation information of the N gene expression data for the target disease includes:

[0030] Perform a preset statistical process on the pathway contribution values of the at least one ith data sample under the mth signal pathway to obtain the ith statistical contribution value of the mth signal pathway; where i is a positive integer, i ≤ I;

[0031] Perform a preset analysis process on the 1st to Ith statistical contribution values of the mth signal pathway to obtain the contribution analysis result of the mth signal pathway to the target disease.

[0032] In a possible implementation manner, the method further includes:

[0033] Determine at least one gene usage pathway corresponding to each of the genes from the at least one signal pathway;

[0034] The step of determining the contribution analysis result of at least one gene to the target disease according to the weighted parameters of the at least one signal pathway and the at least one gene expression data includes:

[0035] Determine the contribution analysis result of each of the genes to the target disease according to the weighted parameters of at least one gene usage pathway corresponding to each of the genes and the gene characterization values of each of the genes in the at least one gene expression data.

[0036] In a possible implementation, the gene expression data includes gene characterization values of K genes; the number of the gene expression data is N; where K is a positive integer and N is a positive integer;

[0037] The process of determining the contribution analysis result of the k-th gene to the target disease includes:

[0038] According to the weighted parameters of at least one gene usage pathway corresponding to the k-th gene and the gene characterization value of the k-th gene in the n-th gene expression data, determine the gene contribution value of the n-th gene expression data under the k-th gene; where k is a positive integer, k ≤ K, n is a positive integer, n ≤ N;

[0039] Average the gene contribution values of the N gene expression data under the k-th gene to obtain the contribution characterization value of the k-th gene under the target disease;

[0040] According to the contribution characterization value of the k-th gene under the target disease, determine the contribution analysis result of the k-th gene to the target disease.

[0041] In a possible implementation, constructing a classification model to be used by using the at least one gene expression data, the annotation information of the at least one gene expression data under the target disease, and at least one signaling pathway includes:

[0042] Perform data mapping processing on each of the gene expression data according to each of the signaling pathways to obtain gene expression samples corresponding to each of the gene expression data;

[0043] Construct a classification model to be used by using the gene expression samples corresponding to the at least one gene expression data and the annotation information of the at least one gene expression data under the target disease.

[0044] In a possible implementation, constructing a classification model to be used by using the gene expression samples corresponding to the at least one gene expression data and the annotation information of the at least one gene expression data under the target disease includes:

[0045] According to the gene expression samples corresponding to the at least one gene expression data and the annotation information of the at least one gene expression data under the target disease, determine at least one training sample, the label information of the at least one training sample, at least one test sample, and the label information of the at least one test sample;

[0046] Update the classification model to be processed by using the at least one training sample and the label information of the at least one training sample;

[0047] Using the at least one test sample and the label information of the at least one test sample, test the classification model to be processed, obtain a model test result, and continue to execute the step of updating the classification model to be processed by using the at least one training sample and the label information of the at least one training sample, until when it is determined that a first preset condition is met, determine the classification model to be used according to the classification model to be processed.

[0048] In a possible implementation manner, the step of updating the classification model to be processed by using the at least one training sample and the label information of the at least one training sample includes:

[0049] Using the classification model to be processed, determine the model classification results of each training sample;

[0050] According to the model classification results of the at least one training sample and the label information of the at least one training sample, update the classification model to be processed, and continue to execute the step of using the classification model to be processed to determine the model classification results of each training sample;

[0051] The step of testing the classification model to be processed by using the at least one test sample and the label information of the at least one test sample to obtain a model test result includes:

[0052] When it is determined that a second preset condition is reached, use the at least one test sample and the label information of the at least one test sample to test the classification model to be processed, and obtain a model test result.

[0053] In a possible implementation manner, the number of signal pathways is M; the gene expression sample includes a set of characterization values corresponding to M signal pathways; the set of characterization values corresponding to the m-th signal pathway includes the gene characterization values of at least one gene in the gene expression data; where m is a positive integer, m ≤ M, and M is a positive integer.

[0054] The embodiments of the present application further provide a gene expression data processing device, and the device includes:

[0055] An information acquisition unit, configured to acquire at least one gene expression data, the annotation information of the at least one gene expression data under a target disease, and at least one signal pathway;

[0056] A model construction unit, configured to construct a classification model to be used by using the at least one gene expression data, the annotation information of the at least one gene expression data under a target disease, and the at least one signal pathway;

[0057] An information analysis unit for determining an information analysis result of the at least one gene expression data under the target disease according to the to-be-used classification model and the at least one signaling pathway.

[0058] An embodiment of the present application also provides a gene expression data processing device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, any implementation manner of the gene expression data processing method provided in the embodiment of the present application is implemented.

[0059] An embodiment of the present application also provides a computer-readable storage medium, in which instructions are stored. When the instructions run on a terminal device, the terminal device is caused to execute any implementation manner of the gene expression data processing method provided in the embodiment of the present application.

[0060] An embodiment of the present application also provides a computer program product. When the computer program product runs on a terminal device, the terminal device is caused to execute any implementation manner of the gene expression data processing method provided in the embodiment of the present application.

[0061] Thus, the embodiments of the present application have the following beneficial effects:

[0062] In the technical solution provided in the embodiment of the present application, after obtaining a large number of signaling pathways, a large number of gene expression data, and their annotation information under the target disease, a to-be-used classification model can be constructed first by using these signaling pathways, these gene expression data, and their annotation information under the target disease, so that the to-be-used classification model has good classification performance under the target disease, and thus the to-be-used classification model carries some biological information required for understanding and analyzing these gene expression data under the target disease; then, based on the to-be-used classification model and these signaling pathways, information analysis processing is performed on these gene expression data to obtain an information analysis result of these gene expression data under the target disease, so that the information analysis result can accurately represent the biological information related to the target disease (for example, which channels and / or which genes are more likely to cause cancer). In this way, biological information related to the target disease (for example, influencing factors of the target disease) can be mined from a large number of gene expression data, and thus biological information analysis of a large number of gene expression data can be achieved.

[0063] In addition, during the construction of the classification model to be used, not only a large amount of gene expression data and its annotation information under the target disease are referred to, but also a large number of signaling pathways are referred to. This enables the classification model to be used to have better prediction performance, and thus have better classification performance under the target disease. Furthermore, the classification model to be used can better represent the biological information required for understanding and analyzing these gene expression data under the target disease. In this way, the information analysis results obtained based on the analysis of the classification model to be used are more accurate, which is conducive to improving the accuracy of the information analysis results. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 It is a flowchart of a method for processing gene expression data provided by an embodiment of the present application;

[0065] Figure 2 It is a schematic diagram of a gene expression matrix provided by an embodiment of the present application;

[0066] Figure 3 It is a schematic diagram of annotation information provided by an embodiment of the present application;

[0067] Figure 4 It is a schematic diagram of a signaling pathway database provided by an embodiment of the present application;

[0068] Figure 5 It is a schematic diagram of the working principle of a classification model to be used provided by an embodiment of the present application;

[0069] Figure 6 It is a schematic diagram of the analysis results of the contributions of multiple signaling pathways to the target disease provided by an embodiment of the present application;

[0070] Figure 7 It is a schematic diagram of the structure of a gene expression data processing device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] To make the above objects, features, and advantages of the present application more obvious and understandable, the following further describes the embodiments of the present application in detail with reference to the drawings and specific embodiments.

[0072] In the research on gene expression data, the inventors found that in some cases, data analysis methods such as clustering algorithms and gene differential expression can be used to perform biological information analysis on a large amount of gene expression data. However, due to the defects of the above data analysis methods, a large amount of information will be lost when using these data analysis methods to perform biological information analysis on a large amount of gene expression data. As a result, the biological information obtained based on the analysis of these data analysis methods is inaccurate, and further leads to the urgent need to solve the technical problem of how to perform biological information analysis on these gene expression data.

[0073] Based on the above findings, to solve the technical problems shown in the background art section, an embodiment of the present application provides a method for processing gene expression data. The method includes: after obtaining a large number of signal pathways, a large amount of gene expression data, and their annotation information under a target disease, a classification model to be used can be constructed first using these signal pathways, this gene expression data, and their annotation information under the target disease, so that the classification model to be used has better classification performance under the target disease, and thus the classification model to be used carries some biological information required for understanding and analyzing these gene expression data under the target disease; then, based on the classification model to be used and these signal pathways, information analysis and processing are performed on these gene expression data to obtain an information analysis result of these gene expression data under the target disease, so that the information analysis result can accurately represent biological information related to the target disease (for example, which channels and / or which genes are more likely to cause cancer). In this way, biological information related to the target disease (for example, influencing factors of the target disease) can be mined from a large amount of gene expression data, and thus biological information analysis of a large amount of gene expression data can be achieved.

[0074] It can be seen that since the above information analysis result is obtained by analyzing all gene expression data, the information analysis result can comprehensively represent the biological information related to the target disease carried by these gene expression data. In this way, the adverse effects caused by the above information loss phenomenon can be effectively avoided, which is beneficial to improving the accuracy of biological information analysis, and further beneficial to improving the biological information analysis effect for a large amount of gene expression data.

[0075] In addition, since the construction process of the classification model to be used not only refers to a large amount of gene expression data and their annotation information under the target disease, but also refers to a large number of signal pathways, the classification model to be used has better prediction performance, so that the classification model to be used has better classification performance under the target disease, and further enables the classification model to be used to better represent the biological information required for understanding and analyzing these gene expression data under the target disease. In this way, the information analysis result obtained based on the classification model to be used is more accurate, which is beneficial to improving the accuracy of the information analysis result.

[0076] In addition, the execution subject of the gene expression data processing method in the embodiments of the present application is not limited. For example, the gene expression data processing method provided in the embodiments of the present application can be applied to data processing devices such as terminal devices or servers. Among them, the terminal device can be a smart phone, a computer, a personal digital assistant (PDA), or a tablet computer, etc. The server can be an independent server, a cluster server, or a cloud server.

[0077] To facilitate the understanding of the present application, the gene expression data processing method provided in the embodiments of the present application will be described below with reference to the accompanying drawings.

[0078] See Figure 1 , which is a flowchart of a gene expression data processing method provided in the embodiments of the present application. The gene expression data processing method may include S1 - S3:

[0079] S1: Obtain at least one gene expression data, the annotation information of the at least one gene expression data under a target disease, and at least one signaling pathway.

[0080] Among them, the gene expression data is used to represent the gene activity information in the body of an object (such as a person, an animal, etc.), so that the gene expression data can reflect the physiological state of the cells in the body of the object (for example, whether the cell is in a normal or deteriorated state, whether the drug is effective for the tumor cell, etc.).

[0081] In addition, the gene expression data in the embodiments of the present application is not limited. For example, it may include the gene characterization values of at least one gene; moreover, the representation method of the gene expression data in the embodiments of the present application is not limited. For example, it can be represented in the form of a vector (such as, [1.882, 1.361, 1.723,...]).

[0082] It should be noted that the above "at least one gene" in the embodiments of the present application is not limited. For example, it may include Figure 2 genes such as DPM1, FGR,... shown in Figure 2 . In addition, the above "gene characterization value" in the embodiments of the present application is not limited. For example, it may be

[0083] 1.882 shown in

[0084] It should be noted that the embodiments of the present application do not limit the implementation manner of "determining at least one gene expression data from these RNA-seq data". For example, if these RNA-seq data are all binary probe data, in order to facilitate the representation of the gene information carried by these RNA-seq data, some data conversion methods (such as RMA (Robust Multi-Array Average), MAS5.0 (MicroArray Suite 5.0), etc.) can be used to perform conversion processing on these probe data to obtain at least one gene expression data. Another example is that if these RNA-seq data are represented according to the Figure 2 gene expression matrix shown, then the respective gene expression data can be directly extracted from this expression matrix.

[0085] The annotation information of the nth gene expression data under the target disease is used to describe the actual state of an object with the nth gene expression data under the target disease (for example, whether having cancer or the cancer stage). Here, n is a positive integer, n ≤ N, N is a positive integer, and N represents the number of gene expression data. It should be noted that the embodiments of the present application do not limit the target disease. For example, it can be cancer.

[0086] In addition, the embodiments of the present application do not limit the representation manner of the above "annotation information of the nth gene expression data under the target disease". For example, it can be represented in the manner Figure 3 shown. It should be noted that for Figure 3 if the annotation information is 0, it means not having cancer; if the annotation information is 1, it means cancer stage 1; if the annotation information is 2, it means cancer stage 2.

[0087] Furthermore, the embodiments of the present application do not limit the acquisition manner of the above "annotation information of the nth gene expression data under the target disease". For example, it can be manually annotated by medical staff. Also, it can be automatically extracted from the diagnosis and treatment documents.

[0088] A signal pathway refers to the phenomenon that when a certain reaction is to occur in a cell, a signal is transmitted from outside the cell to inside the cell to convey a piece of information, so that the cell makes a reaction according to this information.

[0089] In addition, the embodiments of the present application do not limit the representation manner of the signal pathway. For example, as Figure 4 shown, a signal pathway can be represented by the pathway name of the signal pathway, the pathway information website of the signal pathway, and the pathway member genes of the signal pathway (that is, all gene names involved in the signal pathway).

[0090] In addition, the embodiments of the present application do not limit the acquisition method of the above "at least one signal pathway". For example, it can be extracted from any existing or future signal pathway database (such as the BIOCARTA signal pathway database, etc.).

[0091] Based on the relevant content of S1 above, it can be known that if one wants to analyze the biological information related to a target disease from a large amount of gene expression data, it is necessary to pre-acquire a large number of signal pathways, a large amount of gene expression data, and the annotation information of these gene expression data under the target disease, so that subsequent research and analysis of these gene expression data can be realized based on these three.

[0092] S2: Construct a classification model to be used by using at least one gene expression data, the annotation information of the at least one gene expression data under the target disease, and at least one signal pathway.

[0093] Among them, the classification model to be used is used to perform classification processing under the target disease for the input data of the classification model to be used (for example, whether suffering from cancer, or what stage of cancer, etc.).

[0094] In addition, the embodiments of the present application do not limit the classification model to be used. For example, it can be any machine learning model (such as a classification model).

[0095] Furthermore, the embodiments of the present application do not limit the construction process of the classification model to be used. For example, it can be implemented by using any of the following implementation methods for constructing the classification model to be used.

[0096] Based on the relevant content of S2 above, it can be known that after obtaining at least one gene expression data, the annotation information of the at least one gene expression data under the target disease, and at least one signal pathway, these signal pathways, these gene expression data, and the annotation information of these gene expression data under the target disease can be used to construct a classification model to be used, so that the classification model to be used can learn the association relationship between genes, signal pathways, and annotation information, so that the classification model to be used has better classification performance under the target disease, and further enables the classification model to be used to better express the classification rules of the target disease, so that the classification model to be used can effectively represent the biological information required for understanding and analyzing these gene expression data under the target disease, so that the classification model to be used can be used subsequently to interpret these gene expression data.

[0097] S3: Determine the information analysis result of at least one gene expression data under the target disease according to the classification model to be used and at least one signal pathway.

[0098] Among them, the information analysis result is used to represent the biological information related to the target disease carried by the above-mentioned "at least one gene expression data" (for example, biological information such as which channels and / or which genes are more likely to cause the target disease).

[0099] It can be seen that in the embodiment of the present application, after obtaining the classification model to be used, the classification rules of the target disease expressed by the classification model to be used and a large number of signal pathways can be referred to, and the biological information of a large number of gene expression data can be interpreted to obtain the information analysis result of these gene expression data under the target disease, so that the information analysis result can represent the biological information related to the target disease (such as which channels and / or which genes are more likely to cause cancer, etc.). In this way, the influencing factors of the target disease can be mined from a large number of gene expression data, and thus the analysis and processing of a large number of gene expression data can be realized.

[0100] In addition, the embodiment of the present application does not limit the implementation manner of S3. For example, it can be implemented by any of the implementation manners of S3 shown below.

[0101] Based on the relevant content of S1 to S3 above, for the gene expression data processing method provided by the embodiment of the present application, after obtaining a large number of signal pathways, a large number of gene expression data and their annotation information under the target disease, a classification model to be used can be constructed first by using these signal pathways, these gene expression data and their annotation information under the target disease, so that the classification model to be used has good classification performance under the target disease, and thus the classification model to be used carries some biological information required for understanding and analyzing these gene expression data under the target disease; then, based on the classification model to be used and these signal pathways, information analysis and processing are performed on these gene expression data to obtain the information analysis result of these gene expression data under the target disease, so that the information analysis result can accurately represent the biological information related to the target disease (for example, which channels and / or which genes are more likely to cause cancer). In this way, the biological information related to the target disease (such as the influencing factors of the target disease) can be mined from a large number of gene expression data, and thus the biological information analysis of a large number of gene expression data can be realized.

[0102] In addition, during the construction of the classification model to be used, a large amount of gene expression data and their annotation information under the target disease are not only referred to, but also a large number of signaling pathways are referred to. As a result, the classification model to be used has better prediction performance, so that the classification model to be used has better classification performance under the target disease. Furthermore, the classification model to be used can better represent the biological information required for understanding and analyzing these gene expression data under the target disease. In this way, the information analysis results obtained based on the classification model to be used are more accurate, which is beneficial to improving the accuracy of the information analysis results.

[0103] In fact, although the genes involved in different signaling pathways are different, there may be intersections between the genes involved in some signaling pathways. For example, for Figure 5 the shown signaling pathways P1 and P2, there is an intersection between the pathway member genes of signaling pathway P1 and the pathway member genes of signaling pathway P2, and this intersection includes gene G2 and gene G3.

[0104] Based on the characteristics of the above-mentioned signaling pathways, in order to further improve the classification performance of the classification model to be used, an embodiment of the present application also provides a possible implementation manner of the model structure of the classification model to be used. In this implementation manner, the classification model to be used may include a pathway layer, a splicing layer, D hidden layers, and an output layer. Wherein, D is a positive integer. For the convenience of understanding, the following will introduce each layer separately.

[0105] Pathway layer

[0106] Each node in the pathway layer represents a signaling pathway. It can be seen that the number of nodes in the pathway layer is the same as the number of signaling pathways of "at least one signaling pathway" above.

[0107] Each node in the pathway layer is only connected to the input nodes of the pathway member genes of the signaling pathway represented by the node, so that the node is only used to process the gene representation values of the pathway member genes. For example, when the m-th node in the pathway layer represents the m-th signaling pathway, and the pathway member genes of the m-th signaling pathway include the first gene (such as Figure 5 the shown G1 gene), the second gene (such as Figure 5 the shown G2 gene), and the third gene (such as the shown G3 gene), the m-th node is only connected to the input node of the first gene, the input node of the second gene, and the input node of the third gene, so that the m-th node is only used to process the gene representation values of the first gene, the second gene, and the third gene in the input data of the pathway layer. Wherein, m is a positive integer, m ≤ M, M is a positive integer, and M represents the number of nodes in the pathway layer (that is, the number of signaling pathways).

[0108] In addition, the pathway layer is used to determine the pathway contribution value of the input data of the pathway layer under each signal pathway; moreover, the embodiments of the present application do not limit the working principle of the pathway layer. For example, when the m-th node in the pathway layer represents the m-th signal pathway, the m-th node in the pathway layer can be implemented using the working principle shown in formula (1). Where m is a positive integer and m ≤ M.

[0109]

[0110] In the formula, p m represents the pathway contribution value of the input data of the pathway layer under the m-th signal pathway (that is, the output data of the m-th node in the pathway layer); x represents the input data of the pathway layer; g m (x) represents the gene characterization value of the pathway member genes of the m-th signal pathway extracted from x, and t m represents the number of genes in the pathway member genes of the m-th signal pathway (that is, the number of genes involved in the m-th signal pathway); Relu(·) represents the linear rectification function; represents the weighted parameter of the m-th signal pathway, and represents the bias parameter of the m-th signal pathway, and both belong to the layer parameters used by the m-th node in the pathway layer, and both can be determined during the construction of the classification model to be used.

[0111] It should be noted that the working principle of each node in the pathway layer is similar to that of the m-th node above. For the sake of brevity, it will not be elaborated here.

[0112] Concatenation layer

[0113] The concatenation layer is used to perform concatenation processing on the input data of the concatenation layer; and the input data of the concatenation layer includes the output data of the pathway layer. It can be seen that the concatenation layer is used to perform concatenation processing on the output data of the pathway layer; and the working principle of the concatenation layer is shown in formula (2).

[0114]

[0115] In the formula, P represents the output data of the concatenation layer (that is, the concatenation vector of the output data of the pathway layer); p m represents the pathway contribution value of the input data of the pathway layer under the m-th signal pathway (that is, the output data of the m-th node in the pathway layer), m is a positive integer, and m ≤ M.

[0116] Hidden layer

[0117] The input data of the first hidden layer includes the output data of the splicing layer; the input data of the d-th hidden layer includes the output data of the (d - 1)-th hidden layer, where d is a positive integer and 2 ≤ d ≤ D.

[0118] In addition, the embodiments of the present application do not limit the implementation manner of the D hidden layers. For example, it can be implemented using D fully connected layers.

[0119] Furthermore, the embodiments of the present application do not limit the working principle of the D hidden layers. For example, it can be implemented using formulas (3)-(4).

[0120]

[0121]

[0122] In the formula, h1 represents the output data of the first hidden layer; P represents the output data of the splicing layer (that is, the input data of the first hidden layer); represents the weighted parameter of the first hidden layer; represents the bias parameter of the first hidden layer; both belong to the layer parameters of the first hidden layer; h d represents the output data of the d-th hidden layer; h d-1 represents the output data of the (d - 1)-th hidden layer (that is, the input data of the d-th hidden layer); represents the weighted parameter of the d-th hidden layer; represents the bias parameter of the d-th hidden layer; both belong to the layer parameters of the d-th hidden layer; d is a positive integer and 2 ≤ d ≤ D; both can be determined during the construction of the classification model to be used; Relu(·) represents the rectified linear unit function.

[0123] Output layer

[0124] The input data of the output layer includes the output data of the D-th hidden layer; and the embodiments of the present application do not limit the working principle of this output layer. For example, it can be implemented using formula (5).

[0125] out = softmax(W out h D + b out ) (5)

[0126] In the formula, out represents the output data of the output layer; h D represents the output data of the D-th hidden layer (that is, the input data of the output layer); Wout Represents the weighted parameter of the output layer; b out Represents the bias parameter of the output layer; W out and b out Both belong to the layer parameters of the output layer, and W out and b out Both can be determined during the construction of the classification model to be used; softmax(·) represents the softmax logistic regression function.

[0127] Based on the relevant content of the model structure of the above-mentioned classification model to be used, for this classification model to be used, after obtaining the input data of this classification model to be used, first, the path layer determines the path contribution value of the input data under each signal path; then, the splicing layer splices the path contribution values under all signal paths; then, D hidden layers perform a fully connected process on the splicing result; finally, the output layer performs a classification process on the fully connected process result to obtain the output data of this classification model to be used, so that the output data can represent the predicted classification result of the input data under the target disease (for example, whether having cancer, or which stage of cancer, etc.).

[0128] Actually, since each node in the path layer represents a signal path, making the genes concerned by each node in this path layer fixed, therefore, in order to better adapt to the characteristics of this path layer, the embodiment of the present application also provides a possible implementation manner for constructing the classification model to be used (that is, a possible implementation manner of S2), which specifically may include S21 - S22:

[0129] S21: Perform data mapping processing on each gene expression data according to each signal path to obtain gene expression samples corresponding to each gene expression data.

[0130] Among them, the gene expression sample corresponding to the nth gene expression data is used to represent the gene expression information of the nth gene expression data under each signal path. n is a positive integer, n ≤ N, N is a positive integer, and N represents the number of gene expression data.

[0131] In addition, the embodiment of the present application does not limit the above-mentioned "gene expression sample corresponding to the nth gene expression data". For example, when the number of signal paths is M, the "gene expression sample corresponding to the nth gene expression data" may include a set of characterization values corresponding to M signal paths (such as formulas (6) - (7)), and each set of characterization values corresponding to the signal paths includes the gene characterization values of at least one gene in the nth gene expression data.

[0132]

[0133]

[0134] Wherein, S n represents the gene expression sample corresponding to the nth gene expression data; represents the set of characterization values corresponding to the mth signaling pathway in the above-mentioned "gene expression sample corresponding to the nth gene expression data", so that the can represent the gene expression information of the nth gene expression data under the mth signaling pathway, and the has a dimension equal to the number of genes in the pathway member genes of the mth signaling pathway (that is, the number of genes involved in the mth signaling pathway); E n represents the nth gene expression data; g m (E n ) represents extracting the gene characterization values of the pathway member genes of the mth signaling pathway from E n ; n is a positive integer, n ≤ N, N is a positive integer, and N represents the number of gene expression data; m is a positive integer, m ≤ M, M is a positive integer, and M represents the number of signaling pathways.

[0135] It should be noted that the embodiments of the present application do not limit the above For example, it can be represented in a vector manner, so that the can represent the gene characterization vector of the nth gene expression data under the mth signaling pathway.

[0136] Based on the relevant content of S21, after obtaining at least one gene expression data and at least one signaling pathway, the gene expression samples corresponding to each gene expression data can be obtained by referring to the pathway member genes of all signaling pathways, so that the gene expression samples can represent the gene expression information of the gene expression data under each signaling pathway.

[0137] S22: Construct a classification model to be used by using the gene expression samples corresponding to at least one gene expression data and the annotation information of the at least one gene expression data under the target disease.

[0138] In the embodiments of the present application, after obtaining the gene expression samples corresponding to a large number of gene expression data and the annotation information of these gene expression data under the target disease, these gene expression samples can be used as model input data, and these annotation information can be used as model guiding information (that is, prior information) to construct the classification model to be used, so that the constructed classification model to be used can learn the association relationship between the gene expression samples and the annotation information.

[0139] In addition, the embodiments of the present application do not limit the implementation manner of S22. For example, any existing or future model construction method can be used for implementation.

[0140] In addition, to further improve the classification performance of the classification model to be used, another possible implementation manner of S22 is provided in the embodiments of the present application, which may specifically include S221-S225:

[0141] S221: Determine at least one training sample, label information of the at least one training sample, at least one test sample, and label information of the at least one test sample according to gene expression samples corresponding to at least one gene expression data and annotation information of the at least one gene expression data under a target disease.

[0142] Among them, the training sample is used to represent the gene expression sample used in the model training process.

[0143] The label information of the training sample refers to the annotation information of the gene expression data represented by the training sample under the target disease, so that the "label information of the training sample" is used to represent the actual state of the object with the training sample under the target disease (for example, whether having cancer or the cancer stage); and the acquisition process of the "label information of the training sample" can specifically be: after determining a gene expression sample corresponding to a gene expression data as a training sample, the annotation information of the gene expression data under the target disease can be directly determined as the label information of the training sample.

[0144] The test sample is used to represent the gene expression sample used in the model testing process.

[0145] The label information of the test sample refers to the annotation information of the gene expression data represented by the test sample under the target disease, so that the "label information of the test sample" is used to represent the actual state of the object with the test sample under the target disease (for example, whether having cancer or the cancer stage); and the acquisition process of the "label information of the test sample" can specifically be: after determining a gene expression sample corresponding to a gene expression data as a test sample, the annotation information of the gene expression data under the target disease can be directly determined as the label information of the test sample.

[0146] In addition, the embodiments of the present application do not limit the number of training samples and the number of test samples. For example, to ensure that the model has high interpretability and generalization ability, the following relationship can be satisfied between the two: the number of training samples = the number of test samples, and the number of training samples + the number of test samples = the number of gene expression samples (that is, the number of gene expression data).

[0147] Based on the relevant content of S221 above, after obtaining the gene expression samples corresponding to at least one gene expression data, these gene expression samples can be divided into a training set (i.e., the above-mentioned "at least one training sample") and a test set (i.e., the above-mentioned "at least one test sample") according to a preset ratio (e.g., 1:1) set in advance, so that the training set is used to train the model and the test set is used to test the model, so as to subsequently construct a classification model to be used based on the training set and the test set.

[0148] S222: Update the classification model to be processed by using at least one training sample and the label information of the at least one training sample.

[0149] Among them, the classification model to be processed is used to represent the classification model that needs to be trained and updated; and the model structure of the classification model to be processed is the same as the model structure of the above-mentioned "classification model to be used". For the sake of brevity, it will not be elaborated here.

[0150] In addition, the working principle of the classification model to be processed is the same as the working principle of the above-mentioned "classification model to be used" (e.g., it can be implemented by using the working principle shown in formulas (1)-(5)). For the sake of brevity, it will not be elaborated here.

[0151] In addition, the embodiments of the present application do not limit the update process of the classification model to be processed (i.e., the implementation manner of S222). For example, it may specifically include S2221-S2222:

[0152] S2221: Use the classification model to be processed to determine the model classification results of each training sample.

[0153] Among them, the model classification result of the qth training sample is used to represent the predicted state of the qth training sample under the target disease (e.g., whether having cancer or the cancer stage). q is a positive integer, q ≤ Q, Q is a positive integer, and Q represents the number of training samples.

[0154] In addition, the embodiments of the present application do not limit the acquisition process of the above-mentioned "model classification result of the qth training sample". For example, it may specifically be: input the qth training sample into the classification model to be processed to obtain the model classification result of the qth training sample output by the classification model to be processed.

[0155] S2222: Update the classification model to be processed according to the model classification results of at least one training sample and the label information of the at least one training sample.

[0156] In the embodiments of the present application, after obtaining the model classification results of at least one training sample, the model loss value of the classification model to be processed can be determined first by using the model classification results of these training samples and the label information, so that the model loss value can represent the classification performance of the classification model to be processed; then, based on the model loss value, the classification model to be processed is updated to obtain an updated classification model to be processed, so that the updated classification model to be processed has better classification performance.

[0157] It should be noted that the embodiments of the present application do not limit the determination process of the "model loss value of the classification model to be processed" described above. For example, it can be implemented by using any existing or future loss function (such as, cross-entropy loss function, etc.). In addition, the embodiments of the present application do not limit the update process of the "classification model to be processed" described above. For example, it can be implemented by using any existing or future model update method (such as, Adaptive moment estimation (adam) optimizer).

[0158] Based on the relevant content of S2221 to S2222 above, it can be known that at least one training sample and its label information can be used to implement a round of update process for the classification model to be processed, so that the classification performance of the classification model to be processed after a round of update can be tested subsequently, and thus a model training mode of one update and one test can be achieved.

[0159] In fact, for a round of update process of the classification model to be processed, it usually slightly improves the model classification performance. Therefore, in order to effectively improve the model construction efficiency, the classification performance of the classification model to be processed can be tested only after multiple rounds of updates of the classification model to be processed.

[0160] Based on this, the embodiments of the present application also provide another possible implementation manner of S222, which may specifically include Step 11 - Step 13:

[0161] Step 11: Use the classification model to be processed to determine the model classification results of each training sample.

[0162] It should be noted that for the relevant content of Step 11, please refer to S2221 above.

[0163] Step 12: Update the classification model to be processed according to the model classification results of at least one training sample and the label information of the at least one training sample.

[0164] It should be noted that for the relevant content of Step 12, please refer to S2222 above.

[0165] Step 13: Determine whether the second preset condition is met. If so, execute S223 below; if not, return and continue to execute Step 11 and its subsequent steps above.

[0166] Among them, the second preset condition can be set in advance. For example, specifically, it can be: during the current training process of the to-be-processed classification model, the update times of the to-be-processed classification model reach the first number threshold. It should be noted that the first number threshold can be set in advance. For example, the first number threshold = 6.

[0167] It can be seen that for the current training process of the to-be-processed classification model, after completing the update process of the to-be-processed classification model, it can be determined whether the update times of the to-be-processed classification model reach the first number threshold. If the first number threshold is reached, it means that all update processes for the to-be-processed classification model have been completed during the current training process. Therefore, S223 below can be directly executed to implement a test process for the to-be-processed classification model. If the first number threshold is not reached, it means that all update processes for the to-be-processed classification model have not been completed during the current training process. Therefore, Step 11 and its subsequent steps above can be continued to implement the next round of update process for the to-be-processed classification model.

[0168] It should be noted that in order to avoid interference from the current training process to the next training process, after determining that the second preset condition is met, the update times of the to-be-processed classification model can be reset to 0, so that during the next training process of the to-be-processed classification model, the update times of the to-be-processed classification model can be re-counted.

[0169] Based on the relevant content of the above Step 11 to Step 13, during the current training process of the to-be-processed classification model, these training samples and their label information can be used to implement multiple rounds of update processes for the to-be-processed classification model, so as to be able to perform classification performance tests on the to-be-processed classification model after multiple rounds of updates. In this way, a model training mode of multiple updates and one test can be realized, which can effectively reduce the occurrence times of the model test process, thereby being beneficial to improving the model training efficiency of the to-be-processed classification model.

[0170] Based on the relevant content of the above S222, after obtaining at least one training sample and its label information (or, after determining that the first preset condition is not met), these training samples and their label information can be used to complete at least one round of update process for the to-be-processed classification model, so that the updated to-be-processed classification model has better classification performance. It should be noted that the embodiments of the present application do not limit the learning rate used in the above model update process. For example, specifically, it can be 8e -5 .

[0171] S223: Use at least one test sample and the label information of the at least one test sample to test the classification model to be processed, and obtain a model test result.

[0172] Among them, the model test result is used to represent the classification performance of the classification model to be processed; moreover, the embodiments of the present application do not limit the model test result. For example, it may include: the prediction accuracy rate of the classification model to be processed for at least one test sample.

[0173] In addition, the embodiments of the present application do not limit the determination process of the model test result. For example, it may specifically include S2231 - S2233:

[0174] S2231: Use the classification model to be processed to determine the model classification result of each test sample.

[0175] Among them, the model classification result of the y-th test sample is used to represent the predicted state of the y-th test sample under the target disease (for example, whether having cancer or the cancer stage). y is a positive integer, y ≤ Y, Y is a positive integer, and Y represents the number of test samples.

[0176] In addition, the embodiments of the present application do not limit the acquisition process of the above "model classification result of the y-th test sample". For example, specifically, it may be: input the y-th test sample into the classification model to be processed, and obtain the model classification result of the y-th test sample output by the classification model to be processed.

[0177] S2232: Determine the model prediction accuracy rate according to the model classification results of at least one test sample and the label information of the at least one test sample.

[0178] Among them, the model prediction accuracy rate is used to represent the prediction accuracy rate of the classification model to be processed for all test samples.

[0179] In addition, the embodiments of the present application do not limit the determination process of the model prediction accuracy rate. For example, any existing or future accuracy rate determination method can be used for implementation.

[0180] S2233: Determine the model prediction accuracy rate as the model test result.

[0181] In the embodiments of the present application, after obtaining the model prediction accuracy rate, the model prediction accuracy rate can be determined as the model test result, so that the model test result can represent the prediction accuracy rate of the classification model to be processed for all test samples, so that the model test result can represent the prediction performance of the classification model to be processed for all test samples, and further so that the model test result can represent the classification performance of the classification model to be processed.

[0182] Based on the relevant content of S223 above, after obtaining the to-be-processed classification model that has been trained, at least one test sample and its label information can be used to test the to-be-processed classification model, and a model test result can be obtained, so that the model test result can represent the classification performance of the to-be-processed classification model.

[0183] S224: Determine whether the first preset condition is met. If so, execute S225 below; if not, return and continue to execute S222 above and its subsequent steps.

[0184] Among them, the first preset condition can be set in advance. For example, the first preset condition can specifically include: the model test result reaches the third preset condition; or, the number of model tests reaches the second number threshold (for example, 100).

[0185] It can be seen that if it is determined that the model test result reaches the third preset condition, it can be determined that the first preset condition has been met, and thus it can be determined that the to-be-processed classification model has good classification performance. Therefore, based on the to-be-processed classification model, S225 below can be executed; if it is determined that the number of model tests reaches the second number threshold, it can be determined that the first preset condition has been met, and thus it can be determined that the to-be-processed classification model has good classification performance. Therefore, based on the to-be-processed classification model, S225 below can be executed; however, if it is determined that the model test result does not reach the third preset condition and the number of model tests does not reach the second number threshold, it can be determined that the first preset condition is not met, and thus it can be determined that the classification performance of the to-be-processed classification model is still relatively poor. Therefore, S222 above and its subsequent steps can be continued to implement the next update test process for the to-be-processed classification model.

[0186] It should be noted that the above "third preset condition" can be set in advance. For example, when the model test result is the above "model prediction accuracy rate", the third preset condition can specifically be: reaching 90%. That is, if the above model prediction accuracy rate reaches 90%, it can be determined that the model test result reaches the third preset condition; if the above model prediction accuracy rate is lower than 90%, it can be determined that the model test result does not reach the third preset condition.

[0187] Based on the relevant content of S224 above, after completing a test process for the classification model to be processed, it can be determined whether the classification model to be processed meets the first preset condition. If it meets the first preset condition (for example, the model prediction accuracy of the classification model to be processed reaches 90%, or the number of model tests for the classification model to be processed reaches 100 times), it can be determined that the classification model to be processed has good classification performance. Therefore, based on this classification model to be processed, S225 below can be executed. However, if it does not meet the first preset condition (for example, the model prediction accuracy of the classification model to be processed is lower than 90%, and the number of model tests for the classification model to be processed is lower than 100 times), it can be determined that the classification performance of the classification model to be processed is poor. Therefore, S222 above and its subsequent steps can be continued to implement the next round of update and test process for the classification model to be processed.

[0188] S225: Determine the classification model to be used according to the classification model to be processed.

[0189] In the embodiments of the present application, after determining that the classification model to be processed meets the first preset condition, it can be determined that the classification model to be processed has good classification performance. Therefore, according to this classification model to be processed, the classification model to be used can be determined (for example, the classification model to be processed can be directly determined as the classification model to be used; or, the model parameters of the classification model to be used can be configured using the model parameters of the classification model to be processed so that the model parameters of the classification model to be used are consistent with the model parameters of the classification model to be processed), so that the classification model to be used also has good classification performance, so that the biological information related to the target disease carried by the above "at least one gene expression data" can be interpreted using the classification model to be used subsequently.

[0190] Based on the relevant content of S21 to S22 above, after obtaining at least one signaling pathway, at least one gene expression data, and its annotation information under the target disease, these signaling pathways can be used to perform data mapping processing on these gene expression data to obtain gene expression samples corresponding to these gene expression data, so that these gene expression samples conform to the gene attention characteristics of these signaling pathways; then, these gene expression samples and their corresponding annotation information under the target disease are used to construct the classification model to be used, so that the classification model to be used has good classification performance, so that the biological information related to the target disease carried by these gene expression data can be interpreted using the classification model to be used subsequently.

[0191] In fact, a classification model to be used constructed using a large amount of gene expression data can better explain the biological information related to the target disease carried by these gene expression data. Based on this, an embodiment of the present application also provides a possible implementation manner of the above S3, which may specifically include S31 - S32:

[0192] S31: Determine contribution parameters of at least one signaling pathway from the classification model to be used.

[0193] Among them, the contribution parameter of the m-th signaling pathway is used to represent the influence of the pathway member genes of the m-th signaling pathway on the target disease. m is a positive integer, m ≤ M, M is a positive integer, and M represents the number of signaling pathways.

[0194] In addition, the above "contribution parameter of the m-th signaling pathway" may include the weighted parameter of the m-th signaling pathway (that is, the above ), and / or the bias parameter of the m-th signaling pathway (that is, the above ).

[0195] In addition, the embodiment of the present application does not limit the implementation manner of S31. For example, when the classification model to be used includes a pathway layer, S31 may specifically be: Determine contribution parameters of at least one signaling pathway from the layer parameters of the pathway layer. For ease of understanding, the following is described with examples.

[0196] As an example, when the m-th node in the above pathway layer represents the m-th signaling pathway, the layer parameters of the m-th node and can be determined as the contribution parameters of the m-th signaling pathway. Among them, m is a positive integer, m ≤ M, and M is a positive integer.

[0197] Based on the relevant content of the above S31, after constructing the classification model to be used, the contribution parameters of each signaling pathway can be determined using the layer parameters of the pathway layer in the classification model to be used. Among them, since the pathway layer is the first network layer in the classification model to be used, the pathway layer belongs to the shallow network of the classification model to be used, so that the layer parameters of the pathway layer are relatively easy to be understood by humans, and then humans can explain the reason for the correct prediction of the model based on the layer parameters of the pathway layer. In this way, it can effectively avoid the defect that when extracting data features from the deep network, humans cannot directly understand these data features, resulting in humans being unable to explain the reason for the correct prediction.

[0198] S32: Use the contribution parameters of at least one signaling pathway to determine the information analysis result of at least one gene expression data under the target disease.

[0199] As an example, when the above-mentioned "contribution parameter" includes a weighting parameter and a bias parameter, S32 may specifically include S321 - S323:

[0200] S321: Determine the contribution analysis result of each signaling pathway to the target disease according to the contribution parameters of each signaling pathway, the gene expression samples corresponding to at least one gene expression data, and the annotation information of the at least one gene expression data under the target disease.

[0201] Among them, the contribution analysis result of the m-th signaling pathway to the target disease is used to represent the impact of the m-th signaling pathway on the target disease. m is a positive integer, m ≤ M, M is a positive integer, and M represents the number of signaling pathways.

[0202] In addition, the determination process of the above-mentioned "contribution analysis result of the m-th signaling pathway to the target disease" may specifically include S3211 - S3212:

[0203] S3211: Determine the pathway contribution value of the n-th gene expression data under the m-th signaling pathway according to the contribution parameter of the m-th signaling pathway and the gene expression sample corresponding to the n-th gene expression data.

[0204] Among them, the above-mentioned "pathway contribution value of the n-th gene expression data under the m-th signaling pathway" is used to represent the degree of impact of the m-th signaling pathway on the target disease presented in the n-th gene expression data.

[0205] In addition, the above-mentioned "pathway contribution value of the n-th gene expression data under the m-th signaling pathway" can be calculated using the above formula (1).

[0206] S3212: Determine the contribution analysis result of the m-th signaling pathway to the target disease according to the pathway contribution values of the N gene expression data under the m-th signaling pathway and the annotation information of the N gene expression data under the target disease.

[0207] In fact, the degree of impact of each signaling pathway on different states of the target disease (for example, having cancer, not having cancer; or having cancer stage 1, having cancer stage 2, etc.) is different. Therefore, in order to more accurately analyze the impact of each signaling pathway on the target disease, the determination process of the above-mentioned "contribution analysis result of the m-th signaling pathway to the target disease" may specifically include Step 21 - Step 23:

[0208] Step 21: Determine at least one i-th type of data sample from the N gene expression data according to the annotation information of the N gene expression data under the target disease. Where i is a positive integer, i ≤ I, I is a positive integer, and I represents the number of types of data samples.

[0209] Among them, the i-th type of data sample is used to represent gene expression data with the i-th type of annotation information, so that the above-mentioned "at least one i-th type of data sample" is used to represent all gene expression data under the i-th type of annotation information.

[0210] In addition, the embodiments of the present application do not limit the above-mentioned I types of annotation information. For example, when the "annotation information of gene expression data under a target disease" is used to represent whether an object with the n-th gene expression data has cancer, the I types of annotation information may include one type of annotation information for representing "having cancer" (for example, 1), and one type of annotation information for representing "not having cancer" (for example, 0). Another example is that when the "annotation information of gene expression data under a target disease" is used to represent the cancer stage of an object with the n-th gene expression data, the I types of annotation information may include one type of annotation information for representing "cancer stage 1" (for example, 1), one type of annotation information for representing "cancer stage 2" (for example, 2), one type of annotation information for representing "cancer stage 3" (for example, 3), and one type of annotation information for representing "cancer stage 0 (that is, not having cancer)" (for example, 0).

[0211] In addition, the embodiments of the present application do not limit the determination process of the above-mentioned "i-th type of data sample". For example, specifically, if the annotation information of the n-th gene expression data under a target disease belongs to the i-th type of annotation information, then the n-th gene expression data can be determined as the i-th type of data sample. Wherein, n is a positive integer, and n ≤ N.

[0212] Based on the relevant content of step 21 above, after obtaining the annotation information of at least one gene expression data under a target disease, these gene expression data can be divided into multiple sample sets according to these annotation information, so that each sample set is used to record gene expression data under various annotation information.

[0213] It should be noted that the execution time of step 21 is earlier than that of step 22, and the embodiments of the present application do not limit the execution time of step 21.

[0214] Step 22: Perform a preset statistical process on the pathway contribution values of at least one i-th type of data sample under the m-th signaling pathway to obtain the i-th statistical contribution value of the m-th signaling pathway. Wherein, i is a positive integer, and i ≤ I.

[0215] Among them, the preset statistical process can be set in advance; and the embodiments of the present application do not limit the preset statistical process. For example, it can be taking the average value, taking the maximum value, taking the minimum value, or taking the median value.

[0216] The above-mentioned "the i-th statistical contribution value of the m-th signal pathway" is used to represent the degree of influence of the m-th signal pathway on the i-th state under the target disease (that is, the state represented by the above-mentioned "the i-th annotation information").

[0217] In addition, the embodiments of the present application do not limit the determination process of the above-mentioned "the i-th statistical contribution value of the m-th signal pathway". For example, the pathway contribution values of all the i-th data samples under the m-th signal pathway can be averaged to obtain the i-th statistical contribution value of the m-th signal pathway.

[0218] Based on the relevant content of step 22 above, after obtaining at least one i-th data sample, preset statistical processing (such as taking the average value, etc.) can be performed on the pathway contribution values of these i-th data samples under the m-th signal pathway to obtain the i-th statistical contribution value of the m-th signal pathway, so that the "the i-th statistical contribution value of the m-th signal pathway" can represent the degree of influence of the m-th signal pathway on the i-th state under the target disease.

[0219] Step 23: Perform preset analysis processing on the 1st statistical contribution value to the I-th statistical contribution value of the m-th signal pathway to obtain the contribution analysis result of the m-th signal pathway to the target disease.

[0220] Among them, the preset analysis processing can be preset in advance; moreover, the embodiments of the present application do not limit the preset analysis processing. For example, it can be implemented by using a bar chart comparison method (as Figure 6 shown). It should be noted that for Figure 6 it, a positive value indicates a promoting effect on a certain state of the target disease, and a negative value indicates an inhibitory effect on a certain state of the target disease.

[0221] In addition, the embodiments of the present application do not limit the determination process of the above-mentioned "contribution analysis result of the m-th signal pathway to the target disease". For example, specifically, it can be: comparing the 1st statistical contribution value of the m-th signal pathway, the 2nd statistical contribution value of the m-th signal pathway,..., and the I-th statistical contribution value of the m-th signal pathway to obtain the contribution analysis result of the m-th signal pathway to the target disease, so that the "contribution analysis result of the m-th signal pathway to the target disease" is used to record the different influences of the m-th signal pathway on different states under the target disease (for example, the XXX signal pathway has a greater influence on the first stage of cancer, so that the XXX signal pathway makes the possibility of the first stage of cancer greater).

[0222] Based on the relevant content of the above steps 21 to 23, after obtaining the pathway contribution values of N gene expression data under the m-th signaling pathway, the gene expression data can be divided into different categories by referring to the annotation information of these gene expression data under the target disease first; then, using the pathway contribution values of all gene expression data in each category under the m-th signaling pathway, determine the contribution value of the m-th signaling pathway under the disease state represented by each category; finally, compare the contribution values of the m-th signaling pathway under the disease states represented by all categories to determine the contribution analysis result of the m-th signaling pathway to the target disease, so that the "contribution analysis result of the m-th signaling pathway to the target disease" can more accurately represent the impact of the m-th signaling pathway on the target disease.

[0223] Based on the relevant content of the above S321, after obtaining the contribution parameters of the m-th signaling pathway, the contribution analysis result of the m-th signaling pathway to the target disease can be determined by using the contribution parameters of the m-th signaling pathway from the gene expression samples corresponding to at least one gene expression data and their annotation information under the target disease, so that the "contribution analysis result of the m-th signaling pathway to the target disease" can represent the impact of the m-th signaling pathway on the target disease.

[0224] S322: Determine the contribution analysis result of the at least one gene to the target disease according to the weighted parameters of at least one signaling pathway and at least one gene expression data.

[0225] Among them, the contribution analysis result of the k-th gene to the target disease is used to represent the impact of the k-th gene on the target disease. k is a positive integer, k ≤ K, K is a positive integer, and K represents the number of genes involved in a gene expression data.

[0226] In addition, the determination process of the above "contribution analysis result of the k-th gene to the target disease" can specifically include steps 31 - 32:

[0227] Step 31: Determine at least one gene usage pathway corresponding to each gene from at least one signaling pathway.

[0228] Among them, the gene usage pathway corresponding to the k-th gene refers to the signaling pathway involving the k-th gene.

[0229] In addition, the determination process of the above "gene usage pathway corresponding to the k-th gene" can specifically be: for the m-th signaling pathway, if the pathway member genes of the m-th signaling pathway include the k-th gene (that is, the m-th signaling pathway involves the k-th gene), then the m-th signaling pathway can be determined as the gene usage pathway corresponding to the k-th gene. Among them, m is a positive integer, m ≤ M, M is a positive integer, and M represents the number of signaling pathways.

[0230] It should be noted that the execution time of step 31 is earlier than that of step 32, and the embodiments of the present application do not limit the execution time of step 31.

[0231] Step 32: Determine the contribution analysis result of each gene to the target disease according to the weighted parameters of at least one gene usage pathway corresponding to each gene and the gene characterization value of each gene in at least one gene expression data.

[0232] As an example, step 32 may specifically include steps 321 - 323:

[0233] Step 321: Determine the gene contribution value of the nth gene expression data under the kth gene according to the weighted parameters of at least one gene usage pathway corresponding to the kth gene and the gene characterization value of the kth gene in the nth gene expression data. Wherein, n is a positive integer, and n ≤ N.

[0234] The above "gene contribution value of the nth gene expression data under the kth gene" is used to represent the degree of influence of the kth gene on the target disease presented in the nth gene expression data.

[0235] In addition, the embodiments of the present application do not limit the determination process of the above "gene contribution value of the nth gene expression data under the kth gene". For example, it can be implemented using formula (8).

[0236]

[0237] In the formula, represents the gene contribution value of the nth gene expression data under the kth gene; represents the gene characterization value of the kth gene in the nth gene expression data, and E n represents the nth gene expression data, and represents the weighted parameter of the lth gene usage pathway corresponding to the kth gene; represents the weighted weight value of the kth gene recorded in L k represents the number of gene usage pathways corresponding to the kth gene.

[0238] Step 322: Average the gene contribution values of the N gene expression data under the kth gene to obtain the contribution characterization value of the kth gene to the target disease.

[0239] The above-mentioned "contribution characterization value of the k-th gene under the target disease" is used to represent the degree of influence of the k-th gene on the target disease; moreover, the "contribution characterization value of the k-th gene under the target disease" can be determined by the following formula (9).

[0240]

[0241] In the formula, G k represents the contribution characterization value of the k-th gene under the target disease; represents the gene contribution value of the n-th gene expression data under the k-th gene; represents the gene characterization value of the k-th gene in the n-th gene expression data, and E n represents the n-th gene expression data, and represents the weighted parameter of the l-th gene usage pathway corresponding to the k-th gene; represents the weighted weight value of the k-th gene recorded in, and L k represents the number of gene usage pathways corresponding to the k-th gene.

[0242] Step 323: Determine the contribution analysis result of the k-th gene to the target disease according to the contribution characterization value of the k-th gene under the target disease.

[0243] Actually, for the contribution characterization value of the k-th gene under the target disease, if the "contribution characterization value of the k-th gene under the target disease" is larger, it means that the k-th gene has a greater degree of influence on the target disease; if the "contribution characterization value of the k-th gene under the target disease" is smaller, it means that the k-th gene has a smaller degree of influence on the target disease. It can be seen that there is a positive correlation between the above-mentioned "contribution characterization value of the k-th gene under the target disease" and the degree of influence of the k-th gene on the target disease.

[0244] In addition, the embodiment of the present application does not limit the determination process of the above-mentioned "contribution analysis result of the k-th gene to the target disease". For example, specifically, it can be: after determining that the contribution characterization value of the k-th gene under the target disease reaches a preset contribution threshold, it can be determined that the k-th gene has a relatively large degree of influence on the target disease, so the k-th gene can be directly determined as the main influencing gene of the target disease. Among them, the above-mentioned "main influencing gene of the target disease" is used to represent the gene that causes the target disease.

[0245] Based on the relevant content of the above step 32, after obtaining at least one gene usage pathway corresponding to the k-th gene, the weighted parameters of these gene usage pathways can be used to perform information analysis on the gene characterization value of the k-th gene in all gene expression data, so as to obtain the contribution analysis result of the k-th gene to the target disease, so that the "contribution analysis result of the k-th gene to the target disease" can represent the impact of the k-th gene on the target disease (for example, whether the k-th gene belongs to the main influencing genes of the target disease).

[0246] Based on the relevant content of the above S322, after obtaining the weighted parameters of at least one signaling pathway, the weighted parameters of these signaling pathways can be used to perform gene analysis on all gene expression data, so as to obtain the contribution analysis results of at least one gene to the target disease, so that these contribution analysis results can represent the impact of these genes on the target disease (for example, which genes mainly cause the target disease).

[0247] S323: Determine the information analysis result of at least one gene expression data under the target disease according to the contribution analysis result of at least one signaling pathway to the target disease and the contribution analysis result of at least one gene to the target disease.

[0248] In the embodiment of the present application, after obtaining the contribution analysis result of at least one signaling pathway to the target disease and the contribution analysis result of at least one gene to the target disease, the contribution analysis results of these signaling pathways to the target disease and the contribution analysis results of these genes to the target disease can be summarized to obtain the above "information analysis result of at least one gene expression data under the target disease", so that the "information analysis result of at least one gene expression data under the target disease" can represent which signaling pathways and which genes cause the target disease, so that the "information analysis result of at least one gene expression data under the target disease" can better represent the biological information related to the target disease.

[0249] Based on the relevant content of S31 to S32 above, after obtaining the classification model to be used, the contribution parameters of each signal pathway can be extracted from the shallow network of the classification model to be used, so that these contribution parameters have high interpretability; then, using the contribution parameters of these signal pathways, interpretive analysis of all gene expression data under the target disease is performed to obtain the information analysis result of these gene expression data under the target disease, so that the information analysis result can indicate which signal pathways and which genes lead to the target disease, thereby enabling the information analysis result to better represent the biological information related to the target disease. Among them, since the contribution parameters of each signal pathway are data features extracted from the shallow network of the classification model to be used, these contribution parameters can be directly understood by humans, so that the information analysis result obtained by interpreting based on these contribution parameters is more easily accepted by humans, which is beneficial to improving the research effect on these gene expression data, and thus beneficial to making more full use of the biological information carried by these gene expression data.

[0250] Based on the relevant content of the above gene expression data processing method, the embodiment of the present application also provides a gene expression data processing device, which will be described below with reference to the drawings.

[0251] See Figure 7 , which is a schematic structural diagram of a gene expression data processing device provided by an embodiment of the present application.

[0252] The gene expression data processing device 700 provided by the embodiment of the present application includes:

[0253] An information acquisition unit 701, configured to acquire at least one gene expression data, the annotation information of the at least one gene expression data under the target disease, and at least one signal pathway;

[0254] A model construction unit 702, configured to construct a classification model to be used by using the at least one gene expression data, the annotation information of the at least one gene expression data under the target disease, and the at least one signal pathway;

[0255] An information analysis unit 703, configured to determine the information analysis result of the at least one gene expression data under the target disease according to the classification model to be used and the at least one signal pathway.

[0256] In a possible implementation manner, the information analysis unit 703 includes:

[0257] A parameter extraction subunit, configured to determine the contribution parameters of the at least one signal pathway from the classification model to be used;

[0258] An information analysis subunit, configured to use the contribution parameters of the at least one signal pathway to determine the information analysis result of the at least one gene expression data under the target disease.

[0259] In a possible implementation manner, the classification model to be used includes a pathway layer; the pathway layer is configured to determine the pathway contribution value of the gene expression data under each signal pathway.

[0260] The parameter extraction subunit is specifically configured to: determine the contribution parameters of the at least one signal pathway from the layer parameters of the pathway layer.

[0261] In a possible implementation manner, the pathway layer belongs to the shallow network of the classification model to be used.

[0262] In a possible implementation manner, the contribution parameters include a weighting parameter and a bias parameter.

[0263] The information analysis subunit includes:

[0264] A pathway analysis subunit, configured to determine the contribution analysis result of each signal pathway to the target disease according to the contribution parameters of each signal pathway, the gene expression sample corresponding to the at least one gene expression data, and the annotation information of the at least one gene expression data under the target disease.

[0265] A gene analysis subunit, configured to determine the contribution analysis result of at least one gene to the target disease according to the weighting parameter of the at least one signal pathway and the at least one gene expression data.

[0266] An analysis summary subunit, configured to determine the information analysis result of the at least one gene expression data under the target disease according to the contribution analysis result of the at least one signal pathway to the target disease and the contribution analysis result of the at least one gene to the target disease.

[0267] In a possible implementation manner, the number of signal pathways is M; the number of gene expression data is N; where M is a positive integer and N is a positive integer.

[0268] The pathway analysis subunit includes:

[0269] A first determination subunit, configured to determine the pathway contribution value of the nth gene expression data under the mth signal pathway according to the contribution parameter of the mth signal pathway and the gene expression sample corresponding to the nth gene expression data; where m is a positive integer and m ≤ M.

[0270] A second determination subunit, configured to determine a contribution analysis result of the m-th signaling pathway to the target disease according to a pathway contribution value of the N gene expression data under the m-th signaling pathway and annotation information of the N gene expression data for the target disease.

[0271] In a possible implementation manner, the method further includes:

[0272] A sample partitioning unit, configured to determine at least one i-th type of data sample from the N gene expression data according to the annotation information of the N gene expression data for the target disease; where i is a positive integer, i ≤ I, I is a positive integer, and I represents the number of types of the data samples;

[0273] The second determination subunit is specifically configured to: perform a preset statistical process on the pathway contribution values of the at least one i-th type of data sample under the m-th signaling pathway to obtain an i-th statistical contribution value of the m-th signaling pathway; where i is a positive integer, i ≤ I; perform a preset analysis process on the first to I-th statistical contribution values of the m-th signaling pathway to obtain a contribution analysis result of the m-th signaling pathway to the target disease.

[0274] In a possible implementation manner, the method further includes:

[0275] A pathway mapping unit, configured to determine at least one gene usage pathway corresponding to each of the genes from the at least one signaling pathway;

[0276] The gene analysis subunit is specifically configured to: determine a contribution analysis result of each gene to the target disease according to a weighting parameter of at least one gene usage pathway corresponding to each gene and a gene characterization value of each gene in the at least one gene expression data.

[0277] In a possible implementation manner, the gene expression data includes gene characterization values of K genes; the number of the gene expression data is N; where K is a positive integer and N is a positive integer;

[0278] The gene analysis subunit is specifically configured to: determine a gene contribution value of the n-th gene expression data under the k-th gene according to a weighting parameter of at least one gene usage pathway corresponding to the k-th gene and a gene characterization value of the k-th gene in the n-th gene expression data; where k is a positive integer, k ≤ K, n is a positive integer, n ≤ N; perform an averaging process on the gene contribution values of the N gene expression data under the k-th gene to obtain a contribution characterization value of the k-th gene for the target disease; determine a contribution analysis result of the k-th gene to the target disease according to the contribution characterization value of the k-th gene for the target disease.

[0279] In a possible implementation manner, the model construction unit 702 includes:

[0280] A data mapping subunit, configured to perform data mapping processing on each of the gene expression data according to each of the signal pathways, so as to obtain gene expression samples corresponding to each of the gene expression data;

[0281] A model construction subunit, configured to construct a classification model to be used by using the gene expression samples corresponding to the at least one gene expression data and the annotation information of the at least one gene expression data under a target disease.

[0282] In a possible implementation manner, the model construction subunit includes:

[0283] A data determination subunit, configured to determine at least one training sample, label information of the at least one training sample, at least one test sample, and label information of the at least one test sample according to the gene expression samples corresponding to the at least one gene expression data and the annotation information of the at least one gene expression data under a target disease;

[0284] A model update subunit, configured to update a classification model to be processed by using the at least one training sample and the label information of the at least one training sample;

[0285] A model test subunit, configured to test the classification model to be processed by using the at least one test sample and the label information of the at least one test sample, obtain a model test result, and return to the model update subunit to continue to execute the step of updating the classification model to be processed by using the at least one training sample and the label information of the at least one training sample;

[0286] A model determination subunit, configured to determine a classification model to be used according to the classification model to be processed when it is determined that a first preset condition is satisfied.

[0287] In a possible implementation manner, the model update subunit is specifically configured to:

[0288] Use the classification model to be processed to determine the model classification results of each training sample;

[0289] Update the classification model to be processed according to the model classification results of the at least one training sample and the label information of the at least one training sample, and continue to execute the step of using the classification model to be processed to determine the model classification results of each training sample;

[0290] The model testing subunit is specifically configured to: when it is determined that the second preset condition is met, use the at least one test sample and the label information of the at least one test sample to test the classification model to be processed, and obtain a model test result.

[0291] In a possible implementation manner, the number of signal pathways is M; the gene expression sample includes a set of characterization values corresponding to M signal pathways; the set of characterization values corresponding to the m-th signal pathway includes gene characterization values of at least one gene in the gene expression data; where m is a positive integer, m ≤ M, and M is a positive integer.

[0292] Based on the relevant content of the above gene expression data processing device 700, for the gene expression data processing device 700, after obtaining a large number of signal pathways, a large amount of gene expression data and their annotation information under the target disease, it is possible to first use these signal pathways, this gene expression data and their annotation information under the target disease to construct a classification model to be used, so that the classification model to be used has better classification performance under the target disease, so that the classification model to be used carries some biological information required for understanding and analyzing this gene expression data under the target disease; then, based on the classification model to be used and these signal pathways, perform information analysis and processing on this gene expression data to obtain an information analysis result of this gene expression data under the target disease, so that the information analysis result can accurately represent biological information related to the target disease (for example, which channels and / or which genes are more likely to cause cancer), so as to be able to mine biological information related to the target disease (for example, influencing factors of the target disease) from a large amount of gene expression data, and thus be able to achieve biological information analysis for a large amount of gene expression data.

[0293] In addition, since the construction process of the classification model to be used not only refers to a large amount of gene expression data and their annotation information under the target disease, but also refers to a large number of signal pathways, the classification model to be used has better prediction performance, so that the classification model to be used has better classification performance under the target disease, and further enables the classification model to be used to better represent the biological information required for understanding and analyzing this gene expression data under the target disease, so that the information analysis result obtained based on the analysis of the classification model to be used is more accurate, which is beneficial to improving the accuracy of the information analysis result.

[0294] In addition, an embodiment of the present application further provides a gene expression data processing device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, any implementation manner of the gene expression data processing method provided by the embodiment of the present application is implemented.

[0295] In addition, an embodiment of the present application further provides a computer-readable storage medium, in which instructions are stored. When the instructions run on a terminal device, the terminal device is caused to execute any implementation manner of the gene expression data processing method provided by the embodiment of the present application.

[0296] In addition, an embodiment of the present application further provides a computer program product. When the computer program product runs on a terminal device, the terminal device is caused to execute any implementation manner of the gene expression data processing method provided by the embodiment of the present application.

[0297] It should be noted that the various embodiments in this specification are described in a progressive manner. The focus of each embodiment is on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description of the method part for the relevant parts.

[0298] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0299] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0300] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in a software module executed by a processor, or in a combination thereof. The software module may be disposed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0301] The foregoing description of the disclosed embodiments enables those skilled in the art to make or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for processing gene expression data, characterized in that The method includes: obtaining at least one gene expression data, annotation information of the at least one gene expression data under a target disease, and at least one signaling pathway; using the at least one gene expression data, the annotation information of the at least one gene expression data under the target disease, and the at least one signaling pathway to construct a classification model to be used, where the classification model to be used includes a pathway layer for determining a pathway contribution value of the gene expression data under each signaling pathway; determining an information analysis result of the at least one gene expression data under the target disease according to the classification model to be used and the at least one signaling pathway; The process of determining the information analysis result includes: determining contribution parameters of the at least one signaling pathway from layer parameters of the pathway layer, where the contribution parameters include a weighting parameter and a bias parameter; determining a contribution analysis result of each signaling pathway to the target disease according to the contribution parameters of each signaling pathway, gene expression samples corresponding to the at least one gene expression data, and the annotation information of the at least one gene expression data under the target disease; determining a contribution analysis result of at least one gene to the target disease according to the weighting parameters of the at least one signaling pathway and the at least one gene expression data; and determining the information analysis result of the at least one gene expression data under the target disease according to the contribution analysis result of the at least one signaling pathway to the target disease and the contribution analysis result of the at least one gene to the target disease; The number of the signaling pathways is M; the number of the gene expression data is N; where M is a positive integer and N is a positive integer; the process of determining the contribution analysis result of the m-th signaling pathway to the target disease includes: determining a pathway contribution value of the n-th gene expression data under the m-th signaling pathway according to the contribution parameters of the m-th signaling pathway and the gene expression sample corresponding to the n-th gene expression data; where m is a positive integer and m ≤ M; and determining the contribution analysis result of the m-th signaling pathway to the target disease according to the pathway contribution values of the N gene expression data under the m-th signaling pathway and the annotation information of the N gene expression data under the target disease.

2. The method according to claim 1, wherein The pathway layer belongs to a shallow network of the classification model to be used.

3. The method according to claim 1, wherein The method further includes: determining at least one i-th data sample from the N gene expression data according to the annotation information of the N gene expression data under the target disease; where i is a positive integer, i ≤ I, I is a positive integer, and I represents the number of types of the data samples; The step of determining the contribution analysis result of the m-th signaling pathway to the target disease according to the pathway contribution values of the N gene expression data under the m-th signaling pathway and the annotation information of the N gene expression data under the target disease includes: Perform a preset statistical process on the pathway contribution value of the at least one type-i data sample under the m-th signal pathway to obtain the i-th statistical contribution value of the m-th signal pathway; where i is a positive integer and i ≤ I; Perform a preset analysis process on the 1st to I-th statistical contribution values of the m-th signal pathway to obtain the contribution analysis result of the m-th signal pathway to the target disease.

4. The method according to claim 1, characterized in that The method further includes: Determine at least one gene usage pathway corresponding to each of the genes from the at least one signal pathway; The determining of the contribution analysis result of at least one gene to the target disease according to the weighted parameters of the at least one signal pathway and the at least one gene expression data includes: Determine the contribution analysis result of each gene to the target disease according to the weighted parameters of at least one gene usage pathway corresponding to each gene and the gene characterization value of each gene in the at least one gene expression data.

5. The method according to claim 4, wherein The gene expression data includes the gene characterization values of K genes; where K is a positive integer; The process of determining the contribution analysis result of the k-th gene to the target disease includes: Determine the gene contribution value of the n-th gene expression data under the k-th gene according to the weighted parameter of at least one gene usage pathway corresponding to the k-th gene and the gene characterization value of the k-th gene in the n-th gene expression data; where k is a positive integer and k ≤ K; Perform an averaging process on the gene contribution values of N gene expression data under the k-th gene to obtain the contribution characterization value of the k-th gene under the target disease; Determine the contribution analysis result of the k-th gene to the target disease according to the contribution characterization value of the k-th gene under the target disease.

6. The method according to any one of claims 1-5, characterized in that, The constructing of the classification model to be used by using the at least one gene expression data, the annotation information of the at least one gene expression data under the target disease, and the at least one signal pathway includes: Perform data mapping processing on each gene expression data according to each signal pathway to obtain gene expression samples corresponding to each gene expression data; Construct a classification model to be used by using the gene expression samples corresponding to the at least one gene expression data and the annotation information of the at least one gene expression data under the target disease.

7. The method according to claim 6, wherein The constructing of the classification model to be used by using the gene expression samples corresponding to the at least one gene expression data and the annotation information of the at least one gene expression data under the target disease includes: Determine at least one training sample, the label information of the at least one training sample, at least one test sample, and the label information of the at least one test sample according to the gene expression samples corresponding to the at least one gene expression data and the annotation information of the at least one gene expression data under the target disease; Update the classification model to be processed by using the at least one training sample and the label information of the at least one training sample; Test the classification model to be processed by using the at least one test sample and the label information of the at least one test sample, obtain a model test result, and continue to execute the step of updating the classification model to be processed by using the at least one training sample and the label information of the at least one training sample until, when it is determined that a first preset condition is satisfied, determine a classification model to be used according to the classification model to be processed.

8. The method according to claim 7, characterized in that, The step of updating the classification model to be processed by using the at least one training sample and the label information of the at least one training sample includes: Use the classification model to be processed to determine the model classification results of each training sample; Update the classification model to be processed according to the model classification results of the at least one training sample and the label information of the at least one training sample, and continue to execute the step of using the classification model to be processed to determine the model classification results of each training sample; The step of testing the classification model to be processed by using the at least one test sample and the label information of the at least one test sample to obtain a model test result includes: When it is determined that a second preset condition is reached, test the classification model to be processed by using the at least one test sample and the label information of the at least one test sample to obtain a model test result.

9. The method according to claim 6, characterized in that, The gene expression sample includes a set of characterization values corresponding to M signal pathways; the set of characterization values corresponding to the m-th signal pathway includes the gene characterization values of at least one gene in the gene expression data.

10. A gene expression data processing device, characterized in that, The device includes: An information acquisition unit, configured to acquire at least one gene expression data, the annotation information of the at least one gene expression data under a target disease, and at least one signal pathway; A model construction unit, configured to construct a classification model to be used by using the at least one gene expression data, the annotation information of the at least one gene expression data under a target disease, and the at least one signal pathway, where the classification model to be used includes a pathway layer, and the pathway layer is configured to determine the pathway contribution value of the gene expression data under each signal pathway; An information analysis unit, configured to determine an information analysis result of the at least one gene expression data under the target disease according to the classification model to be used and the at least one signal pathway; The process of determining the information analysis result includes: determining the contribution parameters of the at least one signaling pathway from the layer parameters of the pathway layer, where the contribution parameters include a weighting parameter and a bias parameter; determining the contribution analysis result of each signaling pathway to the target disease according to the contribution parameters of each signaling pathway, the gene expression samples corresponding to the at least one gene expression data, and the annotation information of the at least one gene expression data under the target disease; determining the contribution analysis result of at least one gene to the target disease according to the weighting parameters of the at least one signaling pathway and the at least one gene expression data; and determining the information analysis result of the at least one gene expression data under the target disease according to the contribution analysis result of the at least one signaling pathway to the target disease and the contribution analysis result of the at least one gene to the target disease. The number of the signaling pathways is M; the number of the gene expression data is N; where M is a positive integer and N is a positive integer; the process of determining the contribution analysis result of the m-th signaling pathway to the target disease includes: determining the pathway contribution value of the n-th gene expression data under the m-th signaling pathway according to the contribution parameters of the m-th signaling pathway and the gene expression sample corresponding to the n-th gene expression data; where m is a positive integer and m ≤ M; and determining the contribution analysis result of the m-th signaling pathway to the target disease according to the pathway contribution values of the N gene expression data under the m-th signaling pathway and the annotation information of the N gene expression data under the target disease.

11. A gene expression data processing device, characterized in that including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the gene expression data processing method according to any one of claims 1-9 is implemented.

12. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when the instructions run on the terminal device, the terminal device is caused to execute the gene expression data processing method according to any one of claims 1-9.

13. A computer program product, characterized in that, When the computer program product runs on the terminal device, the terminal device is caused to execute the gene expression data processing method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Redundancy removal feature selection method LLRFC score+ based on LLRFC and correlation analysis

    CN105740653A

  • Biological state characterization method and device, equipment and storage medium

    CN112802546A