Tumor key gene identification method based on particle swarm optimization and marking criterion
A particle swarm optimization, key gene technology, applied in the fields of genomics, special data processing applications, instruments, etc., can solve the problems of lack of interpretability, the classification performance needs to be improved, and the subset of key tumor genes is large, so as to improve the tumor subgroup. The effect of type recognition
Patent Information
- Authority / Receiving Office
- CN · China
- Current Assignee / Owner
- Publication Date
- 2017-07-14
Smart Images

Figure 1 
Figure 2 
Figure 3
Abstract
Description
technical field
[0001] The invention belongs to the application field of computer analysis technology of tumor gene expression spectrum data, and in particular relates to a tumor key gene identification method based on particle swarm optimization and scoring criteria. Background technique
[0002] Statistical studies in recent years have shown that tumors have become one of the major diseases that endanger human health, and their prevalence is increasing year by year. Different subtypes of tumors have great differences in treatment methods. The primary key to whether the disease can be cured. However, studies have shown that there are usually a few to dozens of therapeutic genes for tumors, and the characteristics of high-dimensional and small samples of microarray data have become a huge challenge in screening disease-causing genes. Therefore, disease-causing genes are selected from tens of millions of genes The characteristic gene is the key problem to be solved. [0003...
Examples
Embodiment Construction
[0055] A key tumor gene identification method based on particle swarm optimization and scoring criteria, including optimizing Particle Swarm Optimization (PSO) through semi-initialization and Metropolis criteria, and using ELM extreme learning machine as an evaluation gene subset for correct classification The classifier of the rate, the step of obtaining the quantitative data of the classification performance of the algorithm comprises the following steps:
[0056] Step 1. Preprocessing of tumor gene expression profile data, including normalization and preliminary dimensionality reduction of tumor gene expression profile data sets, and simultaneously dividing tumor gene expression profile data sets into training sets and test sets;
[0057] Step 2 defines the scoring criteria and evaluates each gene in combination with the extreme learning machine, and screens out the top-scoring genes to establish a candidate gene pool;
[0058] Step 3 Combined with gene scoring information,...