The invention belongs to the technical field of depression
risk identification methods, and particularly relates to a depression
risk identification method based on multi-model fusion and
gene feature analysis, depression-related
transcriptome expression data is obtained from a public
database, the
transcriptome expression data comprises
RNA expression matrixes of depression patients and normal controls, and the
RNA expression matrixes are used for identifying the depression risk. The
RNA expression matrix at least covers a
sample number, a
gene number, a
signal intensity value and a clinical grouping tag; 2, performing validity
verification on input
transcriptome expression data, identifying missing values, null columns and
abnormal expression quantity problems, prompting correction, outputting an expression matrix and a tag vector before
standardization, performing Z-
score standardization, missing value
processing and ComBat batch effect correction on the expression matrix output in the step 1 in sequence, and performing Z-
score standardization on the expression matrix output in the step 2; and sequentially performing
differential expression analysis, statistical test screening and
machine learning driven screening on the expression matrix corrected in the step 2, and obtaining an optimal feature
gene subset Gfinal through intersection extraction and recursive feature
elimination.