The invention discloses an inherent disordered
protein region prediction method based on multi-
modal feature fusion. The method comprises the following steps: firstly, carrying out feature embedding and preprocessing on a
protein sequence by utilizing a pre-trained TAPE model, screening and preprocessing existing 544
amino acid physicochemical features in
biology, and extracting evolutionary conservative features from the
protein sequence; splicing the preprocessed protein embedding
feature matrix, the preprocessed
amino acid physicochemical
feature matrix and the preprocessed evolutionary conservative
feature matrix to obtain a multi-
modal fusion feature matrix; then, carrying out layer normalization on the multi-
modal fusion feature matrix, and carrying out
dimensionality reduction on protein embedding features in the multi-modal fusion feature matrix to obtain a multi-modal fusion feature matrix after local
dimensionality reduction; performing position coding on the multi-modal fusion feature matrix after local dimension reduction; and finally, inputting the position-coded multi-modal fusion feature matrix into the prediction model, and predicting the residues. According to the method, high
complementation and synergistic interaction of multi-modal features are realized, and the prediction precision is improved.