The invention discloses a
deep learning prediction method fused with multi-
modal features, which can accurately identify
protein hidden binding sites in a ligand-free (apoo) state. The method comprises the following steps of: firstly, constructing a
protein graph by taking residues as nodes and taking C alpha distance less than or equal to 14 as edges, wherein node feature sets comprise
amino acid one-hot, secondary structures, atomic attributes,
protein language model embedding and BLOSUM62 evolutionary information, and edge features comprise distance and angle similarity; then capturing three-dimensional geometric equivariant features by adopting an equivariant graph neural network (EGNN), and modeling a chemical topological relation by using a graph isomorphic network (GINE) with edge features; eGNN and GINE double-
branch feature fusion and global dependence integration are realized through gating cross attention and gating multi-head attention; and finally, inputting the fusion features into a Kolmogorov-Arnold network (KAN) classifier, and predicting whether each residue belongs to a hidden
binding site or not. The method can adapt to large-scale conformation change without coordinate alignment, AUC and F1 on a standard
data set are remarkably superior to those of an existing method, high robustness and generalization are kept for multi-chain protein and complex conformation, and the method can be widely applied to
drug target discovery and structure-driven
drug design.