The present invention discloses a method and device for recommending mutatable sites from small-sample experimental data, comprising the following steps: obtaining small-sample experimental data, including the sequences and data of wild enzymes and mutants, and obtaining the optimal
mutant sequence according to the
data type of the mutants; predicting the
mutant structure based on the optimal
mutant sequence; predicting the substrate-
binding pocket based on the mutant structure; predicting
single mutation sites based on the
protein structure model, and selecting residues with a distance from the center of the substrate-
binding pocket less than a first threshold as the recommended
single mutation sites; selecting sites from the small-sample experimental data, mutating each site in the site set into 19 other amino acids, and pairwise combining them to construct a
double mutation set, predicting the
mutation results, and obtaining the recommended
double mutation sites according to the sorting results; selecting sites from the small-sample experimental data, obtaining the coordinates of the sites in the variant structure for clustering analysis, and selecting 1 site from each cluster to combine with other clusters to construct multi-mutations; predicting the
mutation results of the multi-mutations, and obtaining the recommended multi-
mutation sites according to the sorting results; the method of the present invention operates effectively under the condition of small-sample data; through the powerful generalization ability of the large
language model, the structural and functional information in biomolecules can be captured, so as to effectively
encode the sequence of the
enzyme for more accurate
inference.