The invention is applicable to the technical field of
drug research and development, and provides a
lead compound discovery method based on a large
language model, which comprises the following steps: determining a target spot and constructing an initial compound
library; generating fragments and forming an initial fragment
library; calculating the
frequency of occurrence of the fragments and the affinity contribution degree of the fragments to the
target protein, and updating the fragment
library; fragment sampling weight calculation and candidate fragment selection; the large
language model generates molecules according to the candidate fragments; evaluating the quality of the generated molecules and updating the molecule library; and repeating the cycle until a preset iteration round number is reached. A fragment-based
drug design method is combined with the large
language model, rich
domain knowledge of the large language model is used for replacing human experts to guide
drug design, fragments are used as prompts to stimulate the internal potential of the large language model, the advantages of the two methods complement each other, exploration of the whole chemical space is completed, and the
drug design efficiency is improved. The
lead compound reaching the human expert level is designed in
actual use, and a new technical means is provided for drug research and development.