A complex bait
data set construction method based on structure and sequence collaborative redundancy
elimination belongs to the field of
bioinformatics, and comprises the following steps: screening an initial
protein complex structure set, removing entries containing nucleic acids, small molecules or non-
protein chains, and selecting binary complexes meeting integrity and resolution requirements; secondly, structure clustering and
sequence clustering are carried out based on three-dimensional structure similarity and
sequence homology, combined comparison is carried out on the two results, and redundant compound entries which are highly similar in structure and sequence are removed; then, taking each cluster representative compound as a target, generating a plurality of groups of bait structures by using a molecular docking or prediction modeling method, and calculating a quality index; and finally, performing
stratified sampling and proportion balance based on the
score interval of the quality index, and constructing a high-quality
protein complex bait
data set with structure and sequence collaborative redundancy
elimination and balanced quality distribution. The
data set generated by the method has the advantages of low redundancy, high diversity and quality distribution
controllability.