An
alloy performance prediction method fusing vision-language-process multi-
modal data comprises the steps that a multi-
modal data set of an
alloy material is constructed, and the multi-
modal data set comprises three kinds of modal input including process parameters such as temperature and time of
solid solution and aging treatment, SEM image visual information and SEM
image description text information and corresponding
alloy mechanical property true values; training a visual
encoder ResNet50 model through comparative learning and training a language
encoder BERT model through
mask language modeling to obtain an SEM image and vector codes of language description of the SEM image; splicing and fusing process parameters such as temperature and time of
solid solution and aging treatment and the vector codes obtained in the second step, and training a
random forest regression device according to corresponding alloy mechanical properties; and the
random forest regression device obtained through training is used for alloy
performance prediction. According to the method, the structured process data, the unstructured SEM image data and the derived text description data are subjected to collaborative fusion and joint modeling for the first time, complementarity among different
modal data is fully utilized, and more comprehensive and more three-dimensional digital representation of the alloy state is constructed; the multi-modal fusion framework and the
feature learning mechanism are suitable for wide material systems. Meanwhile, the adopted'depth representation +
random forest '
hybrid model has the
advantage of high training efficiency while ensuring high prediction precision.