The present application relates to the field of display technology and human-computer interaction, and provides a
flat panel display visual comfort prediction method and
system based on a multi-
modal fusion model, wherein the method comprises the following steps: step one, collecting an original image set, for each original image, using a
generative adversarial network model to convert an input
random sequence into enhanced materials, and then fusing the enhanced materials and the original image at a preset ratio to obtain a plurality of new images; step two, collecting physiological features and physical features to obtain training samples; step three, training a multi-
modal fusion model based on a stacked ensemble framework; and step four, inputting
test data to obtain a visual comfort prediction result. The present application aims to solve the problems of a lack of high-quality labeled samples, a lack of
coupling relationship in single-
modal analysis, and poor scene adaptability in the field of
flat panel display visual comfort prediction.