The invention belongs to the field of
large model training, and particularly relates to a multi-
modal data annotation and
large model thinking chain training method based on
visual interaction. The method comprises the following steps: 1, acquiring multi-
modal data; 2, displaying the multi-
modal data on a display, displaying the multi-
modal data to an expert through the display, recording an
eye movement track when the expert watches the multi-
modal data on the display by adopting an
eye movement data acquisition device, and acquiring voice information of the expert by adopting a voice acquisition device; 3, according to the
eye movement track and the voice information, forming multi-modal
annotation data; and 4, preprocessing the multi-modal
annotation data, and inputting the preprocessed multi-modal
annotation data into the
large model for training to obtain a trained large model. According to the method, the multi-modal labeling data of the experts in the labeling process can be collected, dominant
processing is carried out on the thinking chain and the thinking process of the experts, logic fusion is carried out, and the information breadth and accuracy of a large model are improved.