The invention discloses a human preference alignment method based on a self-improvement large visual
language model. The method comprises the following steps: automatically constructing a
preference data set; acquiring an image-question pair from the visual question and answer
data set; applying various visual enhancements to the image, and driving a
reference model to generate a group of candidate answers in combination with a question; the group of candidate answers are summarized and extracted into a more comprehensive'win 'text, and meanwhile, the text answers of the
reference model to the original picture and the question are taken as'fall-fail' texts, so that
preference data in an'image-inquiry-'win 'text-'fall-fail' text 'format are formed; secondly, based on the
preference data set constructed in the first step, a direct preference optimization
algorithm is applied to conduct alignment fine adjustment on a target visual
language model, in the fine adjustment process, parameters of a visual
encoder are kept frozen, and a low-rank
adaptation layer (LoRA) is only introduced into an extended mode alignment module and a language decoder for training; carrying out iterative self-improvement on the model; after
fine tuning is completed, taking the optimized model as a new
reference model, and repeating the data construction process in the step 1 to generate preference data with higher quality for the next round of optimization; and the circulation is repeated, so that the continuous self-improvement of the model alignment capability is realized. According to the method, self-supervised preference alignment without
manual annotation is realized, and a more comprehensive and high-quality'win 'text is generated.