The invention discloses a training method, device and equipment for a large intelligent recognition model of archive files and a medium, and relates to the technical field of
document recognition. The training method comprises the following steps: constructing a self-supervised
diffusion model for first-stage training: carrying out random
mask processing on image samples to generate
mask image samples, respectively inputting the
mask image samples into an image
encoder to extract high-dimensional information, and further enhancing the discrimination of an attention map by using a token selection module, the weight of task related parameters is dynamically adjusted through an attention refocusing mechanism, the perceptual ability of the model to a task target is improved, a text
encoder embedded by an empty text is combined to serve as condition input of a
diffusion model, and the
encoder is optimized by using generation feedback of the
diffusion model; and constructing a second-stage fine-tuning Qwen-vl
large model: freezing the image encoder trained in the first stage, and finely tuning the Qwen-vl
large model by using a small number of samples. According to the method, the visual reasoning and fine-grained sensing capabilities of a large archive identification model in a complex scene are realized, and the generalization and precision of archive identification are improved.