The application discloses a training method and device of an
archive file intelligent identification
large model, equipment and a medium, and relates to the technical field of document identification. The training method comprises the following steps: a self-supervised
diffusion model for first-stage training is built; image samples are subjected to random
mask processing to generate
mask image samples, which are respectively input into an image
encoder to extract high-dimensional information; a tokens selection module is used to further enhance the discriminability of an attention map, and an attention refocusing mechanism is used to dynamically adjust the weight of task-related parameters, improve the
perception ability of the model to the task target, and combine a text
encoder with empty text embedding as the conditional input of the
diffusion model, and the feedback generated by the
diffusion model is used to optimize the
encoder; a Qwen-vl
large model for second-stage
fine tuning is built; the image encoder for first-stage training is frozen, and the Qwen-vl
large model is fine tuned by using a small amount of samples. The application realizes the visual reasoning and fine-grained
perception ability of the archive identification large model in a complex scene, and improves the generalization and precision of archive identification.