The invention discloses a deep
forgery detection method based on a visual
language model, and relates to the field of
image forensics. The deep
forgery detection method based on the visual
language model aims to combine multi-source information to improve the discrimination capability of the model on a
real image and a generated image. The method comprises the following steps: firstly, extracting image features through an image
encoder of a pre-trained CLIP model; meanwhile, a
frequency domain enhanced counterfeit
perception adapter is embedded in the image
encoder to mine potential anomalies of counterfeit images in the
image domain and the
frequency domain. Secondly, a manual
feature extraction module is provided, discriminative low-dimensional features are extracted from the four aspects of the edge, the texture, the frequency and the symmetry of the image, and the discriminative low-dimensional features are used as auxiliary information input in the
forgery detection process, so that the robustness and the
interpretability of the model are improved; meanwhile, the text cue words are converted into feature vectors through a text
encoder of a pre-training CLIP model; and finally, the model predicts a forgery
score by calculating the
cosine similarity between the image features and the text features so as to realize the discrimination of the authenticity of the image. According to the method, the problem that the detection capability of the model on the cross-dataset is insufficient is effectively improved.