Image copying recognition method and device, computer equipment and storage medium

By extracting sub-pictures from images and using ViT models for training and recognition, the problem of low recognition accuracy and insufficient ability to apply complex scenes in image remake recognition is solved, and higher recognition accuracy and generalization capabilities are achieved.

CN119992187APending Publication Date: 2025-05-13SHENZHEN GUANJIAN ZHILIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510066088.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has problems in image remake recognition with low recognition accuracy and insufficient ability to apply to complex scenes, especially in cases of poor image quality or complex background.

Method used

By extracting sub-maps at appropriate positions from the remake and the original image, a pair of training data sets are formed, and these sub-maps are trained based on the ViT model to obtain a finely tuned ViT model. Then, this model is used to process the sub-map of the recognized image to determine whether the image is a remake image.

Benefits of technology

It improves the accuracy of image remake recognition and the ability to apply to complex scenes, significantly reduces the misjudgment rate, and improves the generalization ability of the model on new data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992187A_ABST
    Figure CN119992187A_ABST
Patent Text Reader

Abstract

The invention relates to an image copying recognition method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring a duplicated training image and an original image corresponding to the duplicated training image, acquiring a first part sub-image with a preset area from the duplicated training image, acquiring a second part sub-image with a preset area and a corresponding position from the original image, and forming a paired training data set; training the ViT model based on the training data set to obtain a fine-tuned ViT model; and acquiring a to-be-identified image, acquiring a third part of sub-image with a preset area from the to-be-identified image, and processing the third part of sub-image by using the fine-tuned ViT model to judge whether the to-be-identified image is a copied image. According to the method, the interference of global noise is effectively avoided, and the recognition accuracy is improved, so that the method can be suitable for the condition of poor image quality or complex background, and the misjudgment rate can be remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image remake recognition method, device, computer equipment and storage medium. Background Art

[0002] In the field of image duplication recognition, the prior art mainly focuses on the application of image processing and computer vision algorithms, aiming to effectively solve the problem of recognition and analysis of duplication images. In recent years, deep learning methods have been widely used in image recognition tasks. These methods can automatically extract features and show significant performance advantages in duplication recognition tasks. For example, the prior art discloses a solution: high-frequency and low-frequency filtered images of an image are obtained by filtering; then the high-frequency and low-frequency filtered images are downsampled to obtain multi-scale filtered images; finally, a neural network is used for classification to obtain the duplication classification result. The prior art discloses another solution: a method of using the normalized intensity value in the image as the frequency intensity weight of the moire fringe. The CNN model created by this method is robust to high background frequencies other than the moire fringe. However, the applicant found that there are still many shortcomings and limitations in the traditional duplication recognition algorithm, which not only affect the recognition accuracy, but also limit its application in complex scenes. Summary of the invention

[0003] Based on this, it is necessary to provide an image copy recognition method, device, computer equipment and storage medium that can improve recognition accuracy and be applicable to complex scenes in response to the above technical problems.

[0004] A method for identifying a duplicate image comprises the following steps:

[0005] Acquire a re-shot training image and an original image corresponding to the re-shot training image, collect a first partial sub-image of a preset area from the re-shot training image, and collect a second partial sub-image of a preset area and corresponding position from the original image to form a paired training data set; wherein the first partial sub-image includes moiré patterns; and the positions of the first partial sub-image and the second partial sub-image in the image are in a one-to-one correspondence;

[0006] Train the ViT model based on the training data set to obtain a fine-tuned ViT model;

[0007] The image to be identified is obtained, a third sub-image of a preset area is collected from the image to be identified, and the third sub-image is processed using the fine-tuned ViT model to determine whether the image to be identified is a re-photographed image.

[0008] In one embodiment, the steps of obtaining a re-shot training image and an original image corresponding to the re-shot training image, collecting a first partial sub-image of a preset area from the re-shot training image, and collecting a second partial sub-image of a preset area and corresponding in position from the original image to form a paired training data set include the steps of:

[0009] Randomly collect a first sub-image of a preset area from the re-shot training image; the distance between the center points of any two first sub-images is greater than the product of the side length of the first sub-image and the square root of two;

[0010] Collecting a second sub-image in the original image that corresponds to the position of the first sub-image one by one;

[0011] The first sub-image is compared with the corresponding second sub-image. If the first sub-image has moiré patterns, the first sub-image is retained in the first part of the sub-image, and the second sub-image corresponding to the position of the first part of the sub-image is retained in the second part of the sub-image to form a training data set.

[0012] In one embodiment, the step of randomly collecting a first sub-image of a preset area from the re-photographed training image includes the steps of:

[0013] Randomly collect the first sub-image from the re-shot training image as the first sampling reference;

[0014] After collecting the first sampling reference, randomly collect a second first sub-image from the re-shot training image, and calculate the distance between the center point of the first sampling reference and the second first sub-image;

[0015] If the distance between the first sampling reference and the center point of the second first sub-graph is greater than the product of the side length of the first sub-graph and the square root of two, then the second first sub-graph is retained and added to the sampling reference, otherwise the second first sub-graph is deleted;

[0016] Randomly collect the Nth first sub-image from the re-shot training image, calculate the distance between all the first sub-images in the sampling reference and the center point of the Nth first sub-image, and determine whether the distance is greater than the product of the side length of the first sub-image and the square root of two. If the distance is greater than the product of the side length of the first sub-image and the square root of two, add the Nth first sub-image to the sampling reference; until the number of sampling references randomly collected from the re-shot training image reaches a preset number, the sampling ends, and all the first sub-images in the sampling reference are stored in the first part of the sub-image.

[0017] In one embodiment, the step of training the ViT model based on the training data set to obtain a fine-tuned ViT model includes the following steps:

[0018] Use the embedding layer of the ViT model to convert the training data input in batches into a vector sequence;

[0019] The ViT encoder of the ViT model is used to extract features of the vector sequence;

[0020] The features of the vector sequence obtained by processing the ViT encoder of the ViT model are input into the classifier, and the features of the vector sequence are classified by the classifier. During the classification process, the ViT model is trained based on the difference between the first part sub-graph and the second part sub-graph to obtain a fine-tuned ViT model.

[0021] In one embodiment, the step of obtaining an image to be identified, collecting a third sub-image of a preset area from the image to be identified, and processing the third sub-image using a fine-tuned ViT model to determine whether the image to be identified is a re-photographed image includes the following steps:

[0022] Input a third part sub-image into the fine-tuned ViT model for inference, and obtain the inference result and the corresponding prediction probability; the inference result is that the third part sub-image contains moiré or the third part sub-image does not contain moiré;

[0023] If the inference result is that the third sub-image contains moiré, then the prediction probability of the third sub-image containing moiré is verified, and the third sub-image corresponding to the prediction probability greater than the preset value is used as a judgment reference, and the prediction probability greater than the preset value is used as the prior probability that the image to be identified is a re-photographed image; if the inference result is that the third sub-image does not contain moiré, then the next third sub-image is predicted;

[0024] A fourth sub-image is collected around the judgment reference, and the fine-tuned ViT model is used to predict the probability that the fourth sub-image contains moiré patterns;

[0025] Based on the predicted probability and prior probability that the fourth sub-image contains moiré, the posterior probability that the image to be identified is a re-photographed image is obtained. If the posterior probability is greater than a preset probability threshold, the image to be identified is determined to be a re-photographed image.

[0026] In one embodiment, the step of obtaining an image to be identified, collecting a third sub-image of a preset area from the image to be identified, and processing the third sub-image using a fine-tuned ViT model to determine whether the image to be identified is a re-photographed image includes the following steps:

[0027] Input the third part of the sub-image into the fine-tuned ViT model for inference to obtain an inference result; the inference result is that the third part of the sub-image contains moiré or the third part of the sub-image does not contain moiré;

[0028] If the inference result is that the number of sub-images containing moiré patterns in the third part of sub-images is greater than a preset number threshold, it is determined that the image to be identified is a re-photographed image.

[0029] In one embodiment, the preset area is 224 pixels×224 pixels.

[0030] An image remake recognition device, comprising:

[0031] A data acquisition module is used to obtain a re-shot training image and an original image corresponding to the re-shot training image, collect a first partial sub-image of a preset area from the re-shot training image, and collect a second partial sub-image of a preset area and corresponding position from the original image to form a paired training data set; wherein the first partial sub-image includes moiré patterns; and the positions of the first partial sub-image and the second partial sub-image in the image are in a one-to-one correspondence;

[0032] The model training module is used to train the ViT model based on the training data set to obtain a fine-tuned ViT model;

[0033] The image judgment module is used to obtain the image to be identified, collect a third sub-image of a preset area from the image to be identified, and process the third sub-image using a fine-tuned ViT model to determine whether the image to be identified is a re-photographed image.

[0034] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0035] A computer-readable storage medium stores a computer program, which implements the steps of the above method when executed by a processor.

[0036] One of the above technical solutions has the following advantages and beneficial effects:

[0037] The image remake identification method provided by each embodiment of the present application is through the following steps: obtaining a remake training image and an original image corresponding to the remake training image, collecting a first part of the sub-image of a preset area from the remake training image, and collecting a second part of the sub-image of a preset area and corresponding position from the original image to form a paired training data set; training the ViT model based on the training data set to obtain a fine-tuned ViT model; obtaining an image to be identified, collecting a third part of the sub-image of a preset area from the image to be identified, and processing the third part of the sub-image using the fine-tuned ViT model to determine whether the image to be identified is a remake image. The present application extracts sub-images at appropriate positions from the remake image and the original image, effectively avoiding the interference of global noise, and improving the accuracy of recognition, so that the present application can be applied to situations with poor image quality or complex background, and can significantly reduce the misjudgment rate. The present application can integrate the prediction results of multiple sub-images, output the final remake judgment in the form of statistical probability, and further improve the reliability of the recognition result. In addition, this application adopts the ViT model to fully extract features of the input sub-image through a fine-tuning process, so that the model can accurately identify moiré patterns and original image features, thereby improving the model's generalization ability on new data. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 Schematic diagram of the process of the image copy recognition method in the embodiment of the present application.

[0039] Figure 2 Schematic diagram of the process of obtaining a training data set in an embodiment of the present application.

[0040] Figure 3 It is a schematic diagram of the process of collecting the first initial sub-image step in the embodiment of the present application.

[0041] Figure 4 Schematic diagram of the model training process in the embodiment of the present application.

[0042] Figure 5 A flowchart of the copy-shooting determination step in an embodiment of the present application.

[0043] Figure 6 Another flowchart diagram of the copy-shooting determination step in the embodiment of the present application.

[0044] Figure 7 Schematic diagram of the internal structure of a computer device in an embodiment of the present application. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0046] Traditional copy recognition algorithms have the following shortcomings and limitations, which not only affect the recognition accuracy, but also limit their application in complex scenarios:

[0047] 1. Use the entire reproduced image as the recognition target: In reproduced images, moiré patterns are a local feature that often only appear in local areas of the image. This global analysis method may mislead the algorithm, causing noise in irrelevant areas of the image to interfere with the final recognition results. As a result, the algorithm may regard noise as a valid feature, thereby reducing the accuracy of recognition, especially in the case of poor image quality or complex background.

[0048] 2. Remake recognition algorithm using an end-to-end framework: The remake recognition framework using an end-to-end framework simplifies the model building process in theory, but faces problems in practical applications. This framework often forces the model to forcibly fit the features of the training data instead of extracting the specific features required. This limitation of feature learning may cause the model to be unable to capture truly meaningful patterns when identifying remake images, but instead learn noise features that are irrelevant to the remake features. In addition, since remake images often have larger pixels, the model requires stronger computing power for reasoning, resulting in slower calculation speeds. This not only increases processing time, but may also affect the model's ability to extract features, thereby further reducing the model's accuracy and generalization ability on new data, making it difficult for the algorithm to adapt to diverse remake scenarios.

[0049] 3. Insufficient training data limits algorithm performance: Without rich and diverse training samples, the model cannot fully learn important features for copy recognition. In particular, when the model faces unseen moiré patterns or complex backgrounds, the impact of insufficient training data is particularly evident. This data scarcity not only limits the model's learning ability, but can also lead to unstable recognition results.

[0050] In order to solve the above problem, in one embodiment, Figure 1 As shown, a method for identifying a duplicate image is provided, comprising the following steps:

[0051] Step S110, obtain the reshot training image and the original image corresponding to the reshot training image, collect the first partial sub-image of a preset area from the reshot training image, and collect the second partial sub-image of a preset area and corresponding position from the original image to form a paired training data set. The reshot training image is an image obtained by directly shooting the original image with a camera. Collect the first partial sub-image and the second partial sub-image from the reshot training image and the original image to obtain the training data set. It should be noted that the number of the first partial sub-images is multiple, which is a partial image intercepted from the reshot training image, and the first partial sub-image includes moiré. The number of the second partial sub-images is multiple, which is a partial image intercepted from the original image, and the positions of the first partial sub-image and the second partial sub-image in the picture are in a one-to-one correspondence, wherein the one-to-one correspondence means that the acquisition position is the same, and the area of ​​the image is the same. In one example, the preset area is 224 pixels × 224 pixels.

[0052] In one example, if Figure 2 As shown, the steps of obtaining a reshot training image and an original image corresponding to the reshot training image, collecting a first partial sub-image of a preset area from the reshot training image, and collecting a second partial sub-image of a preset area and corresponding position from the original image to form a paired training data set include the steps of:

[0053] Step S210, randomly collect a first sub-image of a preset area from the re-shot training image, wherein the distance between the center points of any two first sub-images is greater than the product of the side length of the first sub-image and the square root of two, so as to ensure that the sub-images do not overlap with each other, avoid introducing redundant information in the model training, and improve the accuracy and efficiency of the algorithm.

[0054] Step S220, collect a second sub-image in the original image that corresponds to the first sub-image position. After obtaining the first sub-image, based on the collection position and collection area of ​​the first sub-image, collect a second sub-image with the same area as the first sub-image at the same collection position in the original image.

[0055] Step S230, compare the first sub-image with the corresponding second sub-image. If the first sub-image has moiré, the first sub-image is retained in the first sub-image, and the second sub-image corresponding to the position of the first sub-image is retained in the second sub-image, forming a training data set. Since the original image does not contain moiré, the part of the re-shot training image contains moiré. By comparing the first sub-image with the corresponding second sub-image, the first sub-image containing moiré can be found from the first sub-image.

[0056] In order to ensure that the collected sub-images do not overlap, in an example, Figure 3 As shown, the step of randomly collecting a first sub-image of a preset area from a re-shot training image includes the following steps:

[0057] Step S310: randomly collect a first sub-image from the re-shot training image as a first sampling reference.

[0058] Step S320: After collecting the first sampling reference, randomly collect a second first sub-image from the re-shot training image, and calculate the distance between the center point of the first sampling reference and the second first sub-image.

[0059] The present invention uses the following formula to calculate the distance between the center points of each small block:

[0060]

[0061] Among them, (x1, y1) and (x2, y2) are the center coordinates of the two first initial subgraphs respectively.

[0062] Specifically, for the newly collected first sub-images, the following steps are performed to ensure that the collected sub-images do not overlap. First, the center coordinates of each first sub-image are calculated:

[0063]

[0064] Among them, center_new represents the coordinates of the center point of the first sub-image newly collected after the first sub-image, (x new ,ynew ) is the upper left corner coordinate of the position of the first sub-image, L is the side length of the first sub-image, and here, the selected first sub-image is a square.

[0065] For each sampled first sub-image, its position is (p x ,p y ), it is necessary to calculate the distance between the center point of the newly acquired first sub-image and the center points of all the sampled first sub-images:

[0066]

[0067] If for each sampled first initial sub-graph (p x ,p y ), both Then the newly acquired first sub-image can be sampled and added to the sampling reference.

[0068] Step S330 , if the distance between the first sampling reference and the center point of the second first sub-image is greater than the product of the side length of the first sub-image and the square root of two, retain the second first sub-image and add the second first sub-image to the sampling reference; otherwise, delete the second first sub-image.

[0069] Step S340, randomly collect the Nth first sub-image from the re-shot training image, calculate the distance between all the first sub-images in the sampling reference and the center point of the Nth first sub-image, and determine whether the distance is greater than the product of the side length of the first sub-image and the square root of two. If the distance is greater than the product of the side length of the first sub-image and the square root of two, then add the Nth first sub-image to the sampling reference; until the number of sampling references randomly collected from the re-shot training image reaches a preset number, the sampling ends, and all the first sub-images in the sampling reference are stored in the first part of the sub-image. It should be noted that the preset number is set according to actual needs.

[0070] Specifically, a sub-image is cropped from the re-photographed image, and then a sub-image at the same position is cropped from the corresponding original image. By comparing the two types of sub-images, sub-images with moiré features can be identified and used as part of the training data set, while those without moiré features are discarded. The benefit of this pre-processing method is that it can significantly improve the efficiency of model training and the accuracy of recognition. By ensuring that the sampling does not overlap, it avoids introducing too much repeated information in model training, which helps the model learn more accurate feature representations. At the same time, by focusing on image blocks with moiré features, the model can focus more on identifying the key features of the re-photographed image, thereby improving the accuracy of recognition. In addition, this method also helps to reduce the computing resources required for model training because it reduces the complexity of the data by removing irrelevant image blocks. In actual operation, factors such as image resolution, moiré distribution density, and image block size are also considered to ensure the quality of the data set and the effectiveness of model training. For example, for low-resolution images, a denser sampling strategy will be set to directly sample the adjacent grids of the image to ensure that as much data as possible is obtained. For image data with uneven moiré distribution, a trained simplified moiré detector is used to filter the data to obtain truly meaningful paired data (the first part of the sub-image and the second part of the sub-image) to reduce the impact of noise data.

[0071] In the prior art, the recognition of reproduced images often relies on a global analysis method, that is, taking the entire image as the recognition target. This method may mislead the algorithm due to noise in irrelevant areas of the image, especially in the case of poor image quality or complex background, resulting in a significant decrease in recognition accuracy. The present application converts the problem into the recognition of local features by extracting small sub-images at appropriate positions from the reproduced images and the original images, effectively avoiding the influence of global noise. This method not only improves the recognition accuracy, but also enables the algorithm to maintain high performance even in the case of poor image quality or complex background. Specifically, the present application obtains image pairs with moiré images and original images through non-overlapping sampling, thereby ensuring the validity and diversity of the input data, thereby improving the model's ability to recognize local features.

[0072] Step S120, training the ViT model based on the training data set to obtain a fine-tuned ViT model. The purpose of training the ViT model is to enable the ViT model to accurately identify moiré features to determine whether the image is a remake. The core of the ViT model is a ViT encoder, which receives an input image and converts it into a series of embedded vectors that can capture the deep features of the image.

[0073] In one example, if Figure 4As shown, the steps of training the ViT model based on the training data set to obtain the fine-tuned ViT model include the following steps:

[0074] Step S410, using the embedding layer of the ViT model to convert the training data input in batches into a vector sequence.

[0075] Step S420: extract features of the vector sequence using the ViT encoder of the ViT model.

[0076] Step S430, input the features of the vector sequence obtained by processing the ViT encoder of the ViT model into the classifier, and classify the features of the vector sequence through the classifier. During the classification process, the ViT model is trained based on the difference between the first part sub-graph and the second part sub-graph to obtain a fine-tuned ViT model.

[0077] Specifically, the ViT encoder of the ViT model migrates the weights of the open source ViT pre-trained on the ImageNet-21k dataset. The subgraph is converted into a vector representation that the ViT model can handle through the embedding layer. The embedding layer can be expressed as:

[0078] E(p)=Linear(p)+E pos

[0079] Where p is the image block, Linear(p) is the linear transformation, and convolution is used here to put the sub-image into the embedding space. pos is the positional encoding. Finally, the embedding layer output includes the CLS tag and the positional encoding:

[0080]

[0081] where y0 is the embedding of the CLS marker, y i is the i-th embedding of the subgraph, and N is the number of embeddings of the subgraph. The output z of this embedding layer is then input into the ViT encoder as a vector sequence for processing. The ViT encoder is used to extract the moiré features (features of the vector sequence) of the subgraph. It consists of multiple identical layers, each of which includes a multi-head self-attention mechanism (MHSA) and a feedforward neural network (MLP), as shown below:

[0082] MHSA(x)=x+MHSA(LN(x))

[0083] MLP(x)=x+MLP(LN(x))

[0084] ViT-Encoder(x)=MHSA(MLP(x))

[0085] Among them, x is the input of this layer and LN is layer normalization. These layers can capture the dependencies between subgraphs and the global context information.

[0086] The above steps are designed to allow the trained ViT model to truly learn the image features that it is expected to learn, namely the moiré features. Therefore, in the training dataset used, data with and without moiré exist in pairs. The ViT model can learn the features of a clean image itself, as well as the features when it is covered with moiré, so that it can learn the difference between a clean image and an image with moiré (the difference is caused by the image being rephotographed). Through the powerful feature extraction capability of the ViT encoder, the ViT model can learn the deep features of the image, so that it can accurately identify whether a sub-image has moiré. Through this learning method, the ViT model can better generalize to new, unseen images, even if the moiré patterns or background complexity of these images are different.

[0087] Due to the lack of rich and diverse training samples, the performance of existing algorithms is limited when faced with unseen moiré patterns or complex backgrounds. This application expands the data set by extracting an appropriate number of sub-images from the moiré images and the original images, thereby enhancing the diversity and representativeness of the samples. This data augmentation method enables the model to have more available and valid data, allowing the model to more comprehensively learn features that are critical to copy recognition, thereby improving the algorithm's capabilities and the stability of recognition. Through effective data processing and model training, this application aims to improve the accuracy and generalization of copy recognition.

[0088] Traditional end-to-end frameworks may force the model to fit the features of the training data in remake recognition, including noise features that are not related to remakes. This application adopts a Bayesian remake recognition algorithm based on moiré, focusing on extracting and identifying features related to moiré, avoiding the model's learning of noise features. This improvement not only improves the reasoning speed of the model, but also enhances the model's ability to extract features, and improves the model's accuracy and generalization ability on new data. During the fine-tuning process, the model requires sufficient feature extraction of small blocks of input so that the model can learn truly useful discriminant features so that the model can accurately identify moiré and original image features, thereby improving recognition accuracy.

[0089] Step S130 , obtaining the image to be identified, collecting a third sub-image of a preset area from the image to be identified, and processing the third sub-image using the fine-tuned ViT model to determine whether the image to be identified is a re-photographed image.

[0090] In one way, Figure 5As shown, the step of obtaining an image to be identified, collecting a third sub-image of a preset area from the image to be identified, and processing the third sub-image using a fine-tuned ViT model to determine whether the image to be identified is a re-photographed image includes the following steps:

[0091] Step S510, input a third partial sub-image into the fine-tuned ViT model for inference, and obtain an inference result and a corresponding prediction probability; the inference result is that the third partial sub-image contains moiré or the third partial sub-image does not contain moiré.

[0092] Specifically, the embedding layer of the fine-tuned ViT model converts the third subgraph into a vector representation that can be processed by the ViT model. The embedding layer can be expressed as:

[0093] E(p)=Linear(p)+E pos

[0094] Where p is the image block, Linear(p) is the linear transformation, and convolution is used here to put the third sub-image into the embedding space. pos is the positional encoding. Finally, the embedding layer output includes the CLS tag and the positional encoding:

[0095]

[0096] where y0 is the embedding of the CLS marker, y i is the i-th embedding of the third sub-image, and N is the number of embeddings of the third sub-image. The output z of this embedding layer will then be input into the ViT encoder as the vector sequence corresponding to the third sub-image for processing. The ViT encoder is used to extract the moiré features (feature vectors) of the third sub-image. It consists of multiple identical layers, each of which includes a multi-head self-attention mechanism (MHSA) and a feed-forward neural network (MLP), as shown below:

[0097] MHSA(x)=x+MHSA(LN(x))

[0098] MLP(x)=x+MLP(LN(x))

[0099] ViT-Encoder(x)=MHSA(MLP(x))

[0100] Among them, x is the input of this layer, and LN is layer normalization. These layers can capture the dependencies between subgraphs and global context information. Finally, the feature vector processed by the ViT encoder is fed into the classification head, which consists of MLP layers to map the feature vector to the final category output, as shown below:

[0101] Output=σ(MLP(ViT output ))

[0102] Here, Output is the predicted probability after being processed by the sigmoid function, and σ is the sigmoid function, which converts the output of MLP to between 0 and 1. Finally, the fine-tuned ViT model outputs the inference results of each third sub-image and the corresponding predicted probability. These prediction results can be used for further analysis, such as determining whether the image contains moiré patterns, so as to achieve accurate copy recognition.

[0103] Step S520, if the inference result is that the third partial sub-image contains moiré, then verify the prediction probability of the third partial sub-image containing moiré, take the third partial sub-image corresponding to the prediction probability greater than the preset value as a judgment reference, and take the prediction probability greater than the preset value as the prior probability that the image to be identified is a retaken image; if the inference result is that the third partial sub-image does not contain moiré, then continue to predict the next third partial sub-image.

[0104] Step S530 , collecting a fourth sub-image around the judgment reference, and using the fine-tuned ViT model to predict the probability that the fourth sub-image contains moiré.

[0105] Step S540, based on the predicted probability and the prior probability that the fourth partial sub-image contains moiré, obtain the posterior probability that the image to be identified is a re-photographed image, and if the posterior probability is greater than a preset probability threshold, determine that the image to be identified is a re-photographed image.

[0106] After the fine-tuned ViT model infers the third part of the sub-image, the inference results and prediction probabilities of each third sub-image will be obtained. Next, different strategies will be implemented according to different inference results. When a third sub-image is predicted to contain no moiré, the next third sub-image will continue to be predicted. If a third sub-image is predicted to contain moiré, the fine-tuned ViT model will further verify the prediction probability of the third sub-image and the prediction probability of the resampled sub-image around it, that is, the fourth part of the sub-image. When the moiré prediction probability of a sub-image is high, the moiré features around it will also be more obvious. This is mainly because moiré is generally a large-scale feature with a larger coverage range. Therefore, if a more obvious moiré feature appears in a certain position of the image, it is often accompanied by the appearance of moiré features around it, and it is necessary to sample again around this position. When resampling, the center point of the sub-image position needs to be used as the center of the circle. As the radius, sampling is performed on the edge of this circle. The sampling point is used as the center point of the newly sampled sub-image (the fourth sub-image). Each fourth sub-image is sampled at a fixed angle interval to ensure that the fourth sub-images do not overlap. For the image remake recognition problem, if there is a third sub-image in the entire image with a very high probability of moiré, it can be used as a key area and its probability can be used as the initial prior probability, that is, the probability that the entire image is a remake. If this probability is very high, such as reaching 90% or 99% (preset value), this can be used as a prior probability. The discovery of this third sub-image can be regarded as part of the likelihood, that is, the probability of observing moiré when the image is a remake. The prior probability P(F) can be defined as the probability that the image to be identified has moiré, and P(D|F) is defined as the probability that the moiré feature appears in the sub-image. Where D represents the observed data (sub-image) and F represents the hypothesis (the image to be identified has moiré). The posterior prediction distribution can be obtained:

[0107]

[0108] Among them, P(F|D) is the posterior probability, and P(D) is the marginal probability of the data, which can be calculated by the total probability formula:

[0109]

[0110] The posterior probability that the whole image is a remake is updated by sampling more fourth sub-images around the third sub-image with a high probability of moiré. If a large number of the surrounding fourth sub-images are also judged to have moiré, this further provides evidence to support that the image is a remake, thereby increasing the posterior probability that the whole image is a remake. The posterior prediction distribution is the probability distribution of the fourth sub-image after resampling around the third sub-image, which is obtained by marginalizing the posterior probability:

[0111]

[0112] in, are multiple small blocks resampled around the small block of the image. A threshold probability is set. When the posterior probability exceeds this threshold probability, the image is directly judged to be a reprinted image. In general, it is first assumed that the prior probability of the entire image being a reprint is very high; through model analysis, when a sub-image in the image is found to have moiré features, this increases the likelihood of the reprint hypothesis; then further sampling is performed around the sub-image. If these surrounding sampled sub-images also show moiré features, this further increases the posterior probability of the reprint hypothesis; when the posterior probability exceeds the set threshold probability, the entire image is judged to be a reprinted image.

[0113] In another way, Figure 6As shown, the step of obtaining an image to be identified, collecting a third sub-image of a preset area from the image to be identified, and processing the third sub-image using a fine-tuned ViT model to determine whether the image to be identified is a re-photographed image includes the following steps:

[0114] Step S610, input the third partial sub-image into the fine-tuned ViT model for inference to obtain an inference result; the inference result is that the third partial sub-image contains moiré or the third partial sub-image does not contain moiré.

[0115] Step S620: If the inference result is that the number of sub-images containing moiré patterns in the third part of the sub-images is greater than a preset number threshold, it is determined that the image to be identified is a re-photographed image.

[0116] In this method, whether the image to be identified is a re-photographed image is determined by counting the number of third sub-images containing moiré patterns in the third partial sub-images.

[0117] The image remake recognition method provided by each embodiment of the present application is through the following steps: obtaining a remake training image and an original image corresponding to the remake training image, collecting a first part of the sub-image of a preset area from the remake training image, and collecting a second part of the sub-image of a preset area from the original image to form a training data set; training the ViT model based on the training data set to obtain a fine-tuned ViT model; obtaining an image to be identified, collecting a third part of the sub-image of a preset area from the image to be identified, and processing the third part of the sub-image using the fine-tuned ViT model to determine whether the image to be identified is a remake image. The present application extracts sub-images at appropriate positions from the remake image and the original image, effectively avoiding the interference of global noise, and improving the accuracy of recognition, so that the present application can be applied to situations with poor image quality or complex background, and can significantly reduce the misjudgment rate. The present application can integrate the prediction results of multiple sub-images, output the final remake judgment in the form of statistical probability, and further improve the reliability of the recognition result. In addition, the present application adopts the ViT model, and fully extracts the features of the input sub-image through the fine-tuning process, so that the model can accurately identify moiré patterns and original image features, thereby improving the generalization ability of the model on new data. Since the present application can sample sub-images from the entire image for discrimination, it means that the present application can work effectively regardless of the original resolution of the image. This flexibility allows the present application to be widely applied to images of different sources and qualities, including low-resolution images, without the need for preprocessing or adjusting the resolution, because preprocessing or adjusting the resolution may destroy the moiré features of the image, resulting in a decrease in the accuracy of recognition. Therefore, this design of the present application improves the practicality and applicability of the algorithm to a certain extent.

[0118] The Bayesian remake recognition algorithm based on moiré patterns in this application has significant application value in the retail business, especially in the store inspection tasks on the sales side. In this field, real-time and authenticity are crucial for evaluating the work performance of stores and sales representatives. The traditional store inspection process relies on real-time pictures uploaded by sales representatives, which are used to assess store displays and the work performance of sales representatives. However, there is a risk of cheating in this process. Sales representatives may meet the performance evaluation standards by reshooting excellent display pictures on the display screen, thereby affecting the fairness and accuracy of the performance evaluation.

[0119] The application of this application can effectively supervise and eliminate such cheating behaviors. By integrating this algorithm into the store inspection process, it is possible to automatically detect whether the uploaded pictures are re-photographed, thereby ensuring the authenticity of the pictures. The implementation of this technology not only improves the efficiency of store inspection tasks, but also enhances the accuracy and reliability of performance appraisals. Specifically, this algorithm can accurately identify re-photographed pictures by analyzing the moiré features in the image, thereby avoiding performance misjudgments caused by falsified pictures.

[0120] In practical applications, this application can be used as an automated tool to assist store inspectors in quickly verifying the authenticity of images during field inspections. In addition, this algorithm can also be integrated into the performance management system of the retail industry as a background verification mechanism to automatically screen and mark suspicious re-photographed images. In this way, management can make decisions based on more accurate data, optimize store displays, improve the work efficiency of sales representatives, and ultimately improve the operational efficiency and market competitiveness of the entire retail business.

[0121] It should be understood that although Figure 1-Figure 6 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1-Figure 6 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0122] In one embodiment, a device for identifying a duplicate image is provided, the device comprising:

[0123] A data acquisition module is used to obtain a re-shot training image and an original image corresponding to the re-shot training image, collect a first partial sub-image of a preset area from the re-shot training image, and collect a second partial sub-image of a preset area and corresponding position from the original image to form a paired training data set; wherein the first partial sub-image includes moiré patterns; and the positions of the first partial sub-image and the second partial sub-image in the image are in a one-to-one correspondence;

[0124] The model training module is used to train the ViT model based on the training data set to obtain a fine-tuned ViT model;

[0125] The image judgment module is used to obtain the image to be identified, collect a third sub-image of a preset area from the image to be identified, and process the third sub-image using a fine-tuned ViT model to determine whether the image to be identified is a re-photographed image.

[0126] The specific definition of the image duplication identification device can be found in the definition of the image duplication identification method above, which will not be repeated here. Each module in the above-mentioned image duplication identification device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0127] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store image data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an image copy recognition is realized.

[0128] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0129] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:

[0130] Acquire a re-shot training image and an original image corresponding to the re-shot training image, collect a first partial sub-image of a preset area from the re-shot training image, and collect a second partial sub-image of a preset area and corresponding position from the original image to form a paired training data set; wherein the first partial sub-image includes moiré patterns; and the positions of the first partial sub-image and the second partial sub-image in the image are in a one-to-one correspondence;

[0131] Train the ViT model based on the training data set to obtain a fine-tuned ViT model;

[0132] The image to be identified is obtained, a third sub-image of a preset area is collected from the image to be identified, and the third sub-image is processed using the fine-tuned ViT model to determine whether the image to be identified is a re-photographed image.

[0133] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0134] Acquire a re-shot training image and an original image corresponding to the re-shot training image, collect a first partial sub-image of a preset area from the re-shot training image, and collect a second partial sub-image of a preset area and corresponding position from the original image to form a paired training data set; wherein the first partial sub-image includes moiré patterns; and the positions of the first partial sub-image and the second partial sub-image in the image are in a one-to-one correspondence;

[0135] Train the ViT model based on the training data set to obtain a fine-tuned ViT model;

[0136] The image to be identified is obtained, a third sub-image of a preset area is collected from the image to be identified, and the third sub-image is processed using the fine-tuned ViT model to determine whether the image to be identified is a re-photographed image.

[0137] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0138] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0139] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be construed as limiting the scope of the patent application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent application shall be subject to the attached claims.

Claims

1. A method for identifying a duplicated image, characterized in that: The following steps are involved: Acquire a re-shot training image and an original image corresponding to the re-shot training image, collect a first partial sub-image of a preset area from the re-shot training image, and collect a second partial sub-image of a preset area and corresponding position from the original image to form a paired training data set; wherein the first partial sub-image includes moiré patterns; and the positions of the first partial sub-image and the second partial sub-image in the image are in a one-to-one correspondence; Training the ViT model based on the training data set to obtain a fine-tuned ViT model; An image to be identified is obtained, a third partial sub-image of a preset area is collected from the image to be identified, and the third partial sub-image is processed using the fine-tuned ViT model to determine whether the image to be identified is a re-photographed image.

2. The image copy recognition method according to claim 1, characterized in that: The step of obtaining a re-shot training image and an original image corresponding to the re-shot training image, collecting a first partial sub-image of a preset area from the re-shot training image, and collecting a second partial sub-image of a preset area and corresponding position from the original image to form a paired training data set comprises the steps of: A first sub-image of the preset area is randomly collected from the re-shot training image; the distance between the center points of any two of the first sub-images is greater than the product of the side length of the first sub-image and the square root of two; Collecting the second sub-image in the original image in one-to-one correspondence with the first sub-image position; The first sub-image is compared with the corresponding second sub-image. If the first sub-image has moiré patterns, the first sub-image is retained in the first partial sub-image, and the second sub-image corresponding to the position of the first partial sub-image is retained in the second partial sub-image to form the training data set.

3. The image copy recognition method according to claim 2, characterized in that: The step of randomly collecting the first sub-image of the preset area from the re-shot training image comprises the steps of: Randomly collect the first first sub-image from the re-shot training image as a first sampling reference; After collecting the first sampling reference, randomly collecting a second first sub-image from the re-shot training image, and calculating the distance between the center point of the first sampling reference and the second first sub-image; If the distance between the first sampling reference and the center point of the second first sub-graph is greater than the product of the side length of the first sub-graph and the square root of two, retain the second first sub-graph and add the second first sub-graph to the sampling reference; otherwise, delete the second first sub-graph; Randomly collect the Nth first sub-image from the re-shot training image, calculate the distance between all the first sub-images in the sampling reference and the center point of the Nth first sub-image, and determine whether the distance is greater than the product of the side length of the first sub-image and the square root of two; if the distance is greater than the product of the side length of the first sub-image and the square root of two, add the Nth first sub-image to the sampling reference; until the number of the sampling references randomly collected from the re-shot training image reaches a preset number, sampling ends, and all the first sub-images in the sampling reference are stored in the first partial sub-image.

4. The image copy recognition method according to claim 1, characterized in that: The step of training the ViT model based on the training data set to obtain a fine-tuned ViT model includes the following steps: Using the embedding layer of the ViT model, the training data input in batches is converted into a vector sequence; Extracting features of the vector sequence using a ViT encoder of the ViT model; The features of the vector sequence obtained by processing the ViT encoder of the ViT model are input into a classifier, and the features of the vector sequence are classified by the classifier. During the classification process, the ViT model is trained based on the difference between the first partial sub-graph and the second partial sub-graph to obtain the fine-tuned ViT model.

5. The image copy recognition method according to claim 1, characterized in that: The step of obtaining an image to be identified, collecting a third partial sub-image of a preset area from the image to be identified, and processing the third partial sub-image using the fine-tuned ViT model to determine whether the image to be identified is a re-photographed image comprises the following steps: Inputting one of the third partial sub-images into the fine-tuned ViT model for inference, and obtaining an inference result and a corresponding prediction probability; the inference result is that the third partial sub-image contains moiré or the third partial sub-image does not contain moiré; If the inference result is that the third partial sub-image contains moiré, then verify the prediction probability of the third partial sub-image containing moiré, take the third partial sub-image corresponding to the prediction probability greater than a preset value as a judgment reference, and take the prediction probability greater than the preset value as the prior probability that the image to be identified is a re-shot image; if the inference result is that the third partial sub-image does not contain moiré, then predict the next third partial sub-image; Collecting a fourth sub-image around the judgment reference, and using the fine-tuned ViT model to predict the probability that the fourth sub-image contains moiré; Based on the predicted probability that the fourth partial sub-image contains moiré and the prior probability, the posterior probability that the image to be identified is a copied image is obtained; if the posterior probability is greater than a preset probability threshold, the image to be identified is determined to be a copied image.

6. The image copy recognition method according to claim 1, characterized in that: The step of obtaining an image to be identified, collecting a third partial sub-image of a preset area from the image to be identified, and processing the third partial sub-image using the fine-tuned ViT model to determine whether the image to be identified is a re-photographed image comprises the following steps: Inputting the third partial sub-image into the fine-tuned ViT model for inference to obtain an inference result; the inference result is that the third partial sub-image contains moiré or the third partial sub-image does not contain moiré; If the inference result is that the number of sub-images containing moiré patterns in the third part of sub-images is greater than a preset number threshold, it is determined that the image to be identified is a re-photographed image.

7. The image copy recognition method according to any one of claims 1 to 6, characterized in that: The preset area is 224 pixels×224 pixels.

8. An image copy recognition device, characterized in that: The device comprises: A data acquisition module is used to obtain a re-shot training image and an original image corresponding to the re-shot training image, collect a first partial sub-image of a preset area from the re-shot training image, and collect a second partial sub-image of a preset area and corresponding position from the original image to form a paired training data set; wherein the first partial sub-image includes moiré patterns; and the positions of the first partial sub-image and the second partial sub-image in the image are in a one-to-one correspondence; A model training module, used to train the ViT model based on the training data set to obtain a fine-tuned ViT model; The image judgment module is used to obtain the image to be identified, collect a third sub-image of a preset area from the image to be identified, and process the third sub-image using the fine-tuned ViT model to determine whether the image to be identified is a re-photographed image.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.