Method, apparatus, device, and computer-readable storage medium for identifying authenticity of an image
By extracting the histogram of pixel value distribution of the target area in image processing and generating a covariance feature matrix for classification, the problem of low accuracy of image authenticity and false identification in the prior art is solved, and higher identification accuracy is achieved.
Patent Information
- Application Number
- CN202111454061.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-01
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-12-01
AI Technical Summary
When identifying the authenticity of images, the modified area and the unmodified area have the same noise attributes, resulting in low accuracy of the identification results.
By extracting the target area in the image, counting the pixel values, generating a histogram of pixel value distribution, and generating a covariance feature matrix based on the histogram, and finally classifying the feature matrix to determine the authenticity of the image.
It effectively avoids the low discrimination accuracy caused by the same noise characteristics, and improves the accuracy of the authenticity of the image.
Smart Images

Figure CN114140431B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of image processing technologies, and particularly to a method, apparatus, device, and computer-readable storage medium for identifying the authenticity of an image. Background Art
[0002] The progress of image processing technologies has promoted the development and application of digital image editing tools. However, it is difficult to distinguish between an image modified by a digital image editing tool and a real image. Therefore, identifying the authenticity of an image has become an important research topic.
[0003] In related technologies, when identifying the authenticity of an image, after obtaining a target region that may be a modified region, the noise feature map obtained by filtering an RGB (Red Green Blue) image through an SRM (Steganalysis Rich Model) filtering layer is used as an input, and operations such as convolution and pooling are performed to complete the classification of whether the target region is a modified region. When there is a modified region in the target region, it is determined that the image is a modified image, that is, a fake image; otherwise, it is determined that the image is an unmodified image, that is, a real image.
[0004] In the above technology, the method of using the noise feature map obtained by filtering an RGB image through SRM to assist in identifying the authenticity of an image has a low accuracy rate of the result obtained when identifying the authenticity of an image when the modified region and the unmodified region in the image are from images taken by the same device. Since the modified region and the unmodified region have the same noise attributes, the effectiveness of the classification result obtained based on the noise feature map will be reduced, thereby resulting in a low accuracy rate of the result of identifying the authenticity of the image. Summary of the Invention
[0005] Embodiments of the present application provide a method, apparatus, device, and computer-readable storage medium for identifying the authenticity of an image, which can be used to solve the problems in related technologies. The technical solutions are as follows:
[0006] On the one hand, an embodiment of the present application provides a method for identifying the authenticity of an image, the method including:
[0007] Obtain an image to be identified;
[0008] Extract a target region in the image, count the pixel values in the target region, and obtain a pixel value distribution histogram according to the statistical result;
[0009] Generate a pixel value distribution feature map based on the pixel value distribution histogram;
[0010] Classify the pixel value distribution feature map to obtain a classification result, and determine the authenticity result of the image based on the classification result.
[0011] In a possible implementation manner, generating a pixel value distribution feature map based on the pixel value distribution histogram includes: calculating an average value of the number of pixels in each dimension of the pixel value distribution histogram; calculating a covariance between each dimension in the pixel value distribution histogram based on the average value; generating a covariance feature matrix based on the covariance, where the covariance feature matrix has a distribution feature of pixel values in the target region; and obtaining the pixel value distribution feature map based on the covariance feature matrix.
[0012] In a possible implementation manner, determining the authenticity result of the image based on the classification result includes: obtaining a cross-entropy loss function value between the classification result and an annotation result of the image; optimizing the classification result based on the minimum cross-entropy loss function value to obtain an optimization result; and determining the authenticity result of the image based on the optimization result.
[0013] In a possible implementation manner, extracting the target region in the image includes: using a Faster Region-based Convolutional Neural Network (Faster-RCNN) to extract the target region in the image; or using a Single Shot MultiBox Detector (SSD) of a single deep neural network to extract the target region in the image.
[0014] In a possible implementation manner, obtaining the pixel value distribution feature map based on the covariance feature matrix includes: extracting the distribution feature in the covariance feature matrix based on any Residual Neural Network (ResNet) to obtain the pixel value distribution feature map.
[0015] In a possible implementation manner, the target region is a region in the image that contains text.
[0016] On the other hand, a device for authenticating the authenticity of an image is provided. The device includes:
[0017] An acquisition module, configured to acquire an image to be authenticated;
[0018] An extraction module, configured to extract a target region in the image, count pixel values in the target region, and obtain a pixel value distribution histogram according to the statistical result;
[0019] A generation module, configured to generate a pixel value distribution feature map based on the pixel value distribution histogram;
[0020] A classification module, configured to classify the pixel value distribution feature map to obtain a classification result, and determine the authenticity result of the image based on the classification result.
[0021] In a possible implementation manner, the generating module is configured to calculate the average value of the number of pixels in each dimension of the pixel value distribution histogram; based on the average value, calculate the covariance between each dimension in the pixel value distribution histogram; based on the covariance, generate a covariance feature matrix, where the covariance feature matrix has the distribution characteristics of the pixel values in the target region; and obtain the pixel value distribution feature map based on the covariance feature matrix.
[0022] In a possible implementation manner, the classification module is configured to obtain the cross-entropy loss function value between the classification result and the annotation result of the image; optimize the classification result based on the minimum value of the cross-entropy loss function value to obtain an optimization result; and determine the authenticity result of the image based on the optimization result.
[0023] In a possible implementation manner, the extraction module is configured to use the Faster Region-based Convolutional Neural Network (Faster-RCNN) to extract the target region in the image; or use the Single Shot MultiBox Detector (SSD), a single deep neural network, to extract the target region in the image.
[0024] In a possible implementation manner, the generating module is configured to extract the distribution characteristics in the covariance feature matrix based on any Residual Neural Network (ResNet) to obtain the pixel value distribution feature map.
[0025] In a possible implementation manner, the target region is the region in the image that contains text.
[0026] On the other hand, a computer device is provided. The computer device includes a processor and a memory. At least one computer program is stored in the memory and is loaded and executed by the processor so that the computer device implements the method for authenticating the authenticity of an image according to any one of the above.
[0027] On the other hand, a computer-readable storage medium is further provided. At least one computer program is stored in the computer-readable storage medium and is loaded and executed by a processor so that a computer implements the method for authenticating the authenticity of an image according to any one of the above.
[0028] On the other hand, a computer program product is further provided. The computer program product includes a computer program or computer instructions, and the computer program or the computer instructions are loaded and executed by a processor so that a computer implements the method for authenticating the authenticity of an image according to any one of the above.
[0029] The technical solutions provided in the embodiments of the present application at least bring the following beneficial effects:
[0030] The technical solution provided by the embodiments of this application obtains a pixel value distribution histogram by counting the pixel values of the target regions extracted from the image. The pixel value distribution histogram can better retain the difference features between the modified region and the unmodified region in the image, and the difference features are represented by the distribution features of the pixel values. Obtaining a pixel value distribution feature map based on the pixel value distribution histogram can make the distribution features of the pixel values more stable. Then, classifying the pixel value distribution features and determining the authenticity of the image based on the classification results can avoid the low accuracy of authenticity identification caused by the same noise features in the modified region and the unmodified region of the image. Description of the Drawings
[0031] To more clearly illustrate the technical solutions in the embodiments of this application, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0032] Figure 1 It is a schematic diagram of an implementation environment for identifying the authenticity of an image provided by the embodiments of this application;
[0033] Figure 2 It is a flowchart of a method for identifying the authenticity of an image provided by the embodiments of this application;
[0034] Figure 3 It is an overall framework diagram for identifying the authenticity of an image provided by the embodiments of this application;
[0035] Figure 4 It is a schematic diagram of the process for extracting the target region provided by the embodiments of this application;
[0036] Figure 5 It is a schematic diagram of a pixel value distribution histogram provided by the embodiments of this application;
[0037] Figure 6 It is a schematic diagram of a 5-dimensional pixel value distribution histogram provided by the embodiments of this application;
[0038] Figure 7 It is a schematic diagram of a covariance feature matrix provided by the embodiments of this application;
[0039] Figure 8 It is a schematic diagram of a device for identifying the authenticity of an image provided by the embodiments of this application;
[0040] Figure 9 It is a schematic diagram of the structure of a server provided by the embodiments of this application;
[0041] Figure 10It is a schematic structural diagram of a terminal provided by an embodiment of the present application. Detailed implementation manners
[0042] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0043] An embodiment of the present application provides a method for authenticating the authenticity of an image. Please refer to Figure 1 , which shows a schematic diagram of the method implementation environment provided by the embodiment of the present application. The implementation environment may include: a terminal 11 and a server 12.
[0044] Among them, the terminal 11 can acquire an image, store the image, and upload the image to the server 12. The server 12 can receive the image transmitted from the terminal 11, store the image, and transmit the result after authenticating the image to the terminal 11.
[0045] The method for authenticating the authenticity of an image provided by the embodiment of the present application can be executed by the terminal 11, or by the server 12, or jointly by the terminal 11 and the server 12. The embodiment of the present application does not limit this. For the case where the method for authenticating the authenticity of an image provided by the embodiment of the present application is jointly executed by the terminal 11 and the server 12, the server 12 undertakes the main computing work, and the terminal 11 undertakes the secondary computing work; or, the server 12 undertakes the secondary computing work, and the terminal 11 undertakes the main computing work; or, a distributed computing architecture is adopted between the server 12 and the terminal 11 for collaborative computing.
[0046] Optionally, the terminal 11 can be any electronic product that can perform human-computer interaction with a user in one or more ways such as a keyboard, a touchpad, a touch screen, a remote control, voice interaction, or a handwriting device. For example, a PC (Personal Computer), a mobile phone, a smart phone, a PDA (Personal Digital Assistant), a wearable device, a PPC (Pocket PC), a tablet computer, a smart car machine, a smart TV, a smart speaker, etc. The server 12 can be a single server, or a server cluster composed of multiple servers, or a cloud computing service center. The terminal 11 and the server 12 establish a communication connection through a wired or wireless network.
[0047] Those skilled in the art should understand that the above-mentioned terminal 11 and server 12 are only examples. Other existing or future possible terminals or servers that can be applied to the present application should also be included within the protection scope of the present application and are hereby incorporated by reference.
[0048] Based on the above Figure 1In the implementation environment shown, an embodiment of the present application provides a method for authenticating the authenticity of an image. The method for authenticating the authenticity of the image is executed by a computer device, which can be the terminal 11 or the server 12. The embodiment of the present application does not limit this. As Figure 2 shown, the method provided by the embodiment of the present application may include the following steps 201 to 203.
[0049] In step 201, obtain the image to be authenticated, extract the target area in the image, count the pixel values in the target area, and obtain a pixel value distribution histogram according to the statistical result.
[0050] Among them, the image to be authenticated refers to the image whose authenticity needs to be judged. When the image has been modified, the image is a fake image. When the image has not been modified, the image is a genuine image. Optionally, the target area refers to the area that is mainly used when authenticating the authenticity of the image by using the method for authenticating the authenticity of the image provided by the embodiment of the present application. The embodiment of the present application does not limit what specific type of area the target area is. Exemplarily, when the image is a qualification certificate, the modification of the image usually focuses on the text part. Therefore, the target area at this time is the text area in the image.
[0051] In a possible implementation manner, extracting the target area in the image includes, but is not limited to, using a Faster-RCNN (Faster Regions Convolutional Neural Network) neural network to extract the target area in the image, or using an SSD (Single Shot MultiBox Detector) neural network to extract the target area in the image. In an exemplary embodiment, as Figure 3 shown, extracting the target area in the image corresponds to Figure 3 the process from the image to the target area in the figure. In an image, the target area may include multiple area blocks (for example, when there is text in the image, there are usually multiple areas containing text).
[0052] Among them, the purpose of extracting the target area in the image is that when the modified area in the image accounts for a relatively small proportion of the image itself, when authenticating the authenticity of the image based on the entire image, the characteristics of the modified part are easily covered by the characteristics of the entire image. Therefore, it is necessary to extract the target area where modifications may exist in the image for processing to further improve the accuracy of the authentication result.
[0053] Exemplarily, as Figure 4As shown in the figure, the process of using the Faster-RCNN neural network to extract the target region in an image is as follows: The image is input into the convolutional layer for convolution to obtain the convolutional feature map Feature maps of the image; the Feature maps are input into the RPN (Region Proposal Network) layer to generate Regions Proposal that may be the target region; in the RoI pooling (Regions of Interest pooling) layer, the Feature maps and Regions Proposal are processed to obtain a multi-dimensional feature vector; the feature vector is input into the fully connected layer, and the Regions Proposal corresponding to the obtained feature vector after being processed by the fully connected layer is mapped to the image marking the true target region; regression is performed on the Regions Proposal in the image marking the true target region to obtain the target region.
[0054] Exemplarily, when the input image is an image with text and the target region is the text region in the image, the result of using the Faster-RCNN neural network to extract the target region in the image is to obtain multiple text regions in the image.
[0055] In a possible implementation, the pixel values in the target region are statistically analyzed, and a pixel value distribution histogram is obtained according to the statistical result, including: statistically analyzing the pixels in the target region of the image, and then, according to the gray values of the pixels, the pixels are distinguished and the number of pixels under each distinguished gray level is statistically analyzed, and finally, the pixel value distribution histogram is generated in the form of a histogram. The value range of the gray level of the pixels is a settable value. Optionally, the range of the set gray level is from 0 to 255, that is, the dimension of the pixel value distribution histogram is 256 dimensions. The process of statistically analyzing the pixel values in the target region and obtaining the pixel value distribution histogram corresponding to the target region in the pixel distribution histogram. Figure 3 in the target region to the pixel distribution histogram.
[0056] Among them, the purpose of statistically analyzing the pixel values in the target region and obtaining the pixel value distribution histogram is that: when the difference features between the modified region and the unmodified region in the image are not obvious and unstable enough, obtaining the pixel distribution histogram based on the statistical method can well retain the difference features and reflect the difference features through the distribution features of the pixel values, and can enable the method for identifying the authenticity of images provided by the embodiments of the present application to identify the authenticity of images of different sizes. For example, for qualification certificates, the modified parts are concentrated in the text region, and the background of the qualification certificate is simple, and the boundary mutation information between the modified part and the unmodified part is not obvious enough. Therefore, the difference features between the modified part and the unmodified part are not obvious and unstable.
[0057] Exemplarily, when the target area is a text area in an image, the number of pixels in multiple text areas at each gray level is respectively counted, and the statistical results are output in the form of a histogram to obtain multiple pixel value distribution histograms corresponding to each text area. Optionally, when the gray value range is from 0 to 255, the dimension of the obtained multiple pixel value distribution histograms is 256.
[0058] Exemplarily, as Figure 5 shown, Figure 5 the histogram shown is the pixel value distribution histogram corresponding to a region block in the target area, where the gray value range of the pixels is from 0 to 255. As can be seen from Figure 5 it, there are a pixels with a gray level of 0, b pixels with a gray level of 1, c pixels with a gray level of 2, d pixels with a gray level of 3, and e pixels with a gray level of 4 in the region block corresponding to this histogram, where a, b, c, d, and e are all positive integers not less than 0. It should be noted that Figure 5 the histogram shown is only for helping to understand the embodiments of the present application and does not limit the embodiments of the present application.
[0059] In step 202, based on the pixel value distribution histogram, a pixel value distribution feature map is generated.
[0060] In a possible implementation manner, the process of obtaining the pixel value distribution feature map based on the pixel value distribution histogram includes steps 2021 to 2024.
[0061] Step 2021, calculate the average value of the number of pixels in each dimension of the pixel value distribution histogram.
[0062] In a possible implementation manner, according to the result of extracting the target area in the image in step 201, multiple region blocks are obtained, and multiple pixel value distribution histograms can be correspondingly obtained according to the multiple region blocks. Calculating the average value of the number of pixels in each dimension of the pixel value distribution histogram means calculating the average value of the number of pixels included in multiple region blocks at each gray level.
[0063] It should be noted that, in order to simply and clearly describe the process of calculating the average value of the number of pixels in each dimension of the pixel value distribution histogram, the embodiments of the present application will be described with the dimension of the pixel value distribution histogram being 5, and the 5 - dimensional pixel value distribution histogram is as Figure 6 shown.
[0064] Exemplarily, when the number of region blocks is n (n is an integer not less than 1) and the dimension of the pixel value distribution histogram is 5, n 5-dimensional pixel distribution histograms will be obtained according to step 201. Calculating the average value of the number of pixels in each dimension of the n 5-dimensional pixel value distribution histograms means calculating the average value of the sum of the number of pixels corresponding to the gray levels in the n region blocks when the gray levels are 0, 1, 2, 3, and 4. The average value of the number of pixels corresponding to the gray level 0 in the n 5-dimensional histograms is calculated according to the following formula 1:
[0065]
[0066] where a i is the number of pixels corresponding to the gray level 0 in the i-th 5-dimensional pixel distribution histogram, is the average value of the number of pixels corresponding to the gray level 0 in the n 5-dimensional histograms.
[0067] Correspondingly, the average value of the pixels corresponding to the gray level 1 in the n 5-dimensional pixel distribution histograms can be obtained the average value of the pixels corresponding to the gray level 2 the average value of the pixels corresponding to the gray level 3 the average value of the pixels corresponding to the gray level 4
[0068] Step 2022, based on this average value, calculate the covariance between each dimension in the pixel value distribution histogram.
[0069] In a possible implementation manner, based on this average value, calculating the covariance between each dimension in the pixel value distribution histogram includes: based on the average value of the number of pixels in each dimension in the multiple pixel value distribution histograms obtained in step 2021, calculate the covariance between each dimension in the multiple pixel value distribution histograms. Exemplarily, based on the average values obtained in step 2021 and calculate the covariance between each dimension of the n 5-dimensional pixel distribution histograms. Among them, the formula 2 for calculating the covariance is as follows:
[0070]
[0071] where X and Y are the arrays for which the covariance needs to be calculated, X i and Y i are the i-th parameters in X and Y respectively, and are the average values of the parameters included in X and Y, and Cov(X, Y) is the covariance between X and Y. It can be known from formula 2 that Cov(X, Y) = Cov(Y, X).
[0072] The covariance between gray level 0 and gray level 1 is Cov(a, b), and the calculation formula is as follows:
[0073]
[0074] Among them, a represents an array composed of the number of pixels corresponding to n 5D pixel distribution histograms at gray level 0. For example, the number of pixels of the first 5D pixel distribution histogram at gray level 0 is a1, and the number of pixels of the i-th 5D pixel distribution histogram at gray level 0 is a i ; b represents an array composed of the number of pixels corresponding to n 5D pixel distribution histograms at gray level 1; and correspond to the average values of the number of pixels of n 5D pixel value distribution histograms at gray levels 0 and 1.
[0075] Correspondingly, the covariance between gray level 0 and gray level 2 is Cov(a, c), where c represents an array composed of the number of pixels corresponding to n 5D pixel distribution histograms at gray level 2. The covariance between gray level 0 and gray level 3 is Cov(a, d), where d represents an array composed of the number of pixels corresponding to n 5D pixel distribution histograms at gray level 3. The covariance between gray level 0 and gray level 4 is Cov(a, e), where e represents an array composed of the number of pixels corresponding to n 5D pixel distribution histograms at gray level 4. The covariance between gray level 0 and gray level 0 is Cov(a, a).
[0076] Correspondingly, the covariances between gray level 1 and gray levels 0 to 4 are respectively: Cov(b, a), Cov(b, b), Cov(b, c), Cov(b, d), Cov(b, e). The covariances between gray level 2 and gray levels 0 to 4 are respectively: Cov(c, a), Cov(c, b), Cov(c, c), Cov(c, d), Cov(c, e). The covariances between gray level 3 and gray levels 0 to 4 are respectively: Cov(d, a), Cov(d, b), Cov(d, c), Cov(d, d), Cov(d, e). The covariances between gray level 4 and gray levels 0 to 4 are respectively: Cov(e, a), Cov(e, b), Cov(e, c), Cov(e, d), Cov(e, e).
[0077] Step 2023, based on this covariance, generate a covariance feature matrix, and this covariance feature matrix has the distribution characteristics of pixel values in the target area.
[0078] In a possible implementation manner, a covariance feature matrix is generated based on the covariance obtained in step 2022, including: arranging the covariance obtained in step 2022 according to a set arrangement rule to obtain the covariance feature matrix. Wherein, the set arrangement rule is any covariance arrangement rule that can reflect the distribution characteristics of pixel values, and the embodiments of the present application do not limit this. Exemplarily, the set arrangement rule may be that the first row arranges the covariances between the first dimension and other dimensions of the pixel distribution histogram in sequence, the second row arranges the covariances between the second dimension and other dimensions of the pixel distribution histogram in sequence, and then each subsequent row arranges the covariances between the corresponding dimension and other dimensions of the pixel distribution histogram in sequence. Alternatively, the set arrangement rule may be that the first column arranges the covariances between the first dimension and other dimensions of the pixel distribution histogram in sequence, the second column arranges the covariances between the second dimension and other dimensions of the pixel distribution histogram in sequence, and then each subsequent column arranges the covariances between the corresponding dimension and other dimensions of the pixel distribution histogram in sequence.
[0079] Among them, steps 2021 - 2023 correspond to Figure 3 the process from the pixel value distribution histogram to the covariance feature matrix.
[0080] In an exemplary embodiment, the covariance feature matrix generated according to the covariances between each dimension of the n 5 - dimensional pixel distribution histograms obtained in step 2022 is as Figure 7 shown. It should be noted that Figure 7 is only to help understand an arrangement manner of the covariance when the embodiments of the present application generate the covariance feature matrix, and does not limit the covariance feature matrix of the present application. Exemplarily, when the dimension of the pixel value distribution histogram is 256 - dimensional, that is, the range of grayscale values is from 0 to 255, a covariance feature matrix with a size of 256 * 256 can be generated according to the 256 - dimensional pixel value distribution histogram.
[0081] Optionally, the distribution characteristics of pixel values include the characteristics of the number of pixel values at each grayscale. Therefore, the pixel value distribution histogram obtained by counting the number of pixels in the target area at each grayscale has the distribution characteristics of pixel values. Since the number of region blocks obtained by extracting the target area from different images is different, and then the number of pixel value distribution histograms obtained by counting pixels for the region blocks is different, it is impossible to perform batch training based on the pixel value distribution histogram to identify the authenticity of images. Therefore, it is necessary to generate a covariance feature matrix based on the pixel distribution histogram, so that for different numbers of pixel value distribution histograms, a covariance feature matrix can be obtained, and then batch training can be achieved based on the covariance feature matrix to identify the authenticity of images. Because the pixel value distribution histogram has the distribution characteristics of pixel values, the covariance feature matrix obtained based on the pixel value distribution histogram also has the distribution characteristics of pixel values.
[0082] Step 2024: Obtain a pixel value distribution feature map based on the covariance feature matrix.
[0083] In a possible implementation, obtaining a pixel value distribution feature map based on the covariance feature matrix includes: extracting the pixel value distribution feature in the covariance feature matrix obtained in step 2023 using any neural network to obtain a pixel value distribution feature map. Optionally, the neural network for extracting the pixel value distribution feature in the covariance feature matrix is any ResNet (Residual Network) network. Exemplarily, the process of using a ResNet neural network to extract the pixel value distribution feature in the covariance feature matrix is to perform multiple convolutions, activations, and poolings on the covariance feature matrix. The number of times of performing convolutions, activations, and poolings depends on the type of the selected ResNet neural network. Obtaining a pixel value distribution feature map based on the covariance feature matrix corresponds to Figure 3 the process from the covariance feature matrix to the pixel value distribution feature map.
[0084] In step 203, classify the pixel value distribution feature map to obtain a classification result, and determine the authenticity result of the image based on the classification result.
[0085] In a possible implementation, classifying the pixel value distribution feature map to obtain a classification result includes: using a classification neural network to classify the pixel value distribution feature map to obtain the probabilities that the image corresponding to the pixel value distribution feature map is genuine and fake. Exemplarily, the classification neural network is a neural network including a normalized Softmax (an activation function) layer. Among them, the sum of the probability that the image is genuine and the probability that the image is fake is equal to 1. Exemplarily, the probability that the image is genuine is 0.8, and the probability that the image is fake is 0.2. Among them, classifying the pixel value distribution feature map to obtain a classification result corresponds to Figure 3 the process from the pixel value distribution feature map to the classification result of the image.
[0086] In a possible implementation, the process of determining the authenticity result of the image based on the classification result includes: obtaining the cross-entropy loss function value between the classification result and the annotation result of the image corresponding to the pixel value distribution feature map; optimizing the classification result based on the minimum cross-entropy loss function value to obtain an optimized result; and determining the authenticity result of the image corresponding to the pixel value distribution feature map based on the optimized result. Among them, determining the authenticity result of the image based on the classification result corresponds to Figure 3 the process from the classification result of the image to the authenticity of the image.
[0087] In a possible implementation, the annotation result refers to the annotation made in advance by humans on the authenticity of the image. Exemplarily, for an unmodified image, the manually annotated result is that the probability that the image is genuine is 1, and the probability that the image is fake is 0. For an image with modifications in the target area, the manually annotated result is that the probability that the image is genuine is 0, and the probability that the image is fake is 1. The value of the cross-entropy loss function between the classification result and the annotation result represents the degree of difference between the classification result and the annotation result in terms of probability distribution. The smaller the value of the cross-entropy, the smaller the degree of difference, and the more accurate the corresponding classification result.
[0088] The annotation result is a fixed value and does not change. Therefore, when minimizing the value of the cross-entropy loss function between the classification result and the annotation result, the probability distribution corresponding to the classification result approaches the probability distribution corresponding to the annotation result, thereby completing the optimization of the classification result. Exemplarily, the classification result of the image is that the probability that the image is genuine is 0.78, and the probability that the image is fake is 0.22. When any one of the probabilities corresponding to the classification result exceeds the probability threshold of 0.8, the authenticity result of the image can be determined. Therefore, based on this classification result, the authenticity result of the image cannot be determined. Thus, it is necessary to optimize this classification result so that one of the probabilities in this classification result exceeds 0.8. If there is a probability in the classification result that exceeds the probability threshold, the authenticity result of the image can be directly determined based on this classification result.
[0089] Optionally, determining the authenticity result of the image corresponding to the pixel value distribution feature map based on the optimization result means determining the authenticity result of the image based on the probability corresponding to the optimization result obtained by optimizing the classification result. Exemplarily, the set probability threshold is 0.8, and the classification result is that the probability that the image is genuine is 0.78, and the probability that the image is fake is 0.22. At this time, since none of the probabilities corresponding to the classification result exceed 0.8, the authenticity of the image cannot be determined based on the classification result. After optimizing this classification result, the obtained optimization result is that the probability that the image is genuine is 0.93, and the probability that the image is fake is 0.07. At this time, it can be determined that the image is genuine.
[0090] In another exemplary embodiment, based on the method for identifying the authenticity of an image provided in the embodiments of the present application, the accuracy of identifying the authenticity of an image on a self-built dataset (1000 synthetic PS (Photoshop, an image processing software) + 929 real data) is 98.02%.
[0091] In the embodiments of the present application, a pixel value distribution histogram is obtained by counting the pixel values of the target region extracted from the image. The pixel value distribution histogram can better retain the difference features between the modified region and the unmodified region in the image and represent the difference features through the distribution features of the pixel values. Obtaining a pixel value distribution feature map based on the pixel value distribution histogram can make the distribution features of the pixel values more stable. Then, classifying the pixel value distribution features and determining the authenticity of the image based on the classification results can avoid the low accuracy of authenticity identification caused by the same noise features in the modified region and the unmodified region of the image.
[0092] See Figure 8 , embodiments of the present application provide an apparatus for identifying the authenticity of an image. The apparatus includes:
[0093] An acquisition module 801, configured to acquire an image to be identified;
[0094] An extraction module 802, configured to extract the target region in the image, count the pixel values in the target region, and obtain a pixel value distribution histogram according to the statistical results;
[0095] A generation module 803, configured to generate a pixel value distribution feature map based on the pixel value distribution histogram;
[0096] A classification module 804, configured to classify the pixel value distribution feature map to obtain a classification result, and determine the authenticity result of the image based on the classification result.
[0097] In a possible implementation manner, the generation module 803 is configured to calculate the average value of the number of pixels in each dimension of the pixel value distribution histogram; based on the average value, calculate the covariance between each dimension in the pixel value distribution histogram; based on the covariance, generate a covariance feature matrix, and the covariance feature matrix has the distribution features of the pixel values in the target region; obtain a pixel value distribution feature map based on the covariance feature matrix.
[0098] In a possible implementation manner, the classification module 804 is configured to obtain the cross-entropy loss function value between the classification result and the annotation result of the image; optimize the classification result based on the minimum cross-entropy loss function value to obtain an optimization result; determine the authenticity result of the image based on the optimization result.
[0099] In a possible implementation manner, the extraction module 801 is configured to use the Faster Region Convolutional Neural Network (Faster-RCNN) to extract the target region in the image; or use a Single Shot MultiBox Detector (SSD) of a single deep neural network to extract the target region in the image.
[0100] In a possible implementation manner, a generation module 803 is configured to extract distribution features in a covariance feature matrix based on any residual neural network ResNet, and obtain a pixel value distribution feature map.
[0101] In a possible implementation manner, the target area is an area in the image that contains text.
[0102] In the embodiments of the present application, a pixel value distribution histogram is obtained by counting the pixel values of the target area extracted from the image. The pixel value distribution histogram can better retain the difference features between the modified area and the unmodified area in the image and represent the difference features through the distribution features of the pixel values. Obtaining a pixel value distribution feature map based on the pixel value distribution histogram can make the distribution features of the pixel values more stable. Subsequently, classifying the pixel value distribution features and determining the authenticity of the image based on the classification results can avoid the low accuracy of authenticity identification caused by the same noise features in the modified area and the unmodified area of the image.
[0103] It should be noted that when the device provided in the above embodiments implements its functions, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0104] Figure 9 FIG. is a schematic structural diagram of a server provided by an embodiment of the present application. The server may vary greatly due to configuration or performance differences, and may include one or more processors (Central Processing Units, CPUs) 901 and one or more memories 902. Among them, at least one computer program is stored in the one or more memories 902, and the at least one computer program is loaded and executed by the one or more processors 901 to enable the server to implement the method for identifying the authenticity of an image provided by each of the above method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The server may also include other components for implementing the functions of the device, which will not be elaborated here.
[0105] Figure 10It is a schematic structural diagram of a terminal provided by an embodiment of the present application. The terminal may be: a smart phone, a tablet computer, an MP3 (Moving Picture Experts Group Audio Layer III) player, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer or a desktop computer. The terminal may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0106] Generally, the terminal includes: a processor 1501 and a memory 1502.
[0107] The processor 1501 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1501 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor 1501 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1501 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1501 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0108] The memory 1502 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1502 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1502 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1501 so that the terminal implements the method for identifying the authenticity of images provided by the method embodiments in the present application.
[0109] In some embodiments, the terminal may further optionally include: a peripheral device interface 1503 and at least one peripheral device. The processor 1501, the memory 1502, and the peripheral device interface 1503 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1503 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 1504, a display screen 1505, a camera assembly 1506, an audio circuit 1507, a positioning assembly 1508, and a power supply 1509.
[0110] The peripheral device interface 1503 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1501 and the memory 1502. In some embodiments, the processor 1501, the memory 1502, and the peripheral device interface 1503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1501, the memory 1502, and the peripheral device interface 1503 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0111] The radio frequency circuit 1504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1504 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1504 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 1504 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1504 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0112] The display screen 1505 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1505 is a touch display screen, the display screen 1505 also has the ability to collect touch signals on or above the surface of the display screen 1505. The touch signals can be input to the processor 1501 as control signals for processing. At this time, the display screen 1505 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1505, which is provided on the front panel of the terminal; in other embodiments, there may be at least two display screens 1505, which are respectively provided on different surfaces of the terminal or in a folding design; in other embodiments, the display screen 1505 may be a flexible display screen, which is provided on a curved surface or a folding surface of the terminal. Even, the display screen 1505 can also be set to an irregular non-rectangular shape, that is, an irregular-shaped screen. The display screen 1505 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0113] The camera module 1506 is used to capture images or videos. Optionally, the camera module 1506 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, to implement functions such as background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting functions or other fused shooting functions. In some embodiments, the camera module 1506 may also include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0114] The audio circuit 1507 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1501 for processing, or input to the radio frequency circuit 1504 to implement voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 1501 or the radio frequency circuit 1504 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1507 may further include a headphone jack.
[0115] The positioning component 1508 is used to locate the current geographical location of the terminal to implement navigation or LBS (Location Based Service). The positioning component 1508 may be a positioning component based on the GPS (Global Positioning System) of the United States, the Beidou system of China, the GLONASS system of Russia, or the Galileo system of the European Union.
[0116] The power supply 1509 is used to supply power to each component in the terminal. The power supply 1509 may be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 1509 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0117] In some embodiments, the terminal further includes one or more sensors 1510. The one or more sensors 1510 include, but are not limited to: an acceleration sensor 1511, a gyroscope sensor 1512, a pressure sensor 1513, a fingerprint sensor 1514, an optical sensor 1515, and a proximity sensor 1516.
[0118] The acceleration sensor 1511 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the terminal. For example, the acceleration sensor 1511 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1501 can control the display screen 1505 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1511. The acceleration sensor 1511 can also be used for collecting game or user movement data.
[0119] The gyroscope sensor 1512 can detect the body direction and rotation angle of the terminal. The gyroscope sensor 1512 can cooperate with the acceleration sensor 1511 to collect the 3D actions of the user on the terminal. Based on the data collected by the gyroscope sensor 1512, the processor 1501 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0120] The pressure sensor 1513 can be disposed on the side frame of the terminal and / or the lower layer of the display screen 1505. When the pressure sensor 1513 is disposed on the side frame of the terminal, it can detect the holding signal of the user on the terminal, and the processor 1501 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 1513. When the pressure sensor 1513 is disposed on the lower layer of the display screen 1505, the processor 1501 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1505. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0121] The fingerprint sensor 1514 is used to collect the fingerprints of the user. The processor 1501 can identify the user's identity according to the fingerprints collected by the fingerprint sensor 1514, or the fingerprint sensor 1514 can identify the user's identity according to the collected fingerprints. When the identity of the user is identified as a trusted identity, the processor 1501 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 1514 can be disposed on the front, back, or side of the terminal. When there are physical buttons or a manufacturer's logo (trademark) on the terminal, the fingerprint sensor 1514 can be integrated with the physical buttons or the manufacturer's logo.
[0122] The optical sensor 1515 is used to collect the ambient light intensity. In one embodiment, the processor 1501 can control the display brightness of the display screen 1505 according to the ambient light intensity collected by the optical sensor 1515. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1505 is increased; when the ambient light intensity is low, the display brightness of the display screen 1505 is decreased. In another embodiment, the processor 1501 can also dynamically adjust the shooting parameters of the camera module 1506 according to the ambient light intensity collected by the optical sensor 1515.
[0123] The proximity sensor 1516, also known as the distance sensor, is usually disposed on the front panel of the terminal. The proximity sensor 1516 is used to collect the distance between the user and the front of the terminal. In one embodiment, when the proximity sensor 1516 detects that the distance between the user and the front of the terminal is gradually decreasing, the processor 1501 controls the display screen 1505 to switch from the lit state to the off state; when the proximity sensor 1516 detects that the distance between the user and the front of the terminal is gradually increasing, the processor 1501 controls the display screen 1505 to switch from the off state to the lit state.
[0124] Those skilled in the art can understand that Figure 10 the structure shown in
[0125] In an exemplary embodiment, a computer device is further provided, which includes a processor and a memory, and at least one computer program is stored in the memory. The at least one computer program is loaded and executed by one or more processors, so that the computer device implements any one of the above methods for authenticating the authenticity of an image.
[0126] In an exemplary embodiment, a computer-readable storage medium is further provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by the processor of the computer device, so that the computer implements any one of the above methods for authenticating the authenticity of an image.
[0127] In a possible implementation manner, the above computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0128] In an exemplary embodiment, a computer program product is further provided, which includes a computer program or computer instructions. The computer program or the computer instructions are loaded and executed by the processor, so that the computer implements any one of the above methods for authenticating the authenticity of an image.
[0129] It should be noted that the terms "first", "second", etc. (if any) in the specification and claims of this application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0130] It should be understood that the term "plurality" as referred to herein means two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0131] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for identifying the authenticity of an image, characterized in that, The method includes: Obtaining an image to be authenticated; Extracting a target region from the image, counting pixel values in the target region, and obtaining a pixel value distribution histogram according to the statistical result. The pixel value distribution histogram includes multiple dimensions and the number of pixels in each dimension. Each dimension is used to indicate a gray value, and the number of pixels in each dimension is the number of pixels under the corresponding gray value; Generating a pixel value distribution feature map based on the pixel value distribution histogram; Classifying the pixel value distribution feature map to obtain a classification result, and determining the authenticity result of the image based on the classification result; The generating a pixel value distribution feature map based on the pixel value distribution histogram includes: Calculating the average value of the number of pixels in each dimension of the pixel value distribution histogram; Calculating the covariance between each dimension in the pixel value distribution histogram based on the average value; Generating a covariance feature matrix based on the covariance. The covariance feature matrix has the distribution characteristics of pixel values in the target region; Obtaining the pixel value distribution feature map based on the covariance feature matrix.
2. The method according to claim 1, characterized in that, The determining the authenticity result of the image based on the classification result includes: Obtaining the cross-entropy loss function value between the classification result and the annotation result of the image; Optimizing the classification result based on the minimum cross-entropy loss function value to obtain an optimization result; Determining the authenticity result of the image based on the optimization result.
3. The method according to claim 1, characterized in that, The extracting the target region from the image includes: Using the Faster Region-based Convolutional Neural Network (Faster-RCNN) to extract the target region from the image; Or, using the Single Shot MultiBox Detector (SSD), a single deep neural network, to extract the target region from the image.
4. The method according to claim 1, wherein The obtaining the pixel value distribution feature map based on the covariance feature matrix includes: Extracting the distribution characteristics in the covariance feature matrix based on any Residual Neural Network (ResNet) to obtain the pixel value distribution feature map.
5. The method according to any one of claims 1-4, characterized in that The target region is the region in the image that contains text.
6. An apparatus for authenticating the authenticity of an image, characterized in that, The device includes: An acquisition module, configured to obtain an image to be authenticated; An extraction module, configured to extract a target region from the image, count pixel values in the target region, and obtain a pixel value distribution histogram according to the statistical result. The pixel value distribution histogram includes multiple dimensions and the number of pixels in each dimension. Each dimension is used to indicate a gray value, and the number of pixels in each dimension is the number of pixels under the corresponding gray value; A generation module, configured to generate a pixel value distribution feature map based on the pixel value distribution histogram. The generating a pixel value distribution feature map based on the pixel value distribution histogram includes: Calculating the average value of the number of pixels in each dimension of the pixel value distribution histogram; Calculating the covariance between each dimension in the pixel value distribution histogram based on the average value; Generating a covariance feature matrix based on the covariance. The covariance feature matrix has the distribution characteristics of pixel values in the target region; Obtaining the pixel value distribution feature map based on the covariance feature matrix; A classification module, configured to classify the pixel value distribution feature map to obtain a classification result, and determine the authenticity result of the image based on the classification result.
7. A computer device, characterized in that, The computer device includes a processor and a memory. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor so that the computer device implements the method for authenticating the authenticity of an image according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor so that a computer implements the method for authenticating the authenticity of an image according to any one of claims 1 to 5.
9. A computer program product, characterized in that, The computer program product includes a computer program or computer instructions, and the computer program or the computer instructions are loaded and executed by a processor so that a computer implements the method for authenticating the authenticity of an image according to any one of claims 1 to 5.
Citation Information
Patent Citations
In vivo detection and face recognition method combining main-side view
CN109543521A
Authenticity identification method and device and electronic equipment
CN112115921A