Image processing method and device and electronic equipment

By sliding cropping and feature similarity analysis on the image, the target image block where the image content has been tampered with is determined, which solves the problem of difficult to identify the tampered text area in the prior art, and realizes high-accurate image content tampering detection.

CN120219935APending Publication Date: 2025-06-27LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510361369.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-27

Smart Images

  • Figure CN120219935A_ABST
    Figure CN120219935A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and apparatus, and an electronic device. The method comprises the steps of obtaining a first image; the first image is cut in a sliding mode through a first cutting window of at least one size, at least one image block set of the first image is obtained, and each image block set of the first image comprises at least one image block obtained by cutting the first image based on the first cutting window of the same size; for each image block set, a feature similarity set corresponding to each image block in the image block set is determined, and the feature similarity set of the image blocks comprises feature similarities of the image blocks and other image blocks in the image block set; and determining a target image block of which the image content is tampered based on the feature similarity set corresponding to each image block in the image block set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular, to an image processing method, apparatus, and electronic device. Background Art

[0002] With the continuous development of technology, there are often situations where the text or other content in an image is tampered with. In particular, it is very difficult to visually distinguish the tampered text in the image. Therefore, how to locate the tampered areas such as text in the image is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention

[0003] On the one hand, this application provides an image processing method, including:

[0004] Obtain a first image;

[0005] Use at least one size of first cropping window to slide-crop the first image to obtain at least one set of image blocks of the first image. Each set of image blocks of the first image includes: at least one image block cropped from the first image based on the first cropping window of the same size;

[0006] For each set of image blocks, determine the set of feature similarities corresponding to each image block in the set of image blocks. The set of feature similarities of the image block includes the feature similarities between the image block and each other image block in the set of image blocks;

[0007] Based on the sets of feature similarities corresponding to the respective image blocks in the set of image blocks, determine the target image blocks with tampered image content.

[0008] In a possible implementation manner, the determining the target image blocks with tampered image content based on the sets of feature similarities corresponding to the respective image blocks in the set of image blocks includes:

[0009] For each image block in each set of image blocks, based on the respective feature similarities in the set of feature similarities of the image block, determine the confidence corresponding to the set of feature similarities of the image block;

[0010] Determine the target image blocks whose corresponding confidences meet the requirements from each set of image blocks to obtain the target image blocks with tampered image content.

[0011] In another possible implementation manner, it further includes:

[0012] Based on the first image, generate at least one second image. The resolution of the second image is different from that of the first image, and the resolutions of different second images are different;

[0013] For each second image, perform sliding cropping on the second image using a second cropping window of at least one size to obtain at least one set of image patches of the second image. Each set of image patches of the second image includes: at least one image patch cropped from the second image based on a second cropping window of the same size.

[0014] In yet another possible implementation, it further includes:

[0015] Determine a target image region in the first image corresponding to the target image patch;

[0016] Determine the target image region as the region in the first image where the image content is tampered with.

[0017] In yet another possible implementation, generating at least one second image based on the first image includes at least one of the following:

[0018] Perform at least one upsampling on the first image to obtain at least one second image;

[0019] Perform at least one downsampling on the first image to obtain at least one second image.

[0020] In yet another possible implementation, determining the target image region in the first image corresponding to the target image patch includes:

[0021] If the target image patch is from a second image, determine the target image region in the first image corresponding to the target image patch based on the scaling ratio between the first image and the second image to which the target image patch belongs.

[0022] In yet another possible implementation, obtaining the first image includes:

[0023] Obtain an image to be detected;

[0024] Convert the image to be detected into a grayscale image;

[0025] Normalize the pixel values of each pixel point in the grayscale image to values between 0 and 1 to obtain the first image.

[0026] In yet another possible implementation, the sizes of the second cropping windows corresponding to different second images are not completely the same, and the size of the second cropping window is different from the size of the first cropping window.

[0027] In another aspect, the present application further provides an image processing apparatus, including:

[0028] An image acquisition unit for acquiring a first image;

[0029] A first cropping unit, configured to perform sliding cropping on the first image by using a first cropping window of at least one size, so as to obtain at least one set of image blocks of the first image, where each set of image blocks of the first image includes: at least one image block cropped from the first image based on the first cropping window of the same size;

[0030] A feature determination unit, configured to, for each set of image blocks, determine a set of feature similarities corresponding to each image block in the set of image blocks, where the set of feature similarities of the image block includes the feature similarities between the image block and each of the other image blocks in the set of image blocks;

[0031] A tampering recognition unit, configured to determine a target image block whose image content is tampered with based on the sets of feature similarities corresponding to the image blocks in the set of image blocks.

[0032] On the other hand, the present application further provides an electronic device, including a memory and a processor;

[0033] The processor is configured to execute the image processing method described in any one of the above;

[0034] The memory is configured to store a program required for the processor to perform operations. Description of the Drawings

[0035] In combination with the drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original elements and elements are not necessarily drawn to scale.

[0036] Figure 1 It is a schematic flowchart of an image processing method provided by the present application;

[0037] Figure 2 It is another schematic flowchart of an image processing method provided by the present application;

[0038] Figure 3 It is another schematic flowchart of an image processing method provided by the present application;

[0039] Figure 4 It is an example diagram of an implementation framework in an application example of an image processing method provided by the present application;

[0040] Figure 5 It is a schematic diagram of a composition structure of an image processing device provided by the present application;

[0041] Figure 6 It is another schematic diagram of a composition structure of an image processing device provided by the present application;

[0042] Figure 7 It is a schematic diagram of a composition architecture of the electronic device provided by this application. Specific embodiments

[0043] The embodiments of this application will be described below in conjunction with the accompanying drawings in the embodiments of this application. The terms used in the embodiments part of this application are only used to explain the specific embodiments of this application, rather than intended to limit this application. Those of ordinary skill in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.

[0044] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinction used when describing objects with the same attributes in the embodiments of this application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0045] As Figure 1 , it shows a schematic flowchart of an image processing method provided by this application. The method of this embodiment can be applied to an electronic device, and the electronic device can be a laptop computer, a desktop computer, a tablet computer, etc., or a device node in a server or a cloud platform, etc., without limitation.

[0046] The method of this embodiment may include the following steps S101 to S104:

[0047] S101, obtain a first image.

[0048] In this application, the first image is an image for which it is necessary to detect whether there is image content tampering.

[0049] Among them, the first image can be an image containing text content, an image including only non-text content, or an image including both text content and non-text content, without limitation.

[0050] S102, perform sliding cropping on the first image using at least one size of first cropping window to obtain at least one set of image blocks of the first image.

[0051] Among them, different sets of image patches of the first image correspond to first cropping windows of different sizes. Correspondingly, each set of image patches of the first image includes: at least one image patch cropped from the first image based on a first cropping window of the same size.

[0052] In an alternative manner, in order to be able to focus on features in different regions of the first image and further improve the accuracy of locating the tampered region of the image content, the present application can also slide-crop the first image respectively using a variety of first cropping windows of different sizes, so as to obtain a variety of sets of image patches corresponding to the first image. Since different sets of image patches of the first image correspond to first cropping windows of different sizes, the sizes of the image patches in different sets of image patches of the first image are different.

[0053] For example, the first cropping windows of multiple sizes may include: a first cropping window with a size of , a first cropping window with a size of , and a first cropping window with a size of , etc. Among them, the sizes of the image patches in the set of image patches collected using the first cropping window of are all , while the sizes of the image patches collected using the first cropping window of are .

[0054] In particular, for any first cropping window of a certain size, during the process of sliding-cropping the first image using the first cropping window, there is an overlapping image area between two adjacent cropped image patches, so as to reduce the situation where the tampered image area in the first image cannot be selected by the image patches, that is, reduce the situation where an image patch cannot cover a certain tampered area in the first image.

[0055] For example, for each first cropping window of a certain size, the first image can be slide-cropped using the first cropping window according to the sliding step corresponding to the first cropping window of this size. Among them, the sliding step is smaller than the length of the first cropping window, so that there is an overlapping area between two adjacent slid-acquired image patches.

[0056] S103. For each set of image patches, determine the set of feature similarities corresponding to each image patch in the set of image patches.

[0057] For any image patch in a set of image patches, the set of feature similarities of the image patch includes the feature similarities between the image patch and other image patches in the set of image patches respectively.

[0058] For example, for each set of image patches, the set of feature similarities corresponding to each image patch can be a feature similarity matrix, where different elements in the feature similarity matrix represent the similarities between the image patch and different image patches in the set of image patches. Of course, the set of feature similarities can also have other forms, and there is no limitation on this.

[0059] Among them, there can be multiple possibilities for the specific implementation of determining the feature similarity between an image patch and other image patches in the set of image patches, and there is no specific limitation.

[0060] For example, in a possible implementation manner, for each set of image patches, a feature extraction model can be used to determine the features of each image patch in the set of image patches. For example, a Transformer model can be used to extract the features of each image patch. Then, based on the features of each image patch, the set of feature similarities corresponding to each image patch is determined.

[0061] For example, for each image patch in each set of image patches, the cosine similarity or Pearson correlation coefficient, etc., between the features of the image patch and the features of other image patches in the set of image patches can be calculated, and the calculated cosine similarity or Pearson correlation coefficient, etc., is used as the feature similarity between the image patch and other image patches.

[0062] S104. Based on the sets of feature similarities corresponding to each image patch in the set of image patches, determine the target image patch with tampered image content.

[0063] It can be understood that if there is tampering in the text or other content of the image patch, then there will also be changes in the features of the image patch, resulting in differences in the features presented by the image patch and the features of other image patches. On this basis, since the set of feature similarities corresponding to the image patch can reflect the similarity in features between the image patch and other image patches in the set of image patches, therefore, by combining the sets of feature similarities corresponding to each image patch, an image patch with a large difference in features from other image patches can be determined, thereby determining the target image patch with tampered image content.

[0064] It can be understood that after determining the target image patch, since the target image patch belongs to the first image, the target image region corresponding to the target image patch in the first image is actually the region in the first image where the image content is tampered.

[0065] It should be noted that in this application, according to actual needs, it may be possible to determine a target image block with the highest possibility of image content being tampered with. Considering that there may be text or patterns in multiple regions of the first image being tampered with, this application may also determine at least one target image block with a risk of image content being tampered with, without specific limitations. Of course, if there is no tampering with the image content in the first image, it may also be detected that there is no target image block with tampered image content, and in this case, the target image block can be empty.

[0066] As can be seen from the above, by sliding and cropping the first image with at least one size of the first cropping window, the first image can be refined into image blocks corresponding to at least one size. On this basis, for each set of image blocks composed of image blocks of each size, this application will determine the set of feature similarities corresponding to each image block in the set of image blocks. Since the set of feature similarities of an image block includes the feature similarities between the image block and each other image block in the set of image blocks where it is located, and if the image content in the image block belongs to the tampered content, then the set of feature similarities corresponding to the image block will necessarily have a large difference from the sets of feature similarities corresponding to other image blocks, so that the image block with tampered image content can be determined, and thus the tampered area in the first image is obtained.

[0067] Moreover, since this application refines the first image into image blocks of at least one size, it realizes local feature analysis and processing of different regions of the first image from at least one fine-grained level, which is conducive to more accurately determining the tampered image area in the first image.

[0068] In this application, there are multiple possible specific implementations for determining the target image block with tampered image content based on the sets of feature similarities corresponding to each image block.

[0069] For example, based on the sets of image similarities corresponding to each image block in each set of image blocks, using a trained recognition model, the target image block with tampered image content can be determined from each set of image blocks. Among them, the recognition model can be a deep learning model such as a neural network or other types of artificial intelligence models, without specific limitations.

[0070] In a possible implementation manner, in order to reduce the resources consumed for training the model and reduce the complexity of identifying the target image block with tampered image content, for each image block in each set of image blocks, this application determines the confidence corresponding to the set of feature similarities based on each feature similarity in the set of feature similarities of the image block, so as to obtain the confidence corresponding to the image block. On this basis, this application can determine the target image block corresponding to the confidence that meets the requirements from each set of image blocks, and obtain the target image block with tampered image content.

[0071] Among them, there are various specific ways to calculate the confidence level based on the feature similarities in the set of feature similarities of image patches. For example, the confidence level can be determined by combining the average value of the feature similarities corresponding to each image patch; or the confidence levels of the feature similarities corresponding to the image patches can be determined by combining the information entropy method, etc., without specific limitations.

[0072] It can be understood that the confidence level corresponding to the set of feature similarities of image patches can reflect the degree of difference in the feature distributions between the image patches and other image patches. The smaller the confidence level, the greater the pixel feature difference between the image patch and other image patches, and the greater the possibility that the image patch has been tampered with. Based on this, the confidence level corresponding to the image patch can characterize the risk degree of the image content of the image patch being tampered with.

[0073] Based on this, the present application can screen out the target image patches whose image content has been tampered with by reasonably setting the requirements that the confidence level needs to meet.

[0074] For example, the target image patches with the confidence level meeting the requirements can be the target image patches with the confidence level lower than the set threshold; it can also be the target image patches with the confidence level lower than the set threshold and the smallest confidence level, without specific limitations.

[0075] It can be understood that in the case where the content tampered with in the image is less or the content tampered with is text, the visual difference before and after the image is tampered with is small, resulting in a relatively weak feature difference between the tampered area in the image before and after tampering. Based on this, in order to more accurately identify the areas such as the text areas tampered with in the image, the present application can also generate at least one second image based on the first image, where the resolution of the second image is different from that of the first image, and the resolutions of different second images are different.

[0076] Correspondingly, for each second image, the present application can also use at least one size of second cropping window to slide-crop the second image to obtain at least one set of image patches of the second image. Each set of image patches of the second image includes: at least one image patch cropped from the second image based on the second cropping window of the same size. For each second image, the sizes of the second cropping windows corresponding to different sets of image patches of the second image are different.

[0077] In the present application, only for the convenience of distinction, the window used to crop the first image is called the first cropping window, and the window used to crop the second image is called the second cropping window. Among them, different sets of image patches of the same second image correspond to second cropping windows of different sizes.

[0078] On this basis, for each set of image patches corresponding to the first image and each second image, the present application will determine the set of feature similarities corresponding to each image patch in the set of image patches. Correspondingly, based on the feature similarities corresponding to each image patch in each set of image patches, the present application can determine the target image patches with tampered image content from the sets of image patches corresponding to the first image and each second image. The specific implementation of determining the target image patches can be referred to the relevant introduction above and will not be elaborated here.

[0079] It can be understood that the first image and multiple second images are multiple images with the same presented content but different scales. By using multiple images with different scales, the feature difference performance of the area where the image content in the first image is modified can be enhanced. On this basis, while slidingly cropping the first image with at least one size of the first cropping window, slidingly cropping the second image with at least one size of the second cropping window can achieve different-granularity segmentation and feature analysis of multiple images with different scales, so as to be able to achieve more fine-grained feature extraction and analysis, which is beneficial to more accurately detecting the image areas with different features, and then more accurately determining the image patches corresponding to the areas where the image content is tampered, and naturally can more accurately determine the tampered image areas.

[0080] Furthermore, considering the different resolutions of the first image and the second image, in order to more reasonably segment the image patches of the first image and the second image, in the present application, the size of the second cropping window corresponding to the second image is different from the size of the first cropping window corresponding to the first image. Correspondingly, considering that the resolutions of the second images are also different, therefore, the sizes of the second cropping windows corresponding to different second images may not be exactly the same. For example, some of the at least one size corresponding to the second cropping window corresponding to different second images may be the same, or all the sizes may be different.

[0081] Among them, there are multiple possible implementations for generating the second image based on the first image. For example, at least one second image can be obtained by performing at least one upsampling on the first image; it can also be that at least one second image is obtained by performing at least one downsampling on the first image. Of course, in practical applications, at least one upsampling and at least one downsampling can be respectively performed on the first image to obtain at least one second image with a resolution lower than that of the first image and at least one second image with a resolution higher than that of the first image.

[0082] It can be understood that since the target image block may be derived from the first image or the second image, in order to determine the tampered image area in the first image, after determining the target image block, the present application can also determine the target image area in the first image corresponding to the target image block. Correspondingly, the target image area is determined as the area in the first image where the image content is tampered with.

[0083] For ease of understanding, the specific implementation of determining the tampered image area in the first image based on the image blocks segmented from the first image and the second image is described below in conjunction with Figure 2 an implementation shown below.

[0084] As Figure 2 shown, another schematic flowchart of the image processing method provided by the present application is shown. The method of this embodiment may include:

[0085] S201, obtain the first image.

[0086] S202, perform at least one upsampling and at least one downsampling on the first image respectively to obtain a plurality of second images.

[0087] Among them, upsampling the first image can increase the resolution of the first image, so that the resolution of the obtained second image is greater than that of the first image. And downsampling the first image can reduce the resolution of the first image, so that the resolution of the obtained second image is less than that of the first image. Therefore, among the plurality of second images obtained in the present application, the resolution of some second images is lower than that of the first image, while the resolution of another part of the second images is higher than that of the first image.

[0088] For example, assuming the first image is image, then the first image can be downsampled to a second image with a resolution of and second image. In addition, the first image can also be upsampled to second image.

[0089] Among them, the method of downsampling the first image can be average pooling or max pooling, etc., and can also be other downsampling methods, which are not limited thereto. The method of upsampling the first image can be bilinear interpolation, nearest neighbor interpolation, etc., and is not specifically limited.

[0090] In this embodiment, the example of performing upsampling and downsampling on the first image respectively is used for illustration. However, in actual applications, only upsampling or downsampling can be performed on the first image, which is not specifically limited.

[0091] S203. Use a first cropping window of at least one size to perform sliding cropping on the first image, obtaining at least one set of image patches of the first image.

[0092] Among them, each set of image patches of the first image includes: at least one image patch cropped from the first image based on a first cropping window of the same size.

[0093] This step can refer to the relevant introduction in the previous embodiments and will not be elaborated here.

[0094] It should be noted that the order of steps S202 and S203 can be interchanged or they can be executed synchronously, without specific restrictions.

[0095] S204. For each second image, use a second cropping window of at least one size to perform sliding cropping on the second image, obtaining at least one set of image patches of the second image.

[0096] Among them, for any one second image, different sets of image patches of the second image correspond to different sizes of second cropping windows. Therefore, the sizes of the image patches in different sets of image patches of the same second image are different. Each set of image patches of the second image includes: at least one image patch cropped from the second image based on a second cropping window of the same size.

[0097] In this application, considering that the resolutions of different second images are different and the sizes of different second images are also different, in order to more reasonably divide the image patches of the second image, the sizes of the second cropping windows corresponding to different second images in this application are not completely the same.

[0098] For example, if the second image is of an image, then sliding cropping can be performed on the second image respectively using of a second cropping window, obtaining a set of image patches composed of multiple sized image patches; and sliding cropping can also be performed on the second image using of a second cropping window, obtaining a set of image patches composed of multiple sized image patches. And for of a second image, sliding cropping can be performed on the second image respectively using of a second cropping window and of a second cropping window, thereby respectively obtaining a set of image patches including multiple sized image patches and a set of image patches including multiple sized image patches.

[0099] Similarly, considering the different resolutions of the first image and the second image, for any second image, the size of the second cropping window corresponding to the second image can also be different from the sizes of the first cropping windows.

[0100] S205. For any set of image patches corresponding to the first image and the second image, determine the set of feature similarities corresponding to each image patch in the set of image patches.

[0101] Wherein, the set of feature similarities of an image patch includes the feature similarities between the image patch and other image patches in the set of image patches respectively.

[0102] In this application, for each set of image patches determined from the first image and the second image, it is necessary to determine the set of feature similarities corresponding to each image patch in the set of image patches. For the specific implementation of determining the set of feature similarities corresponding to an image patch, reference can be made to the relevant introduction in the previous embodiments, which will not be elaborated here.

[0103] S206. Based on the sets of feature similarities corresponding to the image patches in each set of image patches, determine the target image patches with tampered image content.

[0104] It can be understood that for the specific implementation of determining the target image patches from multiple sets of image patches, reference can be made to the relevant introduction in the previous embodiments, which will not be elaborated here. Similar to the previous embodiments, the target image patches can be one or more. If there is no tampering with text or other content in the first image, the target image patches can also be empty.

[0105] S207. Determine the target image region in the first image corresponding to the target image patches, and determine the target image region as the region in the first image where the image content is tampered.

[0106] For example, if the target image patches are from the first image, the region of the target image patches in the first image can be determined as the target image region.

[0107] If the target image patches are from the second image, considering the different scales of the first image and the second image, this application can also determine the target image region in the first image corresponding to the target image patches based on the scaling ratio between the first image and the second image to which the target image patches belong.

[0108] Wherein, for the second image to which the target image patches belong, the scaling ratio between the second image and the second image is related to the resolutions of the first image and the second image respectively. For example, if the first image is image, and the second image is image, then the scaling ratio between the first image and the second image is 2:1.

[0109] For example, after determining the image block area of the target image block in the second image, the second image is scaled based on the scaling ratio, the target coordinate area corresponding to the image block area in the scaled second image in the first image is determined, and the target image area within the target coordinate area in the first image is determined as the image area where the image content is tampered with.

[0110] It can be understood that there are relatively many scenarios where the text in the image is tampered with. Moreover, in practical applications, more attention is paid to the situation where the text in the image is tampered with. The text in the image is mainly black and white, and few texts have colors, and the color information of the text is not important, which will also increase the amount of computation required to identify whether the image is tampered with. Based on this, in this application, the first image can be a grayscale image converted from the image to be detected.

[0111] Furthermore, in order to reduce the amount of data processing, this application can also normalize the pixel values of each pixel point in the grayscale image to values between 0 and 1, and use the normalized grayscale image as the first image.

[0112] The following is described in combination with an implementation method. For example Figure 3 FIG. shows another schematic flowchart of the image processing method provided by this application. The method of this embodiment may include:

[0113] S301, obtain the image to be detected.

[0114] S302, convert the image to be detected into a grayscale image.

[0115] S303, normalize the pixel values of each pixel point in the grayscale image to values between 0 and 1 to obtain the first image.

[0116] S304, perform at least one upsampling and at least one downsampling on the first image respectively to obtain a plurality of second images.

[0117] Combined with Figure 4 shown for illustration. For example Figure 4 FIG. shows a framework example diagram of an implementation process of the solution of this application.

[0118] In Figure 4 , after converting the image to be detected into a grayscale image and normalizing it to obtain the first image, it is necessary to generate and process second images of multiple scales based on the first image. In Figure 4 , taking the first image as the image as an example, the images of multiple scales include: the first image that remains unchanged (i.e., the original image in Figure 4 ), the image obtained by downsampling the first image to , and the image obtained by upsampling the first image to .

[0119] S305, perform sliding cropping on the first image using at least one size of first cropping window to obtain at least one set of image patches of the first image.

[0120] Among them, each set of image patches of the first image includes: at least one image patch cropped from the first image based on the first cropping window of the same size.

[0121] S306, for each second image, perform sliding cropping on the second image using at least one size of second cropping window to obtain multiple sets of image patches of the second image.

[0122] Among them, each set of image patches of the second image includes: at least one image patch cropped from the second image based on the second cropping window of the same size.

[0123] Still in combination with Figure 4 Explanation is as follows:

[0124] From Figure 4 It can be seen that for the first image of size , sliding cropping is performed on this first image respectively using the three sizes of cropping windows , and to obtain three sets of image patches, and the sizes of the image patches in these three sets of image patches are respectively , and .

[0125] For the second image of size , sliding cropping is performed on this second image respectively using the three sizes of cropping windows , and to obtain three sets of image patches with corresponding image patch sizes of , and .

[0126] For the second image of size , sliding cropping is performed on this second image respectively using the three sizes of cropping windows , and to obtain three sets of image patches with corresponding image patch sizes of , and .

[0127] S307. For any set of image patches corresponding to the first image and the second image, use the feature extraction model to extract the image features of each image patch in the set of image patches.

[0128] S308. For each image patch in each set of image patches, based on the image features of each image patch in the set of image patches, determine the set of feature similarities corresponding to the image patch.

[0129] Among them, the set of feature similarities of an image patch includes the feature similarities between the image patch and each of the other image patches in the set of image patches.

[0130] S309. For each image patch in each set of image patches, based on each feature similarity in the set of feature similarities of the image patch, determine the confidence corresponding to the set of feature similarities of the image patch.

[0131] The above steps S307 to S309 can refer to the relevant introduction in the previous embodiments.

[0132] Still combined with Figure 4 It is explained that in Figure 4 the example shown, the Transformer model can be used to calculate the feature vectors of each image patch in each set of image patches respectively. Then, for each set of image patches, based on the feature vectors of each image patch in the set of image patches, the feature similarities between the image patch and each of the other image patches in the set of image patches can be calculated.

[0133] On this basis, for each image patch, the feature similarities in the set of image patches of the image patch can be used to calculate a confidence, so that each image patch corresponds to a confidence.

[0134] S310. Determine the target image patches whose corresponding confidences meet the requirements from each set of image patches, and determine the target image area in the first image corresponding to the target image patches as the area where the image content in the first image is tampered with.

[0135] For example, the image patch with a confidence lower than the set threshold and the smallest confidence can be determined as the target image patch. Of course, there can be other ways to determine the target image patch. For details, please refer to the previous relevant introduction and will not be elaborated here.

[0136] It can be understood that by converting the image to be detected into a grayscale image and normalizing it, the text data in the first image can be made more prominent, which is beneficial for more reliably identifying the tampered text content in the image in the subsequent process and reducing the amount of data processing. On this basis, while preserving the first image, the present application also constructs second images of multiple scales that contain the same content image as the first image but have different resolutions. Moreover, different-sized cropping windows are used to crop the first image and the second images respectively, and the tampering identification of image content such as text is carried out in units of image blocks. Thus, in view of the characteristic that the difference features before and after text tampering in the image are relatively small, the feature difference labels of text and other image regions are enhanced through multi-scale images, and by analyzing the features of various fine-grained image blocks, the feature quality of different regions in the image is improved, which is beneficial for more accurately identifying the tampered text and other image content regions in the image and reducing the amount of data processing.

[0137] For the sake of easy understanding, compared with the existing commonly used image tampering detection models, the present solution has advantages such as higher detection accuracy. The present application explains the advantages of the present solution through the analysis of several indicators.

[0138] For the sake of easy understanding and description, several parameters as shown in Table 1 below are first constructed:

[0139] Table 1

[0140]

[0141] As can be seen from Table 1 above, the pixel points that are actually tampered in the image and are predicted to be tampered are denoted as TP; the pixel points that are actually tampered in the image and are predicted to be untampered are marked as FN; the pixel points that are actually untampered in the image and are predicted to be tampered are marked as FP; the pixel points that are actually untampered in the image and are predicted to be untampered are marked as TN.

[0142] Then, among all the actually tampered pixel points in the image, the pixel ratio TPR that can be correctly identified by the present solution or the existing model is the ratio of the recall rate to the sensitivity, that is, TPR = TP / (TP + FN).

[0143] Among all the untampered pixel points in the image, the ratio TNR that can correctly identify the pixel points that do not belong to the tampered ones by the present solution or the existing model can be expressed as: TNR = TN / (TN + FP).

[0144] On this basis, in order to compare the identification situations of the present solution and the existing tampering detection models for identifying the tampered pixel points in the image, the following several indicators are analyzed:

[0145] The first metric is: Intersection over Union (IOU), and its expression is as shown in Equation (1) below:

[0146] (Equation (1));

[0147] where represents the binary map corresponding to the predicted tampered region predicted by this solution or the tampering detection model, and B represents the ground truth mask of the actual tampered region of the true tampering. represents the overlapping area between the predicted tampered region and the actual tampered region, represents the union of the predicted tampered region and the actual tampered region. If the prediction is relatively conservative and will both decrease; if the prediction is relatively aggressive, then and will both increase; when the predicted tampered region and the actual tampered region completely overlap, the value of IOU is 1. Except for this optimal case, it should be ensured that is as large as possible, is as small as possible, that is, to ensure that all predicted regions are actual tampered regions, and there will be no situation where the tampered region cannot be predicted; or a large area is misdetected while only a part of the region is actually tampered, resulting in a very high false positive rate of the prediction.

[0148] For the accuracy of model segmentation, we use the maximum IOU at certain values of the fixed prediction TNR as the evaluation metric. For example, we calculate the maximum IOU of the tampered image within the threshold range where TNR is from 0.9 to 0.99, that is, calculate as shown in Equation (2) below:

[0149] (Equation (2));

[0150] where represents the maximum IOU when TNR = i, and the values of i are 0.9, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98 or 0.99.

[0151] The second metric is the F1 score, which is the harmonic mean of the recall rate and the precision rate.

[0152] Among them, Precision = TP / (FP + TP), which represents the proportion of truly positive samples among the samples predicted as positive classes. Recall = TNR = TN / (TN + FP).

[0153] The F1 score can be expressed as Equation (3) below:

[0154] (Formula 3);

[0155] The third metric is AUC (Area Under the Curve), which is the area under the ROC curve in full. The method for calculating AUC is the same as the conventional method.

[0156] Through actual tests, in the scenario of identifying the tampered areas in the image by using the two tampering recognition models of this solution, the existing PSCC-Net model, and the MVSS-Net model respectively, the metric values of the above several metrics obtained from the tests can be seen in Table 2 below:

[0157] Table 2

[0158]

[0159] As can be seen from Table 2 above, the performance of this solution is relatively better in several metrics, making the accuracy of identifying tampered areas such as text in the image relatively high.

[0160] Corresponding to an image processing method provided by this application, this application also provides an image processing device.

[0161] As Figure 5 , a schematic diagram of a composition structure of the image processing device provided by this application is shown. The device in this embodiment may include:

[0162] An image acquisition unit 501, configured to acquire a first image;

[0163] A first cropping unit 502, configured to perform sliding cropping on the first image by using at least one size of first cropping window, to obtain at least one set of image blocks of the first image. Each set of image blocks of the first image includes: at least one image block cropped from the first image based on the first cropping window of the same size;

[0164] A feature determination unit 503, configured to, for each set of image blocks, determine a set of feature similarities corresponding to each image block in the set of image blocks. The set of feature similarities of the image block includes the feature similarities between the image block and each other image block in the set of image blocks;

[0165] A tampering recognition unit 504, configured to determine a target image block with tampered image content based on the sets of feature similarities corresponding to the respective image blocks in the set of image blocks.

[0166] In a possible implementation manner, the tampering recognition unit includes:

[0167] A confidence calculation subunit, configured to determine, for each image block in each set of image blocks, a confidence corresponding to the set of feature similarities of the image block based on each feature similarity in the set of feature similarities of the image block;

[0168] A tampering recognition subunit, configured to determine, from each set of image blocks, a target image block whose corresponding confidence meets the requirements, and obtain a target image block whose image content is tampered;

[0169] In another possible implementation, it may be as Figure 6 shown, Figure 6 shows another structural schematic diagram of the image processing device provided in this application. As Figure 6 known, in addition to including Figure 5 the image acquisition unit 501, the first cropping unit 502, the feature determination unit 503, and the tampering recognition unit 504 shown, it further includes:

[0170] An image generation unit 505, configured to generate at least one second image based on the first image, where the resolution of the second image is different from that of the first image, and the resolutions of different second images are different;

[0171] A second cropping unit 506, configured to perform sliding cropping on each second image by using at least one size of second cropping window, and obtain at least one set of image blocks of the second image, where each set of image blocks of the second image includes: at least one image block cropped from the second image based on the second cropping window of the same size.

[0172] In a possible implementation, the image generation unit may include at least one of the following:

[0173] A first generation subunit, configured to perform at least one upsampling on the first image to obtain at least one second image;

[0174] A second generation subunit, configured to perform at least one downsampling on the first image to obtain at least one second image.

[0175] In a possible implementation, the sizes of the second cropping windows corresponding to different second images in the second cropping unit are not completely the same, and the size of the second cropping window is different from the size of the first cropping window.

[0176] In any of the above embodiments of the device in this application, it further includes:

[0177] A region matching unit, configured to determine a target image region in the first image corresponding to the target image block;

[0178] Determine the target image region as the region in the first image where the image content is tampered with.

[0179] Further, on the premise that the device includes an image generation unit and a second cropping unit, the region matching unit includes:

[0180] A region matching subunit, configured to, if the target image block is from a second image, determine the target image region corresponding to the target image block in the first image based on the scaling ratio between the first image and the second image to which the target image block belongs.

[0181] In the embodiment of any one of the above devices in the present application, the image acquisition unit may include:

[0182] An image acquisition subunit, configured to acquire an image to be detected;

[0183] A grayscale conversion subunit, configured to convert the image to be detected into a grayscale image;

[0184] A normalization subunit, configured to normalize the pixel values of each pixel point in the grayscale image to values between 0 and 1 to obtain a first image.

[0185] In the embodiment of the present application, an electronic device is further provided. As Figure 7 shown, it shows a schematic structural diagram of a composition of the electronic device. The electronic device includes at least a processor 701 and a memory 702;

[0186] The processor 701 is configured to execute the image processing method described in any one of the above embodiments;

[0187] The memory 702 is configured to store a program required for the processor to perform operations.

[0188] It can be understood that the electronic device may further include a display unit 703 and an input unit 704.

[0189] Of course, the electronic device may also have Figure 7 more or fewer components, which is not limited herein.

[0190] In the embodiment of the present application, a computer program product is further provided, including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device is enabled to implement any one of the image processing methods provided in the embodiment of the present application.

[0191] In the embodiment of the present application, a computer-readable storage medium is further provided. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device is enabled to implement any one of the image processing methods provided in the embodiment of the present application.

[0192] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.

[0193] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0194] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0195] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

Claims

1. An image processing method, comprising: obtaining a first image; Sliding cropping the first image using a first cropping window of at least one size to obtain at least one image block set of the first image, each image block set of the first image comprising: at least one image block cropped from the first image based on a first cropping window of the same size; For each image block set, determining a feature similarity set corresponding to each image block in the image block set, wherein the feature similarity set of the image blocks includes feature similarities between the image block and each other image block in the image block set; Based on the feature similarity set corresponding to each image block in the image block set, a target image block whose image content has been tampered with is determined.

2. The image processing method according to claim 1, wherein determining the target image block whose image content has been tampered with based on the feature similarity set corresponding to each image block in the image block set comprises: For each image block in each image block set, based on each feature similarity in the feature similarity set of the image block, determine the confidence corresponding to the feature similarity set of the image block; A corresponding target image block whose confidence level meets the requirement is determined from each image block set, and a target image block whose image content is tampered with is obtained.

3. The image processing method according to claim 1, further comprising: Based on the first image, generating at least one second image, wherein the resolution of the second image is different from that of the first image, and the resolutions of different second images are different; For each second image, sliding cropping is performed on the second image using a second cropping window of at least one size to obtain at least one image block set of the second image, and each image block set of the second image includes: at least one image block cropped from the second image based on a second cropping window of the same size.

4. The image processing method according to claim 3, further comprising: Determining a target image area in the first image corresponding to the target image block; The target image area is determined as an area in the first image where image content is tampered with.

5. The image processing method according to claim 3, wherein generating at least one second image based on the first image comprises at least one of the following: Performing at least one upsampling on the first image to obtain at least one second image; Perform at least one downsampling operation on the first image to obtain at least one second image.

6. The image processing method according to claim 4, wherein determining the target image area corresponding to the target image block in the first image comprises: If the target image block originates from the second image, based on a scaling ratio between the first image and the second image to which the target image block belongs, it is determined that the target image block corresponds to a target image region in the first image.

7. The image processing method according to claim 1, wherein obtaining the first image comprises: Obtaining an image to be detected; Converting the image to be detected into a grayscale image; The pixel value of each pixel in the grayscale image is normalized to a value between 0 and 1 to obtain a first image.

8. The image processing method according to claim 3, wherein the sizes of the second cropping windows corresponding to different second images are not completely the same, and the size of the second cropping window is different from the size of the first cropping window.

9. An image processing device, comprising: An image obtaining unit, configured to obtain a first image; A first cropping unit is configured to perform sliding cropping on the first image using a first cropping window of at least one size to obtain at least one image block set of the first image, wherein each image block set of the first image includes: at least one image block cropped from the first image based on a first cropping window of the same size; a feature determination unit, configured to determine, for each image block set, a feature similarity set corresponding to each image block in the image block set, wherein the feature similarity set of the image blocks includes feature similarities between the image block and each other image block in the image block set; The tampering identification unit is used to determine the target image block whose image content has been tampered with based on the feature similarity set corresponding to each image block in the image block set.

10. An electronic device comprising a memory and a processor; The processor is used to execute the image processing method according to any one of claims 1 to 8; The memory is used to store programs required for the processor to perform operations.