An image detection method, device, intelligent device and storage medium
By performing semantic segmentation and gradient calculation on the image to be detected, the gradient change region of the image is determined, and image patchwork detection is carried out in combination with the target image area, which solves the problem that the prior art is difficult to detect image patchwork under low noise conditions, and achieves a more accurate and robust image detection effect.
Patent Information
- Application Number
- CN202111115002.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-09-23
AI Technical Summary
In the prior art, when detecting low-quality images, it is impossible to effectively determine that the imaging device in the image is less noise.
By semantic segmentation and image feature extraction of the image to be detected, the target feature map is obtained, and gradient calculation is performed on it, the gradient change area is determined, and image patchwork detection is performed in combination with the target image area.
The details of the image stitching area are enhanced, the accuracy and robustness of image stitching detection are improved, and it is suitable for image detection under different noise conditions.
Smart Images

Figure CN114936996B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, in particular to the field of image processing technology, and specifically relates to an image detection method, device, intelligent device and storage medium. Background Art
[0002] With the development of image processing technology, two images can be pieced together to generate an image using image processing methods, and such a pieced image is often likely to be a low-quality image; in this case, it is necessary to perform image piecing detection on the image to determine whether the image is an image piecework. Currently, mainly a splicing positioning method based on a noise model is used to perform image piecing detection on a certain image, and its general principle is: since the splicing area and the non-splicing area in the image come from different images, the model noises of the two are usually inconsistent, and it is possible to judge whether there is image piecing by performing noise analysis on the image. However, when the noise of the imaging device used for the two pieced images included in the image is relatively small, the noise model cannot be used to judge whether there is image piecing in the image. Summary of the Invention
[0003] Embodiments of this application provide an image detection method, device, intelligent device and storage medium, which can better perform image piecing detection on an image.
[0004] On the one hand, embodiments of this application provide an image detection method, including:
[0005] Obtain an image to be detected;
[0006] Perform semantic segmentation on the image to be detected to obtain the target image region of the image to be detected, and perform image feature extraction on the image to be detected to obtain a target feature map;
[0007] Perform gradient operation on the target feature map to obtain a gradient feature map corresponding to the target feature map, and determine the gradient change region of the image to be detected according to the gradient feature map;
[0008] Perform image piecing detection on the image to be detected according to the gradient change region and the target image region.
[0009] On the one hand, embodiments of this application provide an image detection device, including:
[0010] An obtaining unit, configured to obtain an image to be detected;
[0011] A processing unit, configured to perform semantic segmentation on the image to be detected to obtain the target image region of the image to be detected, and perform image feature extraction on the image to be detected to obtain a target feature map;
[0012] The processing unit is further configured to perform gradient operation on the target feature map to obtain a gradient feature map corresponding to the target feature map, and determine a gradient change region of the image to be detected according to the gradient feature map;
[0013] The processing unit is further configured to perform image splicing detection on the image to be detected according to the gradient change region and the target image region.
[0014] On the one hand, an embodiment of the present application provides an intelligent device, including a processor and a memory, the processor and the memory are connected to each other, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the above-mentioned image detection method.
[0015] An embodiment of the present application provides a computer-readable storage medium, in which program instructions are stored, and when the program instructions are executed, they are used to implement the above-mentioned image processing method.
[0016] An embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium, and when the computer instructions are executed by the processor of the intelligent device, they execute the above-mentioned image processing method.
[0017] In the embodiment of the present application, the intelligent device can obtain an image to be detected, perform semantic segmentation on the image to be detected to obtain a target image region of the image to be detected; perform image feature extraction on the image to be detected to obtain a target feature map, perform gradient operation on the target feature map to obtain a gradient feature map corresponding to the target feature map, and determine a gradient change region of the image to be detected according to the gradient feature map; perform image splicing detection on the image to be detected according to the gradient change region and the target image region; by performing gradient operation on the target feature map, the detail features of the splicing region in the image to be detected can be enhanced, and then the image can be more accurately detected for image splicing according to the gradient change region and the target image region. Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0019] Figure 1 It is a schematic diagram of the architecture of an image detection system provided by an embodiment of the present application;
[0020] Figure 2 It is a schematic flowchart of an image detection method provided by an embodiment of the present application;
[0021] Figure 3a It is a schematic diagram of obtaining a target image region provided by an embodiment of the present application;
[0022] Figure 3b It is a schematic architecture diagram of an encoding and decoding network provided by an embodiment of the present application;
[0023] Figure 3c It is a schematic diagram of a feature map in a convolutional layer provided by an embodiment of the present application;
[0024] Figure 3d It is a schematic diagram of multiple feature maps provided by an embodiment of the present application;
[0025] Figure 3e It is a schematic diagram of obtaining a gradient change region provided by an embodiment of the present application;
[0026] Figure 4 It is a schematic flowchart of an image detection method provided by an embodiment of the present application;
[0027] Figure 5 It is a schematic diagram of calculating the intersection over union provided by an embodiment of the present application;
[0028] Figure 6 It is a schematic flowchart of an image detection method provided by an embodiment of the present application;
[0029] Figure 7 It is a schematic structural diagram of an image detection device provided by an embodiment of the present application;
[0030] Figure 8 It is a schematic structural diagram of an intelligent device provided by an embodiment of the present application. Detailed implementation manners
[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0032] An embodiment of the present application provides an image detection solution. The general principle of this image detection solution is as follows: When it is necessary to detect a certain image (i.e., the image to be detected), on the one hand, feature extraction can be performed on the image to be detected to extract the feature map of the image to be detected, and then gradient operation is performed on the feature map of the image to be detected to obtain the gradient feature of the image to be detected. This gradient feature is used to indicate the area where the gradient of the image to be detected changes significantly, and this area where the gradient changes significantly is often caused by the splicing area and the non-splicing area; on the other hand, semantic segmentation can be performed on the image to be detected to obtain the target image area of the image to be detected, and then the area where the gradient changes significantly and the target image area are compared to obtain the image splicing result (or called the picture splicing result) of the image to be detected. The image detection solution provided by the present application can calculate the gradient feature by using the characteristic that the pixel points between the splicing area and the non-splicing area change significantly, and use this gradient feature to assist in judging the image to be detected, which can better perform image splicing detection on the image, with strong robustness and a more concise operation process.
[0033] To better implement the above image detection solution, an embodiment of the present application provides an image detection system. Please refer to Figure 1 , Figure 1 which is a schematic diagram of the architecture of an image detection system provided by an embodiment of the present application. This image detection system may include at least one terminal device 101 and a server 102; different types of application programs can be installed on the terminal device 101. For example, an instant messaging application program, a live broadcast application program, a conference communication application program, etc. can be installed on the terminal device 101; the terminal device 101 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, an intelligent vehicle, etc. The server 102 can be used to store the application data and image data generated by different types of application programs on the terminal device 101. The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), and big data and artificial intelligence platforms, etc.
[0034] Among them, the above image detection solution is executed by the terminal device 101 or the server 102. When the image detection solution is executed by the terminal device 101, the images to be detected generated by the terminal device 101 in different types of application programs can be stored in the server. Then, when the terminal device 101 needs to perform image splicing detection on the images to be detected, the terminal device 101 can obtain the images to be detected from the server 102. After the terminal device 101 obtains the images to be detected from the server 102, it can extract features from the images to be detected to obtain the feature map of the images to be detected, calculate the gradient of the feature map to obtain the gradient features of the images to be detected, and determine the gradient change region of the images to be detected based on the gradient features. The terminal device 101 can perform semantic segmentation on the images to be detected to obtain the target image regions of the images to be detected, and then perform image splicing detection on the images to be detected based on the gradient change region and the target image regions.
[0035] Based on the above-provided image detection solution and image detection system, please refer to Figure 2 , Figure 2 which is a schematic flowchart of an image detection method provided by an embodiment of the present application. The image detection method can be executed by an intelligent device, and the intelligent device can be the above terminal device 101 or server 102. The image detection method can include the following steps S201 - S205.
[0036] S201. Obtain an image to be detected.
[0037] Among them, the image to be detected can be any static picture, any frame image in a dynamic image, or any frame image in any video, etc., which is not limited in the embodiments of the present application. The number of images to be detected can be one or more.
[0038] In a specific implementation, the intelligent device can obtain the image to be detected from the image resources or video resources pre - stored in the local space. Among them, the specific implementation of obtaining the image to be detected from the image resources can be: if the image resource is a static image, the image resource can be directly used as the image to be detected; if the image resource is a dynamic image, any frame image can be obtained from the frame images in the dynamic image as the image to be detected. It should be understood that the specific implementation of obtaining the image to be detected from the video resources is similar to that of obtaining the image to be detected from the dynamic image, and will not be elaborated here.
[0039] In one embodiment, the intelligent device can also obtain the image to be detected from the target video, where the target video can be a video obtained in real time. The target video can refer to the video obtained during a video call in an instant messaging scenario. For example, when multiple people conduct a video communication using an instant messaging application, the target video can be a video including multiple people's communications; or, the target video is obtained during a conference communication in a conference communication scenario. For example, when conducting a conference communication, the target video can include videos of multiple conference participants; or, the target video can also be obtained during a live broadcast in a live broadcast communication scenario. For example, when the live broadcast host is conducting a live broadcast, the target video can be a video related to the live broadcast host. Then, any frame of the image is obtained from the target video as the image to be detected.
[0040] In one embodiment, when the target user believes that there may be a spliced area and a non-spliced area in an image, an image detection request can be submitted; the intelligent device can receive the image detection request sent by the target user, and the image detection request carries the image to be detected. The target user can be the user who manages the image to be detected, or the target user can be the user who browses the image to be detected, etc. The embodiments of the present application do not make limitations.
[0041] S202. Perform semantic segmentation on the image to be detected to obtain the target image area of the image to be detected.
[0042] In a specific implementation, the intelligent device can call a semantic segmentation model to perform semantic segmentation on the image to be detected to obtain the target image area of the image to be detected. In one embodiment, the intelligent device can call a semantic segmentation model to perform semantic segmentation on the image to be detected to obtain the semantic segmentation area of the image to be detected, and perform a bounding box processing on the semantic segmentation area to obtain the target image area. Performing a bounding box processing on the semantic segmentation area to obtain the target image area is beneficial for subsequent image splicing detection of the image area to be detected. It should be understood that semantic segmentation is actually to identify whether there is a spliced area in the image to be detected. For example, in Figure 3a , the image to be detected input to the semantic segmentation model is a person image, and the person is like Figure 3a the 31 in; the intelligent device can perform semantic segmentation on the person image and can obtain a semantic segmentation area such as Figure 3a 32; then perform a bounding box processing on the semantic segmentation area 32 to obtain the target image area 33 of the person image.
[0043] In one embodiment, the semantic segmentation model may be a convolutional neural network, and the convolutional neural network may be an encoder-decoder network; when the semantic segmentation model is an encoder-decoder network, the intelligent device may call the encoder-decoder network to perform semantic segmentation on the image to be detected, so as to obtain the target image region. Specifically, the image to be detected is input into the encoder-decoder network, so that the encoder-decoder network performs semantic segmentation on the image to be detected; then the target image region output by the decoding sub-network in the encoder-decoder network based on the semantic segmentation operation is obtained. Among them, the structure of the encoder-decoder network may adopt the U-Net mode, and the encoder-decoder network is as follows Figure 3b shown. The encoder-decoder network mainly includes an encoding sub-network 301 and a decoding sub-network 302; the encoding sub-network 301 and the decoding sub-network 302 are alternately arranged with convolutional layers and pooling layers. In the encoding sub-network 301, the size of the feature map of the image to be detected gradually decreases and the dimension gradually increases after being processed by the convolutional layer and the pooling layer; while in the decoding sub-network 302, it is the opposite; the encoding sub-network 301 may include a pooling layer and a convolutional layer. Among them, the pooling layer in the encoding sub-network may be called the downsampling layer; the downsampling layer may be used to encode the image to be detected to obtain a feature map with a smaller size than the image to be detected (it can be understood that the image to be detected is compressed in the encoding sub-network); the convolutional layer is used to extract important feature information in the image to be detected. In the embodiment of the present application, the convolutional layer is mainly used to extract the feature information about the splicing region in the image to be detected; the decoding sub-network 302 may include a pooling layer and a convolutional layer; among them, the pooling layer in the decoding sub-network may be called the upsampling layer; the upsampling layer is used to restore and decode the feature map with a smaller size than the image to be detected to an image with the same size as the image to be detected. According to the convolutional kernel parameters used in the encoding sub-network, the corresponding convolutional kernel parameters are selected in the decoding sub-network, and the upsampling process is continuously performed to ensure that the sizes of the feature maps are the same. In the encoding sub-network 301 and the decoding sub-network 302, there are cross-line connections between the feature maps of the same size, and the cross-line connection can quickly restore the information loss. It can be understood that after the intelligent device calls the encoder-decoder network to perform semantic segmentation on the image to be detected, the output of the decoding sub-network of the encoder-decoder network may be the target image region of the image to be detected, or the output of the decoding sub-network of the encoder-decoder network may also be the semantic segmentation region, and then the semantic segmentation region is boxed to obtain the target image region.
[0044] In one embodiment, before invoking the encoding and decoding network to perform semantic segmentation on the image to be detected, the encoding and decoding network can be trained first. The intelligent device can obtain a plurality of training sample images, which can include negative sample images and positive sample images. A negative sample image refers to an image with a splicing area, and a positive sample image refers to an image without a splicing area. Among them, the number of negative sample images and the number of positive sample images need to meet certain preset conditions, and then the encoding and decoding network is trained using the plurality of training sample images. Among them, the preset conditions can be set according to requirements or experience; the plurality of training sample images can be obtained from the PASCAL VOC 2011 semantic dataset.
[0045] S203. Extract image features from the image to be detected to obtain a target feature map.
[0046] Among them, the target feature map includes important features in the image to be detected. In the embodiment of the present application, the target feature map mainly includes important feature information of the splicing area of the image to be detected. In a specific implementation, the intelligent device can extract the features of the image to be detected during the process of performing semantic segmentation on the image to be detected to obtain a target feature map.
[0047] In one embodiment, during the process of the intelligent device invoking the encoding and decoding network to perform semantic segmentation on the image to be detected, the target feature map corresponding to the image to be detected can be obtained from the encoding sub-network. More specifically, the intelligent device obtains the target feature map obtained based on the semantic segmentation operation from the target convolutional layer of the encoding sub-network of the encoding and decoding network. Among them, the target convolutional layer can be any convolutional layer in the encoding sub-network.
[0048] In one embodiment, during the process of the intelligent device invoking the encoding and decoding network to perform semantic segmentation on the image to be detected, one or more feature maps corresponding to the image to be detected can be obtained from the target convolutional layer of the encoding sub-network; since the information that appears most frequently in the image to be detected often appears in the same area of the feature map, the intelligent device can select any one of the one or more feature maps as the target feature map. For example, in Figure 3c , the image to be detected is a butterfly image. During the process of the intelligent device invoking the encoding and decoding network to perform semantic segmentation on the butterfly image, it is processed through the downsampling layer, and then the butterfly image processed by the downsampling layer enters the convolutional layer for processing. It can be seen that after being processed by the encoding sub-network, the size of the feature map corresponding to the butterfly image will become smaller and smaller; then the feature map of the butterfly image in the convolutional layer is as Figure 3c ; then the intelligent device extracts the features of the butterfly image in the convolutional layer to obtain multiple feature maps corresponding to the butterfly image, and the multiple feature maps are as Figure 3d shown, among which, in Figure 3dAmong them, the feature map 0 - feature included in 303 Figure 8 is obtained by feature extraction in the Figure 3c convolutional layer 301 shown. The feature map 0 - feature included in 304 Figure 8 is obtained by feature extraction in the Figure 3c convolutional layer 302 shown. Then the intelligent device can select the first feature map from the multiple feature maps included in 303 as the target feature map of the butterfly image; or the intelligent device can also select the first feature map from the multiple feature maps included in 304 as the target feature map of the butterfly image.
[0049] S204. Perform a gradient operation on the target feature map to obtain a gradient feature map corresponding to the target feature map, and determine the gradient change region of the image to be detected according to the gradient feature map.
[0050] In a specific implementation, the above - mentioned target feature map is composed of multiple pixel points. The intelligent device can perform a gradient operation on any two adjacent pixel points in the target feature map to obtain the gradient feature between any two adjacent pixel points. Then, according to the gradient feature between any two adjacent pixel points, the gradient feature map corresponding to the target feature map can be determined. By performing a gradient operation on the target feature map, the detailed features of the image to be detected can be increased (i.e., the features in the splicing area of the image to be detected are increased), while the smooth area of the image to be detected is weakened (i.e., the features in the non - splicing area of the image to be detected are weakened). This gradient feature map can be like the Figure 3e gradient feature map corresponding to the butterfly image shown. It can be seen that compared with the target feature map, the pixel brightness in this gradient feature map is more obvious. The reason for the more obvious pixel brightness is that important features in the image to be detected are increased through gradient operation. Then the intelligent device performs dilation and erosion processing on the gradient feature map to obtain the gradient change region of the image to be detected. Among them, this gradient change region is composed of pixel points with obvious gradient changes, and this gradient change region can be considered as the splicing area. This gradient change region can be like the Figure 3e area enclosed by the right - most white line.
[0051] S205. Perform image splicing detection on the image to be detected according to the gradient change region and the target image region.
[0052] In a specific implementation, the intelligent device can compare the gradient change region with the target image region, and then determine whether the image to be detected is an image patchwork according to the comparison result. Specifically, since the target image region includes the splicing region, the gradient change region is also used to indicate whether the image to be detected includes the splicing region. Therefore, the intelligent device compares the gradient change region with the target image region. If the gradient change region is the same as the target image region, it is determined that the image to be detected is a patchwork image; if the gradient change region is different from the target image region, it is determined that the image to be detected is a complete image. A complete image can be understood as an image without any splicing, which is the original image.
[0053] In one embodiment, since there are many semantics in the non-spliced region and it may be recognized as the target image region, the above-mentioned target image region does not necessarily include the splicing region; in this case, the intelligent device can calculate the intersection-over-union ratio of the gradient change region and the target image region, and perform image patchwork detection on the image to be detected according to the intersection-over-union ratio.
[0054] In the embodiment of the present application, the intelligent device can obtain the image to be detected, perform semantic segmentation on the image to be detected to obtain the target image region of the image to be detected, then perform image feature extraction on the image to be detected to obtain the target feature map, perform gradient operation on the target feature map to obtain the gradient feature map corresponding to the target feature map, and determine the gradient change region according to the gradient feature map, and perform image patchwork detection on the image to be detected according to the gradient change region and the target image region; by performing gradient operation on the target feature map to obtain the gradient change region, the detail features of the splicing region in the image to be detected can be enhanced, so that the image can be more accurately detected for image patchwork in combination with the gradient change region.
[0055] Based on the above-provided image detection scheme and image detection system, please refer to Figure 4 , Figure 4 which is a schematic flowchart of an image detection method provided by an embodiment of the present application. This image detection method can be executed by an intelligent device, and the intelligent device can be the above-mentioned terminal device 101 or server 102. This image detection method can include the following steps S401 - S405:
[0056] S401. Obtain the image to be detected.
[0057] S402. Perform semantic segmentation on the image to be detected to obtain the target image region of the image to be detected.
[0058] S403. Perform image feature extraction on the image to be detected to obtain the target feature map.
[0059] In a specific implementation, when the intelligent device calls the encoding and decoding network to perform semantic segmentation on the image to be detected, the intelligent device can extract the image features of the image to be detected from the convolutional layer of the encoding sub-network in the encoding and decoding network, and obtain the target feature map of the image to be detected.
[0060] Among them, for the specific implementation manners of steps S401 - S403, reference can be made to some or all of the implementation manners of the above steps S201 - S203, which will not be elaborated here.
[0061] S404. Call the target convolution kernel to perform gradient operation on the target feature map, obtain the gradient feature map corresponding to the target feature map, and determine the gradient change region of the image to be detected according to the gradient feature map.
[0062] Among them, in order to unify the gradient operation and the convolutional layer, the target convolution kernel provided by the embodiments of the present application may include one or both of the following: a convolution kernel for performing gradient operation on the target feature map in the horizontal direction and a convolution kernel for performing gradient operation on the target feature map in the vertical direction. The convolution kernel for performing gradient operation on the target feature map in the horizontal direction may be a 3×3 convolution kernel, etc.; among them, the 3×3 convolution kernel can be expressed as:
[0063]
[0064] The convolution kernel for performing gradient operation on the target feature map in the vertical direction may be a 3×3 convolution kernel, etc., and the 3×3 convolution kernel can be expressed as:
[0065]
[0066] In the actual application process, by performing gradient operation on the target feature map in the horizontal or vertical direction, the gradient operation can be regarded as a layer of convolution. The target convolution kernel can be set according to the actual situation, and the embodiments of the present application do not limit it.
[0067] In one embodiment, since an image belongs to a relatively special two-dimensional function, its differential form needs to be represented by partial derivatives. In the horizontal direction, there is the following formula:
[0068]
[0069] In the vertical direction, there is the following formula:
[0070]
[0071] However, compared with continuous functions, an image belongs to a discrete two-dimensional function because an image is composed of multiple pixel points arranged in a certain way. Therefore, the target feature map is also composed of multiple pixel points arranged in a certain way. In this case, the specific implementation of performing gradient operations on the target feature map by invoking the target convolution kernel is as follows: for the target pixel point among multiple pixel points, the intelligent device can invoke the target convolution kernel to perform gradient operations on the target pixel point and the adjacent pixel points adjacent to the target pixel point, obtain the gradient feature between the target pixel point and the adjacent pixel points, and generate the gradient feature map of the target feature map based on the gradient feature between the target pixel point and the adjacent pixel points. Among them, the target pixel point is any pixel point among multiple pixel points; it should be noted that when performing gradient operations on the target pixel point and the adjacent pixel points, what is actually obtained is the gradient value, and this gradient value can reflect the gradient feature between the target pixel point and the adjacent pixel points. The gradient feature map is generated by the gradient features between each pixel point in the target feature map and its adjacent pixel points. Since the target convolution kernel only performs gradient operations on pixel points in the horizontal or vertical direction, the minimum value ε of the difference between any two adjacent pixel points is 1. Therefore, the gradient operation formula in the horizontal direction can be expressed as:
[0072]
[0073] Among them, represents the gradient value corresponding to the target pixel point in the horizontal direction, (x, y) represents the target pixel point, and f(x, y) represents the pixel value of the target pixel point; (x + 1, y) represents the adjacent pixel point of the target pixel point in the horizontal direction, and f(x + 1, y) represents the pixel value of the adjacent pixel point.
[0074] The gradient operation formula in the vertical direction can be expressed as:
[0075]
[0076] Among them, represents the gradient value corresponding to the target pixel point in the vertical direction, (x, y) represents the target pixel point, and f(x, y) represents the pixel value of the target pixel point; (x, y + 1) represents the adjacent pixel point of the target pixel point in the vertical direction, and f(x, y + 1) represents the pixel value of the adjacent pixel point.
[0077] After obtaining the gradient feature map corresponding to the target feature map, the specific implementation of the intelligent device to determine the gradient change region based on the gradient feature map is as follows: The intelligent device performs dilation processing on the gradient features in the gradient feature map to obtain the dilated gradient feature map, and then performs erosion processing on the gradient features in the gradient feature map to obtain the eroded gradient feature map; and determines the gradient change region of the image to be detected based on the dilated gradient feature map and the eroded gradient feature map. Among them, when the intelligent device performs dilation processing on the gradient feature map, it actually means extending the gradient features corresponding to the gradient values that meet the preset threshold in the gradient feature map. After the dilation processing, the average value of the overall brightness of the gradient feature map is higher than that of the non-dilated gradient feature map, and the area of the region where the brightness corresponding to the gradient features in the gradient feature map is greater than the brightness threshold becomes larger, thereby realizing the extension processing of the gradient features, while the area of the region where the brightness corresponding to the gradient features in the gradient feature map is less than or equal to the brightness threshold becomes smaller or even disappears. When the intelligent device performs erosion processing on the gradient feature map, it actually means eroding the gradient features corresponding to the gradient values that do not meet the preset threshold in the gradient feature map; after the erosion processing, the average value of the overall brightness of the gradient feature map is lower than that of the non-eroded gradient feature map, the area of the region where the brightness corresponding to the gradient features in the gradient feature map is greater than the threshold becomes smaller or even disappears, and the area of the region where the brightness corresponding to the gradient features in the gradient feature map is less than or equal to the threshold becomes larger. The intelligent device can determine the gradient change region of the image to be detected based on the dilated gradient feature map and the eroded gradient feature map. Among them, the gradient change region of the image to be detected may include pixel points with obvious gradient changes, and these pixel points with obvious gradient changes are generally pixel points in the splicing region.
[0078] S405. Perform image splicing detection on the image to be detected according to the gradient change region and the target image region.
[0079] In one embodiment, the intelligent device can determine the intersection of the gradient change region and the target image region, and the union of the gradient change region and the target image region; then calculate the ratio between the intersection and the union, and perform image splicing detection on the image to be detected according to the ratio between the intersection and the union. For example, Figure 5 is the process of calculating the intersection over union ratio of the gradient change region and the target image region. In Figure 5 It includes the target image region and the gradient change region, and the target image region is the region indicated by the rectangular frame; then the intelligent device calculates the union between the target image region and the gradient change region, and calculates the intersection between the target image region and the gradient change region, and then can calculate the ratio IoU between the intersection and the union, and perform image splicing detection on the image to be detected according to Iou.
[0080] In one embodiment, the specific implementation of detecting image splicing for the image to be detected according to the ratio between the intersection and the union is as follows: Determine whether the ratio between the intersection and the union is greater than or equal to the target threshold. If the intelligent device determines that the ratio between the intersection and the union is greater than or equal to the target threshold, it can be determined that the image to be detected is a spliced image; if the intelligent device determines that the ratio between the intersection and the union is less than the target threshold, it can be determined that the image to be detected is a complete image. A complete image refers to an image without a splicing area and is the original image. Among them, the target threshold can be set according to requirements.
[0081] In one embodiment, the above target threshold can be set according to the image source of the image to be detected. For example, if the image source of the image to be detected is a picture resource, the target threshold can be set to 0.5; for another example, if the image source of the image to be detected is a live video, the target threshold can be set to 0.45. In another embodiment, the above target threshold can also be set according to the image category of the image to be detected. For example, if the image category of the image to be detected is a human image, the target threshold can be set to 0.3; for another example, if the image category of the image to be detected is an animal image, the target threshold can be set to 0.6.
[0082] In one implementation, the number of images to be detected can be multiple, and the multiple images to be detected are obtained from the target video. That is to say, it is necessary to determine whether there are multiple frames of images in a certain video that are all spliced images. The intelligent device can randomly obtain multiple frames of images from the target video and use all the multiple frames of images as the images to be detected. Then, the intelligent device can count the number of images determined to be spliced images among the multiple images to be detected and determine whether this number exceeds the quantity threshold. If this number exceeds the quantity threshold, it can be considered that the target video is not a real video and belongs to a video composed of spliced images; the intelligent device can add marking information to the target video, where the marking information is used to indicate that the target video is not recommended. If this number does not exceed the quantity threshold, it is considered that the target video is a video without splicing, and then the target video can be presented to the user. Among them, the quantity threshold can be set according to requirements. For example, the quantity threshold can be 3, 6, etc., and the embodiments of the present application do not make limitations.
[0083] In the embodiments of the present application, after the intelligent device obtains the image to be detected, it can perform semantic segmentation on the image to be detected to obtain the target image area of the image to be detected; then perform image feature extraction on the image to be detected to obtain the target feature map; and call the target convolution kernel to perform gradient operation on the target feature map to obtain the gradient feature map corresponding to the target feature map, and determine the gradient change area according to the gradient feature map; perform image splicing detection on the image to be detected according to the gradient change area and the target image area. By using the target convolution kernel to perform gradient operation on the target features, the gradient feature map can be obtained relatively simply and quickly. By performing portrait splicing detection on the image to be detected according to the gradient change area and the target image area, the image can be more accurately detected for image splicing, and the robustness is relatively high.
[0084] Based on the above-provided image detection method, the embodiments of the present application can be specifically applied to various live broadcast scenarios or various video recording scenarios. For example, please refer to Figure 6 , for the live broadcast scenario, the image detection method may include:
[0085] (1) The intelligent device intercepts any frame of the image in the live broadcast scenario as the image to be detected, and then inputs the image to be detected into the encoding and decoding network. The intelligent device calls the encoding and decoding network to perform semantic segmentation on the image to be detected to obtain the semantic segmentation area; then performs a bounding box processing on the semantic segmentation area to obtain the target image area.
[0086] (2) During the process of the intelligent device calling the encoding and decoding network to perform semantic segmentation on the image to be detected, it obtains the target feature map corresponding to the image to be detected from the convolutional layer of the encoding sub-network, and performs gradient calculation on the target feature map to obtain the gradient feature map. The brightness of each pixel point in the gradient feature map is obvious, and then the gradient feature map is processed to obtain the gradient change area.
[0087] (3) Calculate the intersection over union (IoU) of the target image area and the gradient change area, and then perform image splicing detection on the image to be detected according to the IoU; in Figure 6 , it is determined according to the IoU that the detection result of the image to be detected is image splicing.
[0088] Based on the above-provided image detection method, the embodiments of the present application provide a schematic structural diagram of an image detection device. The image detection device can be applied to the intelligent device in the above Figure 2 or Figure 4 corresponding embodiments; specifically, the image detection device can be a computer program (including program code) running on the intelligent device. For example, the image detection device is an application software; the image detection device can be used to execute the corresponding steps in the method provided by the embodiments of the present application. As Figure 7As shown in the figure, the image detection device may specifically include an acquisition unit 701 and a processing unit 702.
[0089] The acquisition unit 701 is configured to acquire an image to be detected;
[0090] The processing unit 702 is configured to perform semantic segmentation on the image to be detected to obtain a target image region of the image to be detected, and perform image feature extraction on the image to be detected to obtain a target feature map;
[0091] The processing unit 702 is further configured to perform a gradient operation on the target feature map to obtain a gradient feature map corresponding to the target feature map, and determine a gradient change region of the image to be detected according to the gradient feature map;
[0092] The processing unit 702 is further configured to perform an image stitching detection on the image to be detected according to the gradient change region and the target image region.
[0093] In one embodiment, when the processing unit 702 performs semantic segmentation on the image to be detected to obtain a target image region of the image to be detected, and performs image feature extraction on the image to be detected to obtain a target feature map, it may specifically be configured to:
[0094] Input the image to be detected into an encoder-decoder network, so that the encoder-decoder network performs semantic segmentation on the image to be detected;
[0095] Obtain a target image region output by a decoding sub-network in the encoder-decoder network based on semantic segmentation operation, and obtain a target feature map obtained based on semantic segmentation operation from a target convolutional layer of an encoding sub-network of the encoder-decoder network.
[0096] In one embodiment, when the processing unit 702 performs a gradient operation on the target feature map to obtain a gradient feature map corresponding to the target feature map, it may specifically be configured to:
[0097] Call a target convolutional kernel to perform a gradient operation on the target feature map to obtain a gradient feature map corresponding to the target feature map, where the target convolutional kernel includes one or both of the following: a convolutional kernel for performing a gradient operation on the target feature map in the horizontal direction and a convolutional kernel for performing a gradient operation on the target feature map in the vertical direction.
[0098] In one embodiment, when the processing unit 702 determines the gradient change region of the image to be detected according to the gradient feature map, it may specifically be configured to:
[0099] Perform dilation processing on the gradient features in the gradient feature map to obtain a dilated gradient feature map;
[0100] Erode the gradient features in the gradient feature map to obtain an eroded gradient feature map;
[0101] Determine the gradient change region of the image to be detected according to the dilated gradient feature map and the eroded gradient feature map.
[0102] In one embodiment, when the processing unit 702 performs image splicing detection on the image to be detected according to the gradient change region and the target image region, it may specifically be used for:
[0103] Determine the intersection of the gradient change region and the target image region, and determine the union of the gradient change region and the target image region;
[0104] Perform image splicing detection on the image to be detected according to the ratio between the intersection and the union.
[0105] In one embodiment, when the processing unit 702 performs image splicing detection on the image to be detected according to the ratio between the intersection and the union, it may specifically be used for:
[0106] If the ratio between the intersection and the union is greater than or equal to the target threshold, determine that the image to be detected is a spliced image;
[0107] If the ratio between the intersection and the union is less than the target threshold, determine that the image to be detected is a complete image.
[0108] In one embodiment, the number of images to be detected is multiple, and the multiple images to be detected are obtained from the target video. The processing unit 702 is further used for:
[0109] Count the number of images determined to be spliced images among the multiple images to be detected;
[0110] If the number exceeds the quantity threshold, add marking information to the target video, and the marking information is used to indicate that the target video is a non-recommendable video.
[0111] In one embodiment, the target video is obtained during a video call in an instant messaging scenario; or, the target video is obtained during a conference call in a conference communication scenario; or, the target video is obtained during a live broadcast in a live broadcast communication scenario.
[0112] It can be understood that the functions of the units of the image detection device in this embodiment can be specifically implemented according to the methods in the above method embodiments, and the specific implementation process can refer to the relevant descriptions of the above method embodiments, which will not be elaborated here.
[0113] In an embodiment of the present application, the intelligent device can acquire an image to be detected, perform semantic segmentation on the image to be detected to obtain the target image region of the image to be detected; then extract image features from the image to be detected to obtain a target feature map, perform gradient operation on the target feature map to obtain a gradient feature map corresponding to the target feature map, and determine a gradient change region according to the gradient feature map; perform image stitching detection on the image to be detected according to the gradient change region and the target image region; by performing gradient operation on the target feature map, the detail features of the stitching region in the image to be detected can be enhanced; then, according to the gradient change region and the target image region, the image can be better detected for image stitching.
[0114] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of an intelligent device provided by an embodiment of the present application. As Figure 8 shown, the intelligent device in this embodiment may include: one or more processors 801; one or more input devices 802, one or more output devices 803, and a memory 804. The above-mentioned processor 801, input device 802, output device 803, and memory 804 are connected through a bus 805. The memory 802 is used to store a computer program, and the computer program includes program instructions. The processor 801 is used to execute the program instructions stored in the memory 802 and perform the following operations: acquire an image to be detected; perform semantic segmentation on the image to be detected to obtain the target image region of the image to be detected, and extract image features from the image to be detected to obtain a target feature map; perform gradient operation on the target feature map to obtain a gradient feature map corresponding to the target feature map, and determine the gradient change region of the image to be detected according to the gradient feature map; perform image stitching detection on the image to be detected according to the gradient change region and the target image region.
[0115] In one embodiment, when the processor 801 performs semantic segmentation on the image to be detected to obtain the target image region of the image to be detected and extracts image features from the image to be detected to obtain a target feature map, it may specifically be used for:
[0116] Input the image to be detected into an encoding and decoding network, so that the encoding and decoding network performs semantic segmentation on the image to be detected;
[0117] Obtain the target image region output by the decoding sub-network in the encoding and decoding network based on the semantic segmentation operation, and obtain the target feature map obtained based on the semantic segmentation operation from the target convolutional layer of the encoding sub-network of the encoding and decoding network.
[0118] In one embodiment, when the processor 801 performs gradient operation on the target feature map to obtain a gradient feature map corresponding to the target feature map, it may specifically be used for:
[0119] Call the target convolution kernel to perform gradient operation on the target feature map to obtain the gradient feature map corresponding to the target feature map. The target convolution kernel includes one or both of the following: a convolution kernel for performing gradient operation on the target feature map in the horizontal direction and a convolution kernel for performing gradient operation on the target feature map in the vertical direction.
[0120] In one embodiment, when the processor 801 determines the gradient change region of the image to be detected according to the gradient feature map, it may specifically be used for:
[0121] Perform dilation processing on the gradient features in the gradient feature map to obtain a dilated gradient feature map;
[0122] Perform erosion processing on the gradient features in the gradient feature map to obtain an eroded gradient feature map;
[0123] Determine the gradient change region of the image to be detected according to the dilated gradient feature map and the eroded gradient feature map.
[0124] In one embodiment, when the processor 801 performs image stitching detection on the image to be detected according to the gradient change region and the target image region, it may specifically be used for:
[0125] Determine the intersection of the gradient change region and the target image region, and determine the union of the gradient change region and the target image region;
[0126] Perform image stitching detection on the image to be detected according to the ratio between the intersection and the union.
[0127] In one embodiment, when the processor 801 performs image stitching detection on the image to be detected according to the ratio between the intersection and the union, it may specifically be used for:
[0128] If the ratio between the intersection and the union is greater than or equal to the target threshold, determine that the image to be detected is a stitched image;
[0129] If the ratio between the intersection and the union is less than the target threshold, determine that the image to be detected is a complete image.
[0130] In one embodiment, the number of images to be detected is multiple, and the multiple images to be detected are obtained from the target video. The processor 801 may specifically be used for:
[0131] Count the number of images determined to be stitched images among the multiple images to be detected;
[0132] If the quantity exceeds the quantity threshold, add marking information to the target video, where the marking information is used to indicate that the target video is a non-recommendable video.
[0133] In one embodiment, the target video is obtained during a video call in an instant messaging scenario; or, the target video is obtained during a conference communication in a conference communication scenario; or, the target video is obtained during a live broadcast in a live broadcast communication scenario.
[0134] It should be understood that in the embodiments of the present application, the so-called processor 801 may be a central processing unit (CPU), and this processor 801 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0135] The memory 804 may include a read-only memory and a random access memory, and provide instructions and data to the processor 801. A part of the memory 804 may also include a non-volatile random access memory. The memory 804 may store the image to be detected.
[0136] In specific implementation, the processor 801, input device 802, output device 803, and memory 804 described in the embodiments of the present application may execute the implementation manners described in the image detection method provided in the embodiments of the present application, and may also execute the implementation manners of the intelligent device described in the embodiments of the present application, which will not be elaborated herein.
[0137] In the embodiments of the present application, the intelligent device may obtain the image to be detected, perform semantic segmentation on the image to be detected to obtain the target image area of the image to be detected, perform image feature extraction on the image to be detected to obtain the target feature map, perform gradient operation on the target feature map to obtain the gradient feature map corresponding to the target feature map, and determine the gradient change area of the image to be detected according to the gradient feature map; perform image stitching detection on the image to be detected according to the gradient change area and the target image area; by performing gradient operation on the target feature map, the detailed features of the stitching area in the image to be detected can be enhanced, and then the image can be accurately detected for image stitching according to the gradient change area and the target image area.
[0138] In an embodiment of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the steps performed in the above image detection embodiment can be executed.
[0139] An embodiment of the present application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions. When the computer instructions are stored in a computer-readable storage medium and executed by a processor of an intelligent device, the methods in all the above embodiments are executed.
[0140] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0141] The above-disclosed is only a preferred embodiment of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the invention.
Claims
1. An image detection method, characterized in that, comprising: obtaining an image to be detected; performing semantic segmentation on the image to be detected to obtain a target image region of the image to be detected, and performing image feature extraction on the image to be detected to obtain a target feature map; performing a gradient operation on the target feature map to obtain a gradient feature map corresponding to the target feature map, and determining a gradient change region of the image to be detected according to the gradient feature map; performing image splicing detection on the image to be detected according to the gradient change region and the target image region.
2. The method according to claim 1, characterized in that, the performing semantic segmentation on the image to be detected to obtain a target image region of the image to be detected, and performing image feature extraction on the image to be detected to obtain a target feature map includes: inputting the image to be detected into an encoder-decoder network to enable the encoder-decoder network to perform semantic segmentation on the image to be detected; obtaining a target image region output by a decoding sub-network in the encoder-decoder network based on a semantic segmentation operation, and obtaining a target feature map obtained based on the semantic segmentation operation from a target convolutional layer of an encoding sub-network of the encoder-decoder network.
3. The method according to claim 1, characterized in that, the performing a gradient operation on the target feature map to obtain a gradient feature map corresponding to the target feature map includes: invoking a target convolutional kernel to perform a gradient operation on the target feature map to obtain a gradient feature map corresponding to the target feature map, the target convolutional kernel including one or both of the following: a convolutional kernel for performing a gradient operation on the target feature map in a horizontal direction and a convolutional kernel for performing a gradient operation on the target feature map in a vertical direction.
4. The method according to claim 3, characterized in that, the determining a gradient change region of the image to be detected according to the gradient feature map includes: performing dilation processing on gradient features in the gradient feature map to obtain a dilated gradient feature map; performing erosion processing on gradient features in the gradient feature map to obtain an eroded gradient feature map; determining a gradient change region of the image to be detected according to the dilated gradient feature map and the eroded gradient feature map.
5. The method according to claim 1, characterized in that, the performing image splicing detection on the image to be detected according to the gradient change region and the target image region includes: determining an intersection of the gradient change region and the target image region, and determining a union of the gradient change region and the target image region; performing image splicing detection on the image to be detected according to a ratio between the intersection and the union.
6. The method according to claim 5, characterized in that, the performing image splicing detection on the image to be detected according to a ratio between the intersection and the union includes: if the ratio between the intersection and the union is greater than or equal to a target threshold, determining that the image to be detected is a spliced image; if the ratio between the intersection and the union is less than the target threshold, determining that the image to be detected is a complete image.
7. The method according to any one of claims 1-6, wherein, the number of the images to be detected is multiple, and the multiple images to be detected are obtained from a target video, and the method further includes: counting the number of the images determined to be pieced-together images among the multiple images to be detected; if the number exceeds a number threshold, adding marking information to the target video, where the marking information is used to indicate that the target video is a non-recommendable video.
8. The method according to claim 7, wherein, the target video is obtained during a video call in an instant messaging scenario; or, the target video is obtained during a conference communication in a conference communication scenario; or, the target video is obtained during a live broadcast in a live broadcast communication scenario.
9. An intelligent device, wherein, comprising: a processor adapted to implement one or more computer programs; and a computer storage medium storing one or more computer programs, and the one or more computer programs are adapted to be loaded and executed by the processor to perform the image detection method according to any one of claims 1-8.
10. A computer storage medium, wherein, the computer storage medium stores a computer program, and when the computer program is executed by a processor, it performs the image detection method according to any one of claims 1-8.
Citation Information
Patent Citations
A spliced image tampering detection method
CN109816676A
Image stitching positioning device and method based on deep clustering
CN112465700A