Method and apparatus for discriminating tampered information of face image, and electronic device
By extracting multi-dimensional feature information and building a generative adversarial network, the problem of inaccurate recognition results of facial information tampering in facial images is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- PCT/CN2024/135285
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-11-28
- Publication Date
- 2025-06-19
AI Technical Summary
The prior art recognizes whether face information in face images has been tampered with, and the recognition results are inaccurate and lacks effective solutions.
By obtaining the face image of tampered information to be detected, background feature information, face feature information, edge feature information and pixel difference information are extracted, and target detection model is constructed in combination with the generation of adversarial networks to detect and discriminate multi-dimensional feature information.
It improves the robustness and adaptability of the model, can more accurately identify tampered information in the face image, and enhances the accuracy of the detection results.
Smart Images

Figure CN2024135285_19062025_PF_FP_ABST
Abstract
Description
Method, device and electronic device for determining facial image tampering information
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 12, 2023, with application number 2023117068462, and application name “Method, device and electronic device for determining face image tampering information”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence, and more specifically, to a method, device, and electronic device for determining tampering information in facial images. Background Art
[0003] With the advancement of digitalization and intelligence and their widespread adoption across all industries, facial recognition, as a key foundational capability for digital risk control, has become a crucial method for identity verification, authentication, and verification in the information age. However, technological advancements also bring new challenges, exposing traditional security measures to new threats. The misuse of face-swapping technology has become a hallmark of emerging information risks in recent years. AI models generate "realistic" images based on training data, blurring the line between authenticity and fraud in identity verification. Therefore, AI face-swapping detection has become an increasingly important risk control capability. Since deep learning technology has demonstrated outstanding performance in computer vision, the detection and identification of AI face-swapping attacks has become a fiercely competitive battle within deep learning scenarios. Existing AI face-swapping detection technologies have proposed their own unique methods based on the research on face-swapping methods. For example, Celeb-DF, based on the ResNet network architecture, learns the difference between face forgeries and real faces on a large-scale dataset; the Face X-ray method explores the differences in image tampering traces reflected by the patterns of facial "edge" information appearing in the face-swapping process; for example, the FaceForensics++ method, which incorporates multiple AI face-swapping architectures, makes the scope of AI face-swapping detection not limited to a single, classic forgery source, thereby improving the generalization performance of detection.
[0004] In recent years, while AI face-swapping technology has achieved certain results, its methodological characteristics and problems have gradually become clear. First, the traditional model of AI face-swapping detection has always had a twin relationship with the specific AI face-swapping method. That is, a detection method is a solution developed for a certain type of classic face-swapping technology, such as deepfake. Some methods attempt to expand the scope of data sources to cover more face-swapping technologies, such as including realistic human faces generated by adversarial means to enrich the data source. However, mining image generation information from a single "face" dimension still cannot meet the demand for "face-swapping diversity." In particular, the influence of general AI large model (Foundation Model) technology has spread in the AIGC field, making the iteration of "generation" capabilities faster and increasing the pressure of traditional conventional iterations. These current situations have exacerbated the diversity and uncertainty of AI face-swapping attacks, making production operations increasingly uncontrollable, and the coverage of a single detection capability has gradually decreased. Secondly, most AI face-changing detection tasks have a coarse granularity in dividing the "tampering" features in the face-changing process, lack guidance for network algorithms, or simply use information dimensions such as faces and edges. This makes the networks trained for AI face-changing detection tasks prone to regional overfitting or focusing on secondary information, making it difficult to meet the diverse distribution and large differences in data scenarios such as financial payment scenarios, and its practicality is poor.
[0005] Currently, no effective solution has been proposed to the problem of inaccurate recognition results when artificial intelligence technology in related technologies identifies whether facial information in facial images has been tampered with. Summary of the Invention
[0006] The main purpose of this application is to provide a method, device and electronic device for determining whether facial image tampering information has been tampered with, so as to solve the problem of inaccurate recognition results when artificial intelligence technology in related technologies recognizes whether facial information in facial images has been tampered with.
[0007] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a method for distinguishing tampering information of a facial image is provided, the method comprising: obtaining a facial image of a person to be detected for tampering information to obtain a target image; extracting feature information from the target image to obtain a feature information set, wherein the feature information at least includes the following feature information: background feature information, facial feature information, edge feature information, and pixel difference information; inputting the feature information of the feature information set into a target detection model for detection to obtain a detection result, wherein the target detection model refers to a model obtained by constructing a generative adversarial network based on N learning tasks and training the generative adversarial network using the feature information set, where N is a positive integer; distinguishing the tampering information of the target image based on the detection result to obtain a distinction result.
[0008] In some embodiments, extracting feature information from the target image to obtain a feature information set includes: using a first segmentation model to extract background feature information in the target image to obtain a first image; using a detection model to extract facial feature information in the target image to obtain a second image, and extracting facial edge information in the target image to obtain a third image; performing edge blurring on the third image, and using an enhanced edge model to extract edge feature information of the processed third image to obtain a fourth image; using a selective reconstruction method filter to extract pixel difference information in the target image to obtain a fifth image; and combining the first image, the second image, the fourth image, and the fifth image to obtain the feature information set.
[0009] In some embodiments, the target detection model is obtained by the following steps: collecting images containing facial information to obtain a training set; extracting the feature information from each image contained in the training set to obtain a feature information set for each image; constructing branches corresponding to the N learning tasks based on the feature information and the generative adversarial network to obtain N branches; constructing a loss function corresponding to each branch in the N branches, and fusing the loss functions corresponding to each branch to obtain a target loss function; inputting each image in the training set and the feature information set of each image into the N branches for iterative calculation until the target loss function converges to obtain the target detection model.
[0010] In some embodiments, branches corresponding to the N learning tasks are respectively constructed based on the feature information and the generative adversarial network, and the obtained N branches include: constructing a first branch based on the generative adversarial network, wherein the first branch is used to identify tampering information of the face image; constructing a second branch based on the detection model and the second segmentation model, wherein the second branch is used to generate a face edge image and detect face information in the face image; constructing a third branch based on the image detection algorithm, wherein the third branch is used to detect tampering information of the face edge image; and determining the first branch, the second branch and the third branch as the N branches.
[0011] In some embodiments, constructing a first branch based on the generative adversarial network includes: constructing a first generator, a second generator, a third generator and a fourth generator according to the extraction method of each feature information in the feature information, wherein the first generator is used to extract the background feature information, the second generator is used to extract the facial feature information, the third generator is used to extract the edge feature information, and the fourth generator is used to extract the pixel difference information; using the first generator, the second generator, the third generator and the fourth generator to extract the feature information of each image in the training set to obtain a feature information set of each image; constructing a discriminator based on the generative adversarial network, wherein the discriminator is used to determine the tampering information of the facial image; constructing the first branch based on the first generator, the second generator, the third generator, the fourth generator and the discriminator, wherein the first branch is used to identify the tampering information of each image in the training set and identify the tampering information of each image in the feature information set of each image.
[0012] In some embodiments, a loss function corresponding to each of the N branches is constructed, and the loss functions corresponding to each branch are fused to obtain a target loss function, including: constructing a first loss function based on the difference between the predicted value and the actual value of each image in the first branch; constructing a second loss function corresponding to the second branch based on a preset loss function of the image segmentation algorithm; constructing a third loss function based on the difference between the predicted value and the actual value of each edge image in the third branch; configuring the weight of the first loss function, the weight of the second loss function, and the weight of the third loss function; and constructing the target loss function based on the first loss function, the second loss function, the third loss function, the weight of the first loss function, the weight of the second loss function, and the weight of the third loss function.
[0013] In some embodiments, after determining the tampering information of the target image based on the detection result and obtaining the determination result, the method further includes: generating an alarm message based on the determination result and the target image when the determination result indicates that the target image contains the tampering information; and sending the alarm message to the target object, wherein the target object processes the target image based on the alarm message.
[0014] In order to achieve the above-mentioned purpose, according to another aspect of the present application, a device for distinguishing facial image tampering information is provided, which includes: an acquisition unit, used to obtain a facial image to be detected for tampering information, and obtain a target image; an extraction unit, used to extract feature information from the target image, and obtain a feature information set, wherein the feature information at least includes the following feature information: background feature information, facial feature information, edge feature information, and pixel difference information; a detection unit, used to input the feature information of the feature information set into a target detection model for detection, and obtain a detection result, wherein the target detection model refers to a model obtained by constructing a generative adversarial network based on N learning tasks and training the generative adversarial network using the feature information set, and N is a positive integer; a discrimination unit, used to distinguish the tampering information of the target image based on the detection result, and obtain a discrimination result.
[0015] In some embodiments, the extraction unit includes: a first extraction subunit, used to extract background feature information in the target image using a first segmentation model to obtain a first image; a second extraction subunit, used to extract facial feature information in the target image using a detection model to obtain a second image, and extract facial edge information in the target image to obtain a third image; a third extraction subunit, used to perform edge blurring on the third image, and use an enhanced edge model to extract edge feature information of the processed third image to obtain a fourth image; a fourth extraction subunit, used to extract pixel difference information in the target image using a selective reconstruction method filter to obtain a fifth image; and a combination subunit, used to combine the first image, the second image, the fourth image and the fifth image to obtain the feature information set.
[0016] In some embodiments, the detection unit is obtained by the following steps: an acquisition subunit, used to acquire images containing facial information to obtain a training set; a fifth extraction subunit, used to extract the feature information from each image contained in the training set to obtain a feature information set of each image; a first construction subunit, used to construct branches corresponding to the N learning tasks based on the feature information and the generative adversarial network, to obtain N branches; a second construction subunit, used to construct a loss function corresponding to each branch in the N branches, and fuse the loss functions corresponding to each branch to obtain a target loss function; a calculation subunit, used to input each image in the training set and the feature information set of each image into the N branches for iterative calculation until the target loss function converges to obtain the target detection model.
[0017] In some embodiments, the first construction subunit includes: a first construction module for constructing a first branch based on the generative adversarial network, wherein the first branch is used to identify tampering information of a facial image; a second construction module for constructing a second branch based on the detection model and the second segmentation model, wherein the second branch is used to generate a facial edge image and detect facial information in the facial image; a third construction module for constructing a third branch based on an image detection algorithm, wherein the third branch is used to detect tampering information of the facial edge image; and a determination module for determining the first branch, the second branch, and the third branch as the N branches.
[0018] In some embodiments, the first construction module includes: a first construction submodule, used to construct a first generator, a second generator, a third generator and a fourth generator based on the extraction method of each feature information in the feature information, wherein the first generator is used to extract the background feature information, the second generator is used to extract the facial feature information, the third generator is used to extract the edge feature information, and the fourth generator is used to extract the pixel difference information; an extraction submodule, used to use the first generator, the second generator, the third generator and the fourth generator to extract the feature information of each image in the training set to obtain a feature information set of each image; a second construction submodule, used to construct a discriminator based on the generative adversarial network, wherein the discriminator is used to determine the tampering information of the facial image; a third construction submodule, used to construct the first branch based on the first generator, the second generator, the third generator, the fourth generator and the discriminator, wherein the first branch is used to identify the tampering information of each image in the training set and identify the tampering information of each image in the feature information set of each image.
[0019] In some embodiments, the second construction subunit includes: a fourth construction module for constructing a first loss function based on the difference between the predicted value and the actual value of each image in the first branch; a fifth construction module for constructing a second loss function corresponding to the second branch based on a preset loss function of the image segmentation algorithm; a sixth construction module for constructing a third loss function based on the difference between the predicted value and the actual value of each edge image in the third branch; a configuration module for configuring the weight of the first loss function, the weight of the second loss function and the weight of the third loss function; and a seventh construction module for constructing the target loss function based on the first loss function, the second loss function, the third loss function, the weight of the first loss function, the weight of the second loss function and the weight of the third loss function.
[0020] In some embodiments, the device further includes: a generating unit for determining the tampering information of the target image based on the detection result, and after obtaining the determination result, generating an alarm information based on the determination result and the target image when the determination result indicates that the target image contains the tampering information; and a sending unit for sending the alarm information to a target object, wherein the target object processes the target image based on the alarm information.
[0021] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a computer-readable storage medium is provided, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned methods for determining facial image tampering information.
[0022] In order to achieve the above-mentioned purpose, according to one aspect of the present application, an electronic device is provided, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by one or more processors, the one or more processors implement any of the above-mentioned methods for determining facial image tampering information.
[0023] Through the present application, the following steps are adopted: obtaining a facial image to be detected for tampering information to obtain a target image; extracting feature information from the target image to obtain a feature information set, wherein the feature information at least includes the following feature information: background feature information, facial feature information, edge feature information, and pixel difference information; inputting the feature information of the feature information set into a target detection model for detection to obtain a detection result, wherein the target detection model refers to a model obtained by constructing a generative adversarial network based on N learning tasks and training the generative adversarial network using the feature information set, where N is a positive integer; judging the tampering information of the target image based on the detection result to obtain a judgment result, which solves the problem of inaccurate recognition results when artificial intelligence technology in related technologies recognizes whether facial information in facial images has been tampered with. By extracting feature information of multiple dimensions from the target image and constructing and training a generative adversarial network, a target detection model is obtained. This can guide the model to focus on deep correlation information of the image, allowing the model to focus on richer and more diverse features, thereby improving the robustness and adaptability of the model. At the same time, by designing multiple learning tasks that intersect and constrain each other, the learning ability of the model is enhanced, thereby achieving the effect of improving the recognition accuracy of the target detection model. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0025] FIG1 is a flow chart of a method for determining facial image tampering information according to a first embodiment of the present application;
[0026] FIG2 is a schematic diagram of an optional method for determining tampering information of a facial image according to the first embodiment of the present application;
[0027] FIG3 is a second schematic diagram of an optional method for determining facial image tampering information provided in accordance with the first embodiment of the present application;
[0028] FIG4 is a schematic diagram of a device for determining facial image tampering information according to the second embodiment of the present application;
[0029] Figure 5 is a schematic diagram of an electronic device for determining facial image tampering information provided in Example 5 of the present application. DETAILED DESCRIPTION
[0030] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0031] It should be noted that the methods and devices determined by the processing methods, devices, storage media and electronic devices of the present application document can be used in the field of financial technology to improve the recognition accuracy in the process of identifying tampered information of facial images, and can also be used in any field other than the field of financial technology. The application fields of the methods and devices of the processing methods, devices, storage media and electronic devices of the present application document are not limited.
[0032] It should be noted that the user information (including but not limited to user device information, user personal information, user information contained in facial images, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, data in facial images, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant areas, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0033] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0035] Example 1
[0036] The present invention is described below in conjunction with preferred implementation steps. FIG1 is a flow chart of a method for determining facial image tampering information provided in accordance with Embodiment 1 of the present application. As shown in FIG1 , the method includes the following steps:
[0037] Step S101: Acquire a face image of a person whose information is to be detected for tampering, and obtain a target image.
[0038] In the first embodiment, tampering information refers to information left in a facial image after a naturally generated facial image is tampered with by synthesizing or modifying an image. The target image refers to the facial image to be detected for tampering information.
[0039] Step S102 , extracting feature information from the target image to obtain a feature information set, wherein the feature information at least includes the following feature information: background feature information, face feature information, edge feature information, and pixel difference information.
[0040] In the first embodiment of the present invention, in order to detect whether the target image contains tampered information, it is necessary to extract feature information of multiple dimensions from the target image, namely the above-mentioned background feature information, facial feature information, edge feature information, and pixel difference information, so as to determine whether the target image contains tampered information through the extracted feature information of multiple dimensions.
[0041] In step S103, the feature information of the feature information set is input into the target detection model for detection to obtain a detection result, wherein the target detection model refers to a model obtained by constructing a generative adversarial network based on N learning tasks and training the generative adversarial network using the feature information set, and N is a positive integer.
[0042] In the first embodiment of the present invention, in order to identify whether the target image contains tampered information, it is necessary to construct a generator and a discriminator of a generative adversarial network (GAN). The generator of the generative adversarial network generates an image, and the discriminator determines whether the image is a naturally generated image or a later synthesized image. The constructed generative adversarial network is trained. The two networks compete with each other. Through continuous competition and learning, the generator can generate more realistic samples, and the discriminator can also more accurately identify real samples and generated samples, thereby using the trained discriminator to build a target detection model for recognition.
[0043] Step S104 : determining the tampering information of the target image based on the detection result to obtain a determination result.
[0044] After obtaining the detection result in the first embodiment of the present invention, if the detection result is 1, it is determined that the target image does not contain tampering information, that is, the judgment result is a facial image naturally generated by the target image; if the detection result is 0, it is determined that the target image contains tampering information, that is, the judgment result is a facial image synthesized later by the target image.
[0045] In summary, the method for distinguishing tampering information of facial images provided in Example 1 of the present application obtains a target image by acquiring a facial image of a person to be detected for tampering information; extracts feature information from the target image to obtain a feature information set, wherein the feature information at least includes the following feature information: background feature information, facial feature information, edge feature information, and pixel difference information; inputs the feature information of the feature information set into a target detection model for detection to obtain a detection result, wherein the target detection model refers to a model obtained by constructing a generative adversarial network based on N learning tasks and training the generative adversarial network using a feature information set, and N is a positive integer; the tampering information of the target image is distinguished based on the detection result to obtain a distinction result, which solves the problem of inaccurate recognition results when artificial intelligence technology in related technologies recognizes whether facial information in facial images has been tampered with. By extracting feature information of multiple dimensions from the target image and constructing and training a generative adversarial network, a target detection model is obtained. This can guide the model to focus on deep correlation information of the image, allowing the model to focus on richer and more diverse features, thereby improving the robustness and adaptability of the model. At the same time, by designing multiple learning tasks that intersect and constrain each other, the learning ability of the model is enhanced, thereby achieving the effect of improving the recognition accuracy of the target detection model.
[0046] In some embodiments, in the method for discriminating face image tampering information provided in Embodiment 1 of the present application, feature information is extracted from a target image, and the obtained feature information set includes: extracting background feature information in the target image by using a first segmentation model to obtain a first image; extracting face feature information in the target image by using a detection model to obtain a second image, and extracting face edge information in the target image to obtain a third image; performing edge blurring processing on the third image, and extracting edge feature information of the processed third image by using a strengthened edge model to obtain a fourth image; extracting pixel difference information in the target image by using a selective reconstruction method filter to obtain a fifth image; combining the first image, the second image, the fourth image, and the fifth image to obtain a feature information set.
[0047] In Embodiment 1, in order to extract rich feature information from the target image, different methods can be used to process the target image, so as to extract the background feature information, face feature information, edge feature information, and pixel difference information of the target image.
[0048] In some embodiments, a portrait segmentation model (i.e., the above-mentioned first segmentation model) can be used to set the pixels in the area of the target image containing face information to 0, and the remaining background or invalid pixels to 1, to obtain a binary mask image, that is, the above-mentioned first image. The process of processing the target image by using the portrait segmentation model is shown in Formula 1: t = (F seg (I)[x > T])·I (1)
[0049] where x represents the pixel value in the target image, T represents a preset pixel threshold, F seg represents the portrait segmentation model, and I represents the input original image. The portrait segmentation model can be implemented by using the OpenCV library. For example, the target image is grayscaled and binarized to obtain Image A; flood filling is performed on Image A to obtain Image B; Image B is inverted to obtain Image C; finally, Image B and Image C are superimposed to obtain the binary mask image of the target image.
[0050] Then, a face detection model (e.g., the YOLO (You Only Look Once) model, the YOLO5 model, the SSD (Single Shot Multibox Detector) model, etc.) is used to obtain pixels within the region of the target image containing facial information, i.e., to capture the facial rectangular frame information, thereby obtaining the aforementioned second image. Simultaneously, the position and boundaries of the facial information in the target image are marked based on the face detection model, thereby obtaining a face mask, i.e., the aforementioned third image. In the field of target detection, a face mask is typically a binary image composed of pixels that contains the outline and position information of the face.
[0051] Next, the third image is blurred with a convolution kernel of 3 to soften the edges of the face and fully enhance the edge features. Then, the feature information of the target image in the face edge region is calculated using Formula 2 (i.e., the enhanced edge model described above) to obtain the fourth image. Formula 2 is as follows: B = 4·M·(1-M) (II)
[0052] Here, M represents the third image after the blur operation, and B represents the fourth image. As the key area for the distribution of the "feature traces" of the face swap, the "edge" part of the calculation is retained as a dimension of information.
[0053] Finally, the SRM filter domain features of the target image, i.e., the above-mentioned pixel difference features, are calculated to obtain the fifth image. For example, an SRM convolution template is constructed, convolved with the target image, and the multi-channel feature map is merged to obtain the pixel difference features of the target image. The SRM convolution template is shown in Formula 3.
[0054] Among them, Kernel SRM Represents the SRM convolution kernel. The SRM convolution kernel calculates the pixel differences in an image and amplifies the traces of external forgery in a natural image. As a dimension of information, it enhances the characteristic representation of the tampered information in the target image.
[0055] By using a portrait segmentation model to calculate the background feature information of the target image, using a face detection model to calculate the facial feature information of the target image, using a convolution operation and formula 2 to calculate the edge feature information of the target image, and using an SRM convolution template to calculate the pixel difference information of the target image, the multidimensional feature information of the target image can be obtained, thereby detecting whether the target image contains tampering information based on the extracted multidimensional feature information, thereby improving the accuracy of the detection results.
[0056] In some embodiments, in the method for determining facial image tampering information provided in Example 1 of the present application, the above-mentioned target detection model is obtained by the following steps: collecting images containing facial information to obtain a training set; extracting feature information from each image contained in the training set to obtain a feature information set of each image; constructing N branches corresponding to learning tasks based on the feature information and the generative adversarial network to obtain N branches; constructing a loss function corresponding to each branch in the N branches, and fusing the loss functions corresponding to each branch to obtain a target loss function; inputting each image in the training set and the feature information set of each image into the N branches for iterative calculation until the target loss function converges to obtain the target detection model.
[0057] In the first embodiment of the present invention, in order to detect whether the target image contains face tampering information, a generative adversarial network containing multiple branches can be constructed, and the constructed generative adversarial network can be trained, so as to judge whether the target image contains tampering information through the main branch among the multiple branches, and improve the accuracy of the main branch through other branches.
[0058] In some embodiments, naturally taken portrait images of various types, environments, and different numbers are obtained as a training set, and the label of each image in the training set is set to true (i.e., 1); then, according to the method of extracting background feature information, facial feature information, edge feature information, and pixel difference information of the target image, the background feature information, facial feature information, edge feature information, and pixel difference information of each image in the training set are extracted to obtain a feature information set for each image; secondly, N learning tasks are constructed, and the N learning tasks at least include: a learning task of extracting feature information and judging whether the image has tampered information based on the feature information and a learning task of auxiliary judgment; corresponding branches and loss functions are constructed according to the N learning tasks, and the loss functions corresponding to the N branches are fused to obtain a target loss function; finally, each image in the training set and each image in the feature information set generated based on each image in the training set are input into the N branches for iterative calculation, a gradually decreasing learning rate is set, and backpropagation is continuously performed to the N branches, and the parameters are adjusted until the target loss function converges to obtain a target detection model.
[0059] By constructing different learning tasks and building branches corresponding to different learning tasks based on the generative adversarial network, a discriminator framework based on the generative adversarial network is obtained. The training set and the generated feature information are used to iteratively train it to obtain a target detection model. The target detection model can then identify the tampered information in the target image, thereby improving the recognition accuracy of tampered information in facial images.
[0060] In some embodiments, in the method for determining facial image tampering information provided in Example 1 of the present application, N branches corresponding to learning tasks are constructed based on feature information and a generative adversarial network, respectively, and the obtained N branches include: constructing a first branch based on a generative adversarial network, wherein the first branch is used to identify tampering information of a facial image; constructing a second branch based on a detection model and a second segmentation model, wherein the second branch is used to generate a facial edge image and detect facial information in the facial image; constructing a third branch based on an image detection algorithm, wherein the third branch is used to detect tampering information of a facial edge image; and determining the first branch, the second branch, and the third branch as N branches.
[0061] In the first embodiment of the present invention, in order to improve the accuracy of identifying tampered information, multiple branches can be constructed to assist in the judgment of tampered information. Specifically, the first branch includes a generator and a discriminator of a generative adversarial network, which are used to identify whether a face image contains tampered information; the second branch includes a face detection model for extracting facial feature information and an occlusion decoder network based on a convolutional layer (i.e., the second segmentation model mentioned above), which is used to generate a binary mask image of the face edge in the face image and detect the probability that the face image contains face information; the third branch includes a detection model of a binary mask image of an edge mask, which is used to identify whether the generated binary mask image of the face edge contains tampered information. By constructing a generative adversarial network containing multiple branches, each branch can be used to learn different features and patterns, thereby generating more diverse samples, enhancing the stability of the model, improving the generation ability of the generative adversarial network, and at the same time improving the quality of the generated samples, reducing errors and noise in the generated samples, and improving the generalization ability of the generative adversarial network, so that it performs better on new samples, thereby achieving the effect of improving the recognition accuracy of the trained target detection model.
[0062] In some embodiments, in the method for distinguishing facial image tampering information provided in Example 1 of the present application, constructing a first branch based on a generative adversarial network includes: constructing a first generator, a second generator, a third generator, and a fourth generator based on a method for extracting each feature information in the feature information, wherein the first generator is used to extract background feature information, the second generator is used to extract facial feature information, the third generator is used to extract edge feature information, and the fourth generator is used to extract pixel difference information; using the first generator, the second generator, the third generator, and the fourth generator to extract feature information of each image in the training set to obtain a feature information set of each image; constructing a discriminator based on a generative adversarial network, wherein the discriminator is used to determine tampering information of the facial image; constructing a first branch based on the first generator, the second generator, the third generator, the fourth generator, and the discriminator, wherein the first branch is used to identify the tampering information of each image in the training set and to identify the tampering information of each image in the feature information set of each image.
[0063] In the first embodiment, in order to construct the first branch for determining whether a facial image contains tampering information, it is necessary to combine the generator and the discriminator of each feature information to construct the first branch.
[0064] In some embodiments, a first generator is constructed based on a portrait segmentation model that generates background feature information. For example, the first generator can be a background image generator G1 based on a transfomer layer and a fully connected layer; a second generator is constructed based on a face detection model that extracts facial feature information. For example, the second generator can be a face image generator G2 based on a convolutional layer and a fully connected layer; a third generator is constructed based on an edge blurring algorithm and an enhanced edge model that extract edge feature information. For example, the third generator can be an edge generator G3 based on a fully connected layer; a fourth generator is constructed based on an SRM filter that extracts pixel difference information. For example, the fourth generator can be a pixel difference generator G4 based on a convolutional and fully connected layer.
[0065] Then, a discriminator of the generative adversarial network is constructed. The discriminator is used to determine whether the face image is naturally generated, that is, it does not contain tampering information, or is a generated face image (the image generated by the first generator, the second generator, the third generator and the fourth generator mentioned above).
[0066] Finally, the first branch of the target detection model is composed of the constructed first generator, second generator, third generator, fourth generator and discriminator.
[0067] By constructing the first branch, the discriminator can be trained to judge the generated face images and the natural face images, so as to identify the tampering information in the face images through the discriminator.
[0068] In some embodiments, in the method for determining facial image tampering information provided in Example 1 of the present application, a loss function corresponding to each of the N branches is constructed, and the loss functions corresponding to each branch are fused to obtain a target loss function, including: constructing a first loss function based on the difference between the predicted value and the actual value of each image in the first branch; constructing a second loss function corresponding to the second branch based on the preset loss function of the image segmentation algorithm; constructing a third loss function based on the difference between the predicted value and the actual value of each edge image in the third branch; configuring the weight of the first loss function, the weight of the second loss function, and the weight of the third loss function; and constructing a target loss function based on the first loss function, the second loss function, the third loss function, the weight of the first loss function, the weight of the second loss function, and the weight of the third loss function.
[0069] In the first embodiment, in order to adjust the model parameters through continuous iterative training, it is necessary to set a corresponding loss function for each branch.
[0070] In some embodiments, the loss function of the first branch is as shown in Formula 4, loss main =-[E(log(D(x)))+E(log(1-D(G(z))))] (4)
[0071] Among them, D(x) is the discriminator's prediction of the real data, that is, the prediction of the natural image in the above-mentioned face image. The labels of all natural images are set to 1. Ideally, for real data, D(x) should be as close to 1 as possible, indicating that the discriminator can correctly identify the real data. G(z) is the generator's mapping of the extracted feature information, that is, the generated image containing feature information. The labels of all generated images are set to 0, and D(G(z)) is the discriminator's prediction of these generated images. Ideally, D(G(z)) should be as close to 0 as possible.
[0072] The loss function of the second branch is constructed based on a loss function commonly used in image segmentation tasks, namely the Dice loss function, as shown in Formula 5.
[0073] Among them, p i represents the predicted value of the i-th pixel, g i Represents the true value of the i-th pixel, smooth represents the smoothing factor to prevent the denominator from being 0. When the predicted value and the true value are exactly the same, the value of the Dice loss function is 0, indicating that the prediction is accurate, otherwise it means that the prediction is wrong.
[0074] The loss function of the third branch can be shown as formula 6,
[0075] Among them, y i represents the true label of the i-th instance, p i represents the predicted probability of the i-th instance, n is the total number of samples, and for each face image, if the true label is 1, the loss function is the negative logarithmic predicted probability; if its true label is 0, then the loss function is the positive logarithmic predicted probability.
[0076] According to the pre-configured weights assigned to the first branch, the second branch and the third branch, the loss functions of the three branches are fused to obtain the target loss function. The target loss function is shown in Formula 7, loss = loss main +τ1loss seg +τ2loss edge (seven)
[0077] Among them, loss main To generate the loss function value of the discriminant branch (i.e. the first branch mentioned above), loss seg is the loss function value of the second branch, lossedge is the loss function value of the third branch, τ1 and τ2 are the influence weights of the first and second branches respectively.
[0078] Optionally, in the method for determining facial image tampering information provided in Example 1 of the present application, after determining the tampering information of the target image based on the detection result and obtaining the determination result, the above method further includes: when the determination result indicates that the target image contains tampering information, generating an alarm information based on the determination result and the target image; and sending the alarm information to the target object, wherein the target object processes the target image based on the alarm information.
[0079] After obtaining the discrimination result in this embodiment 1, if the discrimination result indicates that the target image does not contain tampered information, it means that the customer is using a naturally generated facial image, and the possibility of risk is low; if the discrimination result indicates that the target image contains tampered information, it means that the customer is using a synthetic or face-swapped facial image, and the possibility of risk is high. It is necessary to generate an alarm message and send the alarm message to the business personnel or management personnel (i.e., the target object mentioned above), who will manage the customer, for example, temporarily prohibiting the customer from conducting transactions. By promptly notifying the business personnel of the financial institution if the target image contains tampered information, the customer's transaction security is guaranteed.
[0080] In some embodiments, in the first embodiment, the framework of the generative adversarial network constructed by this scheme can be as shown in Figure 2. Feature information is extracted from the real face images in the training set through multiple generators, and the generated images containing feature information and the real face images in the training set are input into the discriminator for discrimination. The discrimination result of the discriminator is compared with the real label of the image to obtain the loss function value, and the model parameters of the generator and the discriminator are adjusted according to the loss function value. Finally, the constructed generative adversarial network is iteratively trained to obtain the above-mentioned target detection model.
[0081] In some embodiments, in the first embodiment, the process of constructing a generative adversarial network based on multi-dimensional feature information and training it to obtain a target detection model can be shown in FIG3 .
[0082] In this solution, the process of constructing a generative adversarial network based on multi-dimensional feature information and training it to obtain a target detection model can be shown in Figure 3.
[0083] Step 1: Prepare portrait data and collect face images from various open source datasets to form a training set.
[0084] Step 2: traverse each face image in the training set, use the pre-trained portrait segmentation model based on the Resnet network to calculate and generate the portrait binary mask template image Mask1 in each image, and then segment the person and background in the image based on the binary portrait mask template to obtain image 1 containing portrait information and image 2 containing background information; based on the yolov-face face detection model, obtain the face rectangular area in the original image, and then use the non-maximum suppression algorithm to confirm that the image retains the face area, and detect the connectivity between the face rectangular area in this step and the portrait binary mask template image Mask1. If the two are connected, it means that the detection is accurate. By the formula Mask2=(argF l (Mask1,bbox)>threshold)∩bbox retains the face area mask Mask2, where argF l (Mask1,bbox)>threshold indicates that the connectivity between the face area and the portrait area is greater than the required part. By calculating the intersection of the portrait binary mask template image Mask1 and the rectangular area of the face image, the face area mask Mask2 is obtained, which can effectively ensure the segmentation accuracy in the process of automatic segmentation of the image area; traverse the face images in the training set, and perform Gaussian filtering with a convolution kernel size of 3 on the face area mask Mask2 obtained for each image to obtain a smoothed mask MaskG, and then calculate the face smooth edge template and the corresponding original image information; construct an SRM filter template, call the Conv2d function provided by the nn module in the deep learning pytorch library to perform SRM kernel convolution calculation on the face images in the training set, and obtain the SRM filter domain features of the face image.
[0085] Step three: Integrate the feature information of multiple dimensions mentioned above, and call the corresponding convolution and vit modules in the pytorch library and transformer library based on the integration results to build the generator corresponding to each feature information respectively, and use the multi-dimensional feature information extracted from the real images in the training set as the generation learning target of the corresponding generator.
[0086] Step 4: Combine multiple generators and data streams from the real dataset, set the label bits separately, the label of the real image in the training set is 1, and the label of the image generated by the generator is 0, which serves as the data stream of the discriminator.
[0087] Step 5: Call the convolutional layer, linear layer, and normalization layer from the deep learning pytorch library as the main structure, and select the LeakyReLU activation function to build the discriminator backbone network, thereby extending the multi-task branch. It is mainly divided into the traditional adversarial generation discriminator network branch of the main branch, the face segmentation network branch, and the forged area edge prediction branch. Specifically, the shared image feature extraction module contains multiple residual blocks and uses a Bottleneck structure. Each residual block contains three convolution layers and a skip connection. The convolution layers are 1x1 convolution layer, 3x3 convolution layer, and 1x1 convolution layer respectively. The output of each residual block is obtained by adding the output of the skip connection and the convolution layer and then processing it through the ReLU activation function. The main task network mainly contains a convolution layer, which retains more spatial information to further extract image features.
[0088] Step 6: Based on the above multiple branches, construct the multi-dimensional feature adversarial generation task loss function, the face segmentation subtask loss function, and the face edge forgery area generation task loss function respectively, and add each part to form a comprehensive loss function to construct the discriminator module of the multi-task learning architecture.
[0089] Step 7: Call the torch Adam optimization function to perform backpropagation on the overall network model of the constructed generative adversarial network.
[0090] Step 8: Iterate the above calculation process to complete the training and obtain the target detection model.
[0091] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0092] Example 2
[0093] The second embodiment of the present application also provides a device for determining whether a facial image has been tampered with. It should be noted that the device for determining whether a facial image has been tampered with in the second embodiment of the present application can be used to execute the method for determining whether a facial image has been tampered with in the first embodiment of the present application. The device for determining whether a facial image has been tampered with in the second embodiment of the present application is described below.
[0094] 4 is a schematic diagram of a device for identifying facial image tampering information according to Embodiment 2 of the present application. As shown in FIG4 , the device includes: an acquisition unit 401 , an extraction unit 402 , a detection unit 403 , and a determination unit 404 .
[0095] Specifically, the acquisition unit 401 is configured to acquire a face image of a person whose information is to be detected for tampering, and obtain a target image.
[0096] The extraction unit 402 is used to extract feature information from the target image to obtain a feature information set, wherein the feature information at least includes the following feature information: background feature information, face feature information, edge feature information, and pixel difference information.
[0097] The detection unit 403 is used to input the feature information of the feature information set into the target detection model for detection to obtain a detection result, wherein the target detection model refers to a model obtained by constructing a generative adversarial network based on N learning tasks and training the generative adversarial network using the feature information set, and N is a positive integer.
[0098] The determination unit 404 is configured to determine whether the target image has been tampered with based on the detection result, and obtain a determination result.
[0099] The second embodiment of the present application provides a device for distinguishing tampering information of a facial image. The acquisition unit 401 obtains the facial image to be detected for tampering information to obtain a target image; the extraction unit 402 extracts feature information from the target image to obtain a feature information set, wherein the feature information at least includes the following feature information: background feature information, facial feature information, edge feature information, and pixel difference information; the detection unit 403 inputs the feature information of the feature information set into a target detection model for detection to obtain a detection result, wherein the target detection model refers to a model obtained by constructing a generative adversarial network based on N learning tasks and training the generative adversarial network using a feature information set, where N is a positive integer; the discrimination unit 404 discriminates the tampering information of the target image based on the detection result to obtain a discrimination result, which solves the problem in the related art of inaccurate recognition results when artificial intelligence technology recognizes whether facial information in a facial image has been tampered with. By extracting feature information of multiple dimensions from the target image and constructing and training a generative adversarial network, a target detection model is obtained. This can guide the model to focus on deep correlation information of the image, allowing the model to focus on richer and more diverse features, thereby improving the robustness and adaptability of the model. At the same time, by designing multiple learning tasks that intersect and constrain each other, the learning ability of the model is enhanced, thereby achieving the effect of improving the recognition accuracy of the target detection model.
[0100] In some embodiments, in the facial image tampering information identification device provided in Example 2 of the present application, the above-mentioned extraction unit 402 includes: a first extraction subunit, used to use a first segmentation model to extract background feature information in the target image to obtain a first image; a second extraction subunit, used to use a detection model to extract facial feature information in the target image to obtain a second image, and extract facial edge information in the target image to obtain a third image; a third extraction subunit, used to perform edge blurring on the third image, and use an enhanced edge model to extract edge feature information of the processed third image to obtain a fourth image; a fourth extraction subunit, used to use a selective reconstruction method filter to extract pixel difference information in the target image to obtain a fifth image; and a combination subunit, used to combine the first image, the second image, the fourth image and the fifth image to obtain a feature information set.
[0101] In some embodiments, in the device for determining facial image tampering information provided in Example 2 of the present application, the above-mentioned detection unit 403 is obtained by the following steps: an acquisition subunit, used to acquire images containing facial information to obtain a training set; a fifth extraction subunit, used to extract feature information from each image contained in the training set to obtain a feature information set of each image; a first construction subunit, used to construct N branches corresponding to learning tasks based on the feature information and the generative adversarial network to obtain N branches; a second construction subunit, used to construct a loss function corresponding to each branch in the N branches, and fuse the loss functions corresponding to each branch to obtain a target loss function; a calculation subunit, used to input each image in the training set and the feature information set of each image into the N branches for iterative calculation until the target loss function converges to obtain a target detection model.
[0102] In some embodiments, in the device for determining facial image tampering information provided in Example 2 of the present application, the above-mentioned first construction subunit includes: a first construction module for constructing a first branch based on a generative adversarial network, wherein the first branch is used to identify tampering information of a facial image; a second construction module for constructing a second branch based on a detection model and a second segmentation model, wherein the second branch is used to generate a facial edge image and detect facial information in the facial image; a third construction module for constructing a third branch based on an image detection algorithm, wherein the third branch is used to detect tampering information of a facial edge image; and a determination module for determining the first branch, the second branch, and the third branch as N branches.
[0103] In some embodiments, in the device for distinguishing facial image tampering information provided in Example 2 of the present application, the above-mentioned first construction module includes: a first construction sub-module, used to construct a first generator, a second generator, a third generator and a fourth generator based on the extraction method of each feature information in the feature information, wherein the first generator is used to extract background feature information, the second generator is used to extract facial feature information, the third generator is used to extract edge feature information, and the fourth generator is used to extract pixel difference information; an extraction sub-module, used to use the first generator, the second generator, the third generator and the fourth generator to extract feature information of each image in the training set to obtain a feature information set of each image; a second construction sub-module, used to construct a discriminator based on a generative adversarial network, wherein the discriminator is used to determine the tampering information of the facial image; a third construction sub-module, used to construct a first branch based on the first generator, the second generator, the third generator, the fourth generator and the discriminator, wherein the first branch is used to identify the tampering information of each image in the training set and to identify the tampering information of each image in the feature information set of each image.
[0104] In some embodiments, in the device for determining facial image tampering information provided in Example 2 of the present application, the above-mentioned second construction subunit includes: a fourth construction module, used to construct a first loss function based on the difference between the predicted value and the actual value of each image in the first branch; a fifth construction module, used to construct a second loss function corresponding to the second branch based on the preset loss function of the image segmentation algorithm; a sixth construction module, used to construct a third loss function based on the difference between the predicted value and the actual value of each edge image in the third branch; a configuration module, used to configure the weight of the first loss function, the weight of the second loss function and the weight of the third loss function; a seventh construction module, used to construct a target loss function based on the first loss function, the second loss function, the third loss function, the weight of the first loss function, the weight of the second loss function and the weight of the third loss function.
[0105] In some embodiments, in the facial image tampering information identification device provided in Example 2 of the present application, the above-mentioned device also includes: a generation unit, which is used to identify the tampering information of the target image based on the detection result, and after obtaining the identification result, when the identification result indicates that the target image contains tampering information, generates an alarm information based on the identification result and the target image; a sending unit, which is used to send the alarm information to the target object, wherein the target object processes the target image based on the alarm information.
[0106] The device for determining facial image tampering information includes a processor and a memory. The above-mentioned acquisition unit 401, extraction unit 402, detection unit 403 and determination unit 404 are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize corresponding functions.
[0107] The processor contains a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be set, and the accuracy of the judgment results can be improved by adjusting the kernel parameters.
[0108] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0109] A third embodiment of the present invention provides a computer-readable storage medium having a program stored thereon, which implements a method for determining facial image tampering information when the program is executed by a processor.
[0110] A fourth embodiment of the present invention provides a processor, which is used to run a program, wherein the program executes a method for determining facial image tampering information when running.
[0111] As shown in Figure 5, embodiment five of the present invention provides an electronic device, which includes a processor, a memory, and a program stored in the memory and runnable on the processor. When the processor executes the program, the following steps are implemented: obtaining a facial image of a person whose tampering information is to be detected to obtain a target image; extracting feature information from the target image to obtain a feature information set, wherein the feature information includes at least the following feature information: background feature information, facial feature information, edge feature information, and pixel difference information; inputting the feature information of the feature information set into a target detection model for detection to obtain a detection result, wherein the target detection model refers to a model obtained by constructing a generative adversarial network based on N learning tasks and training the generative adversarial network using the feature information set, where N is a positive integer; and discriminating the tampering information of the target image based on the detection result to obtain a discrimination result.
[0112] When the processor executes the program, the following steps are also implemented: extracting feature information from the target image to obtain a feature information set, including: using a first segmentation model to extract background feature information in the target image to obtain a first image; using a detection model to extract facial feature information in the target image to obtain a second image, and extracting facial edge information in the target image to obtain a third image; performing edge blurring on the third image, and using an enhanced edge model to extract edge feature information of the processed third image to obtain a fourth image; using a selective reconstruction method filter to extract pixel difference information in the target image to obtain a fifth image; and combining the first image, the second image, the fourth image, and the fifth image to obtain a feature information set.
[0113] When the processor executes the program, the following steps are also implemented: the above-mentioned target detection model is obtained by the following steps: collecting images containing facial information to obtain a training set; extracting feature information from each image contained in the training set to obtain a feature information set of each image; constructing N branches corresponding to learning tasks based on the feature information and the generative adversarial network to obtain N branches; constructing a loss function corresponding to each branch in the N branches, and fusing the loss functions corresponding to each branch to obtain a target loss function; inputting each image in the training set and the feature information set of each image into the N branches for iterative calculation until the target loss function converges to obtain the target detection model.
[0114] When the processor executes the program, the following steps are also implemented: N branches corresponding to the learning tasks are constructed based on the feature information and the generative adversarial network, and the N branches obtained include: constructing a first branch based on the generative adversarial network, wherein the first branch is used to identify tampering information of the face image; constructing a second branch based on the detection model and the second segmentation model, wherein the second branch is used to generate a face edge image and detect face information in the face image; constructing a third branch based on the image detection algorithm, wherein the third branch is used to detect tampering information of the face edge image; and determining the first branch, the second branch and the third branch as N branches.
[0115] When the processor executes the program, the following steps are also implemented: constructing a first branch based on a generative adversarial network, including: constructing a first generator, a second generator, a third generator, and a fourth generator based on a method for extracting each feature information in the feature information, wherein the first generator is used to extract background feature information, the second generator is used to extract facial feature information, the third generator is used to extract edge feature information, and the fourth generator is used to extract pixel difference information; using the first generator, the second generator, the third generator, and the fourth generator to extract feature information of each image in the training set to obtain a feature information set of each image; constructing a discriminator based on a generative adversarial network, wherein the discriminator is used to determine tampering information of a facial image; constructing a first branch based on the first generator, the second generator, the third generator, the fourth generator, and the discriminator, wherein the first branch is used to identify tampering information of each image in the training set and identify tampering information of each image in the feature information set of each image.
[0116] When the processor executes the program, it also implements the following steps: constructing a loss function corresponding to each of the N branches, and fusing the loss functions corresponding to each branch to obtain a target loss function, including: constructing a first loss function based on the difference between the predicted value and the actual value of each image in the first branch; constructing a second loss function corresponding to the second branch based on the preset loss function of the image segmentation algorithm; constructing a third loss function based on the difference between the predicted value and the actual value of each edge image in the third branch; configuring the weight of the first loss function, the weight of the second loss function, and the weight of the third loss function; and constructing a target loss function based on the first loss function, the second loss function, the third loss function, the weight of the first loss function, the weight of the second loss function, and the weight of the third loss function.
[0117] When the processor executes the program, the following steps are also implemented: after determining the tampering information of the target image based on the detection result and obtaining the determination result, the above method also includes: when the determination result indicates that the target image contains tampering information, generating an alarm information based on the determination result and the target image; sending the alarm information to the target object, wherein the target object processes the target image based on the alarm information.
[0118] The devices in this article can be servers, PCs, PADs, mobile phones, etc.
[0119] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program initialized with the following method steps: obtaining a facial image of a person whose tampering information is to be detected to obtain a target image; extracting feature information from the target image to obtain a feature information set, wherein the feature information at least includes the following feature information: background feature information, facial feature information, edge feature information, and pixel difference information; inputting the feature information of the feature information set into a target detection model for detection to obtain a detection result, wherein the target detection model refers to a model obtained by constructing a generative adversarial network based on N learning tasks and training the generative adversarial network using a feature information set, where N is a positive integer; and discriminating the tampering information of the target image based on the detection result to obtain a discrimination result.
[0120] When executed on a data processing device, it is also suitable for executing a program initialized with the following method steps: extracting feature information from a target image to obtain a feature information set including: using a first segmentation model to extract background feature information in the target image to obtain a first image; using a detection model to extract facial feature information in the target image to obtain a second image, and extracting facial edge information in the target image to obtain a third image; performing edge blurring on the third image, and using an enhanced edge model to extract edge feature information of the processed third image to obtain a fourth image; using a selective reconstruction method filter to extract pixel difference information in the target image to obtain a fifth image; and combining the first image, the second image, the fourth image, and the fifth image to obtain a feature information set.
[0121] When executed on a data processing device, it is also suitable for executing a program initialized with the following method steps: the above-mentioned target detection model is obtained by the following steps: collecting images containing facial information to obtain a training set; extracting feature information from each image contained in the training set to obtain a feature information set of each image; constructing N branches corresponding to learning tasks based on the feature information and the generative adversarial network to obtain N branches; constructing a loss function corresponding to each branch in the N branches, and fusing the loss functions corresponding to each branch to obtain a target loss function; inputting each image in the training set and the feature information set of each image into the N branches for iterative calculation until the target loss function converges to obtain a target detection model.
[0122] When executed on a data processing device, it is also suitable for executing an initialized program having the following method steps: constructing N branches corresponding to learning tasks based on feature information and a generative adversarial network, respectively, to obtain N branches including: constructing a first branch based on a generative adversarial network, wherein the first branch is used to identify tampering information of a face image; constructing a second branch based on a detection model and a second segmentation model, wherein the second branch is used to generate a face edge image and detect face information in the face image; constructing a third branch based on an image detection algorithm, wherein the third branch is used to detect tampering information of a face edge image; and determining the first branch, the second branch, and the third branch as N branches.
[0123] When executed on a data processing device, it is also suitable for executing a program initialized with the following method steps: constructing a first branch based on a generative adversarial network, including: constructing a first generator, a second generator, a third generator and a fourth generator according to the extraction method of each feature information in the feature information, wherein the first generator is used to extract background feature information, the second generator is used to extract facial feature information, the third generator is used to extract edge feature information, and the fourth generator is used to extract pixel difference information; using the first generator, the second generator, the third generator and the fourth generator to extract the feature information of each image in the training set to obtain a feature information set of each image; constructing a discriminator based on a generative adversarial network, wherein the discriminator is used to determine the tampering information of the facial image; constructing a first branch based on the first generator, the second generator, the third generator, the fourth generator and the discriminator, wherein the first branch is used to identify the tampering information of each image in the training set and identify the tampering information of each image in the feature information set of each image.
[0124] When executed on a data processing device, it is also suitable for executing a program initialized with the following method steps: constructing a loss function corresponding to each branch of N branches, and fusing the loss functions corresponding to each branch to obtain a target loss function including: constructing a first loss function based on the difference between the predicted value and the actual value of each image in the first branch; constructing a second loss function corresponding to the second branch based on the preset loss function of the image segmentation algorithm; constructing a third loss function based on the difference between the predicted value and the actual value of each edge image in the third branch; configuring the weight of the first loss function, the weight of the second loss function and the weight of the third loss function; constructing the target loss function based on the first loss function, the second loss function, the third loss function, the weight of the first loss function, the weight of the second loss function and the weight of the third loss function.
[0125] When executed on a data processing device, it is also suitable for executing an initialized program having the following method steps: after determining the tampering information of the target image based on the detection result and obtaining the determination result, the above method also includes: when the determination result indicates that the target image contains tampering information, generating an alarm information based on the determination result and the target image; sending the alarm information to the target object, wherein the target object processes the target image based on the alarm information.
[0126] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0127] The present application is described with reference to the flow chart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.
[0128] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0129] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0130] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0131] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0132] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0133] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0134] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0135] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A method for determining tampering information of a face image, comprising: Obtain a face image of a person whose information is to be tampered with and obtain a target image; Extracting feature information from the target image to obtain a feature information set, wherein the feature information at least includes the following feature information: background feature information, face feature information, edge feature information, and pixel difference information; Inputting the feature information of the feature information set into a target detection model for detection to obtain a detection result, wherein the target detection model refers to a model obtained by building a generative adversarial network based on N learning tasks and training the generative adversarial network using the feature information set, and N is a positive integer; The tampering information of the target image is judged according to the detection result to obtain a judgment result.
2. The method according to claim 1, wherein: Extracting feature information from the target image to obtain a feature information set includes: Extracting background feature information from the target image using a first segmentation model to obtain a first image; Extracting facial feature information in the target image using a detection model to obtain a second image, and extracting facial edge information in the target image to obtain a third image; Performing edge blurring processing on the third image, and extracting edge feature information of the processed third image using an enhanced edge model to obtain a fourth image; extracting pixel difference information in the target image using a selective reconstruction method filter to obtain a fifth image; The first image, the second image, the fourth image and the fifth image are combined to obtain the feature information set.
3. The method according to claim 1, wherein: The target detection model is obtained by the following steps: Collect images containing facial information to obtain a training set; Extracting the feature information from each image included in the training set to obtain a feature information set for each image; Constructing branches corresponding to the N learning tasks respectively according to the feature information and the generative adversarial network to obtain N branches; Constructing a loss function corresponding to each of the N branches, and fusing the loss functions corresponding to each branch to obtain a target loss function; Each image in the training set and the feature information set of each image are input into the N branches for iterative calculation until the target loss function converges to obtain the target detection model.
4. The method according to claim 3, wherein: According to the feature information and the generative adversarial network, branches corresponding to the N learning tasks are respectively constructed, and the obtained N branches include: Constructing a first branch based on the generative adversarial network, wherein the first branch is used to identify tampering information of a facial image; Constructing a second branch based on the detection model and the second segmentation model, wherein the second branch is used to generate a face edge image and detect face information in the face image; Constructing a third branch based on an image detection algorithm, wherein the third branch is used to detect tampering information of the face edge image; The first branch, the second branch and the third branch are determined as the N branches.
5. The method according to claim 4, wherein: Constructing a first branch based on the generative adversarial network includes: Constructing a first generator, a second generator, a third generator and a fourth generator according to the extraction method of each feature information in the feature information, wherein the first generator is used to extract the background feature information, the second generator is used to extract the face feature information, the third generator is used to extract the edge feature information, and the fourth generator is used to extract the pixel difference information; Using the first generator, the second generator, the third generator, and the fourth generator to extract the feature information of each image in the training set, to obtain a feature information set of each image; Constructing a discriminator based on the generative adversarial network, wherein the discriminator is used to determine tampering information of a face image; The first branch is constructed based on the first generator, the second generator, the third generator, the fourth generator and the discriminator, wherein the first branch is used to identify the tampering information of each image in the training set and to identify the tampering information of each image in the feature information set of each image.
6. The method according to claim 4, wherein: Construct the loss function corresponding to each branch of the N branches, and merge the loss functions corresponding to each branch to obtain the target loss function including: constructing a first loss function based on the difference between the predicted value and the actual value of each image in the first branch; Constructing a second loss function corresponding to the second branch according to a preset loss function of the image segmentation algorithm; constructing a third loss function based on the difference between the predicted value and the actual value of each edge image in the third branch; Configure the weight of the first loss function, the weight of the second loss function, and the weight of the third loss function; The target loss function is constructed according to the first loss function, the second loss function, the third loss function, the weight of the first loss function, the weight of the second loss function and the weight of the third loss function.
7. The method according to claim 1, further comprising: After determining the tampering information of the target image based on the detection result and obtaining the determination result, If the discrimination result indicates that the target image contains the tampering information, generating warning information according to the discrimination result and the target image; The warning information is sent to a target object, wherein the target object processes the target image according to the warning information.
8. A device for determining whether a face image has been tampered with, comprising: An acquisition unit, used to acquire a face image of a person whose information is to be detected for tampering, and obtain a target image; An extraction unit, used to extract feature information from the target image to obtain a feature information set, wherein the feature information at least includes the following feature information: background feature information, face feature information, edge feature information, and pixel difference information; A detection unit, configured to input the feature information of the feature information set into a target detection model for detection to obtain a detection result, wherein the target detection model refers to a model obtained by constructing a generative adversarial network based on N learning tasks and training the generative adversarial network using the feature information set, and N is a positive integer; The distinguishing unit is used to distinguish the tampering information of the target image according to the detection result to obtain a distinguishing result.
9. A computer-readable storage medium comprising a stored computer program, wherein: When the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the method for determining facial image tampering information as described in any one of claims 1 to 7.
10. An electronic device comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein: When the one or more programs are executed by the one or more processors, the one or more processors implement the method for determining facial image tampering information as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Incremental learning method and system for detecting fake face image video
CN116310754A
Method and device for judging tampered information of face image and electronic equipment
CN117690173A
Heterogenous Face Recognition System and Method
US20230306732A1
Cited By
Image tampering identification method and device and medium
CN121280738A