A method, system and medium for face authentication based on image patch order loss

By cutting, disrupting and enhancing the face pseudo-recognition data set, and designing multiple loss functions to jointly train the neural network, the problem of high cost of face pseudo-recognition training in the existing technology is solved, and efficient and accurate pseudo-recognition effect is achieved.

CN115147908BActive Publication Date: 2025-05-06XIAMEN MEIYA PICO INFORMATION CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210879588.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-25
Publication Date
2025-05-06
Estimated Expiration
2042-07-25

AI Technical Summary

Technical Problem

The prior art has problems with high training costs and data costs in facial false recognition, which is difficult to effectively reduce and improve the accuracy of the judgment.

Method used

A face pseudo-identification method based on image patch sequence loss is proposed. By uniformly cutting, disrupting and enhancing the data set, multiple loss functions (including cross entropy loss function and binary cross entropy function) are designed to jointly train the neural network classification model.

Benefits of technology

It greatly reduces the training cost and data collection cost, enhances discrimination accuracy, and does not rely on network structure. It can theoretically adapt to any algorithm framework and improves generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147908B_ABST
    Figure CN115147908B_ABST
Patent Text Reader

Abstract

The present application proposes a face authentication method based on image patch sequence loss, including: S1, obtaining a portrait authentication data set containing training images and their label information; S2, performing data processing on the portrait authentication data set, including: uniformly cutting the training images and disrupting the initial patch sequence set of the training images to obtain the first patch sequence set of the training images; enhancing the first patch sequence set to obtain the second patch sequence set; randomly exchanging the patch blocks in the first patch sequence set and the second patch sequence set of different training images, thereby obtaining a processed portrait authentication data set; S3, using the processed portrait authentication data set to train a neural network classification model; S4, inputting the image to be detected into the neural network classification model. The portrait authentication method of the present application greatly saves training costs and data collection costs, and has a high discrimination accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of AI security, and in particular to a method, system and medium for face authentication based on image patch order loss. Background Art

[0002] Deepfake portrait discrimination is the most important algorithm module in image authentication. It is difficult to train, difficult to explain, and has poor generalization ability, which brings great troubles to relevant practitioners. After research and analysis, it is believed that the main reasons for the above difficulties are the prominent semantics of the face subject, the poor data set construction method, and the unclear training objectives, which lead to considerable training and data costs.

[0003] In the prior art, a Chinese patent with application number 202010097995.3 discloses a face-changing detection method, apparatus, device and storage medium, which is essentially an enhanced supervised neural network prediction method that introduces joint constraints on feature combinations and enhances the ability to capture and refine relevant features. However, it still cannot get rid of the problem of excessive training costs.

[0004] The Chinese patent with application number 202010614929.9 discloses a detection method, device, equipment and computer storage medium. It is only an algorithm combination application process. It determines whether it is necessary to further perform face authentication by pre-detecting the human body information in the picture. In principle, it can speed up the reasoning efficiency and reduce false detections, but in practice it increases the reasoning process instead. The final result is also affected by the superposition of algorithm errors at different stages, and the effect is not ideal.

[0005] A Chinese patent with application number 202010806564.X discloses an image detection method, apparatus, device and computer-readable storage medium, which extracts two modalities, namely RGB features and specific high-frequency features, and fuses and enhances the two features to identify fake faces. This algorithm mode is relatively common, the overall framework is large and complex, and lacks practical significance.

[0006] In summary, the existing technology still cannot solve the problem of high training cost and data cost of face authentication. Therefore, it is particularly important to provide a face authentication method that saves training cost and data collection cost and has high discrimination accuracy. Summary of the invention

[0007] In order to solve the above technical problems, the present application proposes a face authentication method, system and medium based on image patch order loss.

[0008] According to the first aspect of the present application, a face authentication method based on image patch order loss is proposed, comprising the following steps:

[0009] S1. Obtain a portrait authentication dataset including training images and label information thereof, wherein the training images at least include a face area;

[0010] S2. Processing the portrait authentication dataset, wherein the data processing includes:

[0011] Uniformly cutting the training image and disrupting the initial patch sequence set of the training image to obtain a first patch sequence set of the training image;

[0012] Performing enhancement processing on the first patch sequence set to obtain a second patch sequence set;

[0013] Randomly exchanging the patch blocks in the first patch sequence set and the second patch sequence set of different training images, so as to obtain a processed portrait counterfeit detection data set;

[0014] S3, using the processed portrait counterfeit detection data set to train a neural network classification model; and

[0015] S4. Input the image to be detected into the neural network classification model to obtain the identification result.

[0016] Preferably, the initial patch sequence set of the training image disrupted in step S2 specifically includes: selecting a first patch block from the initial patch sequence set, and finding a second patch block with the largest similarity distance between the remaining patch blocks in the initial patch sequence set and the first patch block, and exchanging the first patch block with the second patch block; repeating the above operation until all patch blocks have completed an exchange.

[0017] Preferably, the step S2 of enhancing the first patch sequence set specifically includes:

[0018] Performing random rotation, brightness adjustment, color difference adjustment, compression quality change, and grayscale processing on the patch blocks in the first patch sequence set;

[0019] The patch blocks in the first patch sequence set are mixed in pairs according to a preset ratio, and the order and main semantics of one of the patch blocks are retained.

[0020] Preferably, the portrait authentication data set in step S1 includes at least five data sets: a full image generation set, a face exchange set, an expression reenactment set, a beauty set and a face attribute editing set, wherein the ratio of forged images to natural images is 1:1.

[0021] Preferably, the step S3 of using the processed data set to train a neural network classification model specifically includes:

[0022] Obtaining a first loss function, a second loss function, and a third loss function corresponding to the processed training image and the processed label information, wherein the first loss function is used to characterize the loss generated by the processed training image and the processed label information, the second loss function is used to characterize the loss between the patch sequence prediction and the true sequence in the processed training image, and the third loss function is used to determine the attribution of the patch block in the processed training image;

[0023] The neural network classification model is jointly trained according to the first loss function, the second loss function and the third loss function to adjust the parameters of the neural network classification model.

[0024] Preferably, the first loss function and the second loss function are both composed of a cross entropy loss function, and the first loss function and the second loss function are jointly defined as:

[0025]

[0026] Among them, S is the forged image type index set, N p is the sample index subset capacity involved in each forgery method, L c is the first loss function, h is the convolutional neural network, and is the sample before and after the patch block is shuffled, θ f ,θ c With θ d These are the parameters of each part of the convolutional neural network. is the image content label, K p is the sample index subset capacity generated after the patch block is shuffled, α is the weight, L d is the second loss function, is the true sequence label.

[0027] Preferably, the third loss function is a binary cross entropy function, and the formal definition of the third loss function is:

[0028]

[0029] Among them, y i is the true label, Labels obtained by convolutional neural network inference.

[0030] Preferably, the similarity distance between two patch blocks is calculated by an image similarity comparison model, and the formal definition of the image similarity comparison model is:

[0031]

[0032] Among them, G is a deep convolutional neural network, θ is a set of model parameters, and I p For the RGB content of the image, the Euclidean distance is used to compare the CNN linear layer features.

[0033] According to the second aspect of the present application, a face authentication system based on image patch order loss is proposed, comprising:

[0034] A sample acquisition module is configured to acquire a portrait authentication dataset containing training images and label information thereof, wherein the training images at least include a face area;

[0035] A sample processing module is configured to perform data processing on the portrait counterfeit detection data set, wherein the data processing includes: uniformly cutting the training image and disrupting the initial patch sequence set of the training image to obtain a first patch sequence set of the training image; performing enhancement processing on the first patch sequence set to obtain a second patch sequence set; and randomly exchanging patch blocks in the first patch sequence set and the second patch sequence set of different training images to obtain a processed portrait counterfeit detection data set;

[0036] A training module configured to train a neural network classification model using the processed portrait counterfeit detection data set;

[0037] The authentication module is configured to input the image to be detected into the neural network classification model to obtain an authentication result.

[0038] According to a third aspect of the present application, a computer-readable storage medium is proposed, which stores a computer program. When the computer program is executed by a processor, the face authentication method based on image patch order loss as described in the first aspect of the present application is implemented.

[0039] The present application proposes a face authentication method, system and medium based on image patch sequence loss. The present application realizes deep fake portrait detection from the perspectives of self-supervised learning and image subject semantic destruction for the first time. By constructing and processing the data set and designing a combination of multiple loss functions, especially designing the second loss function and the third loss function to predict the patch sequence and patch block attribution in the image, the training cost and data collection cost are greatly reduced, and the discrimination accuracy is enhanced. In addition, the present application does not depend on the network structure and can theoretically adapt to any algorithm framework, thereby improving the generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The accompanying drawings illustrate the embodiments and are used together with the description to explain the principles of the present application. It will be easy to recognize other embodiments and many expected advantages of the embodiments because they become better understood by reference to the following detailed description. The elements of the drawings are not necessarily to scale with each other. The same reference numerals refer to corresponding similar parts.

[0041] Figure 1 is a flow chart of a face authentication method based on image patch order loss according to an embodiment of the present application;

[0042] Figure 2 It is a block diagram of a face authentication system based on image patch order loss according to an embodiment of the present application.

[0043] Explanation of the accompanying drawings: 1. Sample acquisition module; 2. Sample processing module; 3. Training module; 4. Authentication module. DETAILED DESCRIPTION

[0044] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and Examples. It should be understood that the specific embodiments described herein are only configured to explain the present application and are not configured to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.

[0045] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0046] According to the first aspect of the present application, a face authentication method based on image patch order loss is proposed. Figure 1 FIG. 4 shows a flow chart of a face authentication method based on image patch order loss according to an embodiment of the present application. Figure 1 As shown, the method comprises the following steps:

[0047] S1. Obtain a portrait authentication dataset including training images and label information thereof, wherein the training images at least include a face area.

[0048] In a specific embodiment, a portrait authentication dataset is prepared. Where N p >=20000,p=1,...,5,N p represents the capacity of various forged sample datasets, that is, there are 5 datasets here, namely, full image generation set, face exchange set, expression replay set, beauty set and face attribute editing set. The total amount of forged images and natural images in the entire dataset each account for 50%, x i is a 224×224 RGB training image, y i For its label information, y i ∈{0, 1}.

[0049] S2. Process the portrait authentication dataset. The data processing includes:

[0050] The training image is evenly cut and the initial patch sequence set of the training image is disrupted to obtain the first patch sequence set of the training image;

[0051] Performing enhancement processing on the first patch sequence set to obtain a second patch sequence set;

[0052] The patch blocks in the first patch sequence set and the second patch sequence set of different training images are randomly exchanged to obtain a processed portrait counterfeit detection dataset.

[0053] In a specific embodiment, the training image is first cut into N blocks of size w×w. Assuming that the initial patch sequence set of the training image is U, then U contains a total of 9×N! elements. Then the initial patch sequence set is shuffled. The specific method is: uniformly sample and select a first patch block from the initial patch sequence set, and find the second patch block with the largest similarity distance between the remaining patch blocks in the initial patch sequence set and the first patch block, and then exchange the first patch block with the second patch block; repeat the above operation until all patch blocks in the initial patch sequence set have completed an exchange, and finally obtain the first patch sequence set U~1. Among them, the similarity distance between two patch blocks can be calculated by preparing or training an image similarity comparison model. The formal definition of the image similarity comparison model is:

[0054]

[0055] Among them, G is a deep convolutional neural network, θ is a set of model parameters, and I p For the RGB content of the image, the Euclidean distance is used to compare the CNN linear layer features.

[0056] In a specific embodiment, the enhancement processing of the first patch sequence set U~1 mainly includes two methods. The first method is: randomly rotating, adjusting brightness, adjusting color difference, changing compression quality, and graying the patch blocks in the first patch sequence set; the second method is: superimposing and perturbing the patch blocks in the first patch sequence set in pairs, mixing them according to a preset ratio, and retaining the order and main semantics of one of the patch blocks, for example, mixing patch block A and patch block B at a ratio of 90% and 10%, and retaining the order and main semantics of patch block A. Finally, the first patch sequence set U~1 is enhanced to obtain the second patch sequence set U~2.

[0057] In a specific embodiment, according to obtaining the first patch sequence set U~1 and the second patch sequence set U~2, the patch blocks in the first patch sequence set U~1 and the second patch sequence set U~2 of different training images are randomly exchanged, and the exchanged U~1 is retained as the final patch sequence set of the training image. Among them, the final patch sequence set The formal expression is:

[0058]

[0059] Among them, US means uniform sampling, and change means exchanging the pixel content of the patch block. After the exchange and mixing, the corresponding label of each patch needs to be recorded. The label information of the patch block that originally belonged to the image can be set to 1, and the rest can be set to 0. For example, y = [1, 1, 0, 0, 0, 0, 01] means that the 1st, 2nd and 9th in the mixed training image belong to the original content of this image block.

[0060] S3. Use the processed portrait authentication dataset to train a neural network classification model.

[0061] In a specific embodiment, after data processing of a portrait counterfeit detection data set, processed training images and processed label information are obtained, and the processed training images and processed label information are used to design a first loss function, a second loss function, and a third loss function. The neural network classification model is jointly trained by these three loss functions to adjust the parameters of the neural network classification model.

[0062] In a specific embodiment, the first loss function is used to characterize the loss generated by the processed training image and the processed label information, and the second loss function is used to characterize the loss between the patch sequence prediction and the true sequence in the processed training image. The first loss function and the second loss function are both composed of a cross entropy loss function, and the joint definition of the first loss function and the second loss function is:

[0063]

[0064] Among them, S is the forged image type index set, N p is the sample index subset capacity involved in each forgery method, L c is the first loss function, h is the convolutional neural network, and is the sample before and after the patch block is shuffled, θ f ,θ c With θ d These are the parameters of each part of the convolutional neural network. is the image content label, K p is the sample index subset capacity generated after the patch block is shuffled, α is the weight, L d is the second loss function, is the true sequence label.

[0065] In a specific embodiment, the third loss function is used to determine the attribution of the patch blocks in the processed training image, which is equivalent to a multi-label classification problem. The third loss function is a standard binary cross entropy function, and the formal definition of the third loss function is:

[0066]

[0067] Among them, y i is the true label, Labels obtained by convolutional neural network inference.

[0068] S4. Input the image to be detected into the neural network classification model to obtain the identification result.

[0069] In a specific embodiment, a ResNet50 neural network is selected as the classification model, the weight decay is set to 0.0001, the batch size is 256, the initial learning rate is 0.02, and a total of 20 epochs are completed. After the training is completed, the image to be detected is input into the trained neural network classification model to obtain the identification result.

[0070] In summary, the present application provides a face authentication method based on image patch sequence loss, and its beneficial effects are: for the first time, portrait deep fake detection is realized from the perspectives of self-supervised learning and image subject semantic destruction, wherein the construction of the data set, the processing of the data set, and the design of multiple loss functions are all innovative designs, especially the design of the second loss function and the third loss function to predict the patch sequence in the image and the patch block attribution prediction, which greatly reduces the training cost and data collection cost, and enhances the discrimination accuracy, and the present application does not depend on the network structure, and theoretically can be adapted to any algorithm framework, thereby improving the generalization ability.

[0071] According to the second aspect of the present application, a face authentication system based on image patch order loss is proposed. The face authentication system is built based on the above-mentioned face authentication method. Figure 2 FIG. 4 shows a block diagram of a face authentication system based on image patch order loss according to an embodiment of the present application. Figure 2 As shown, the system includes:

[0072] The sample acquisition module 1 is configured to acquire a portrait authentication dataset including training images and label information thereof, wherein the training images at least include a face area;

[0073] The sample processing module 2 is configured to perform data processing on the portrait counterfeit detection data set, and the data processing includes: uniformly cutting the training image and disrupting the initial patch sequence set of the training image to obtain a first patch sequence set of the training image; performing enhancement processing on the first patch sequence set to obtain a second patch sequence set; randomly exchanging the patch blocks in the first patch sequence set and the second patch sequence set of different training images, so as to obtain a processed portrait counterfeit detection data set;

[0074] A training module 3 is configured to train a neural network classification model using the processed portrait counterfeit detection data set;

[0075] The authentication module 4 is configured to input the image to be detected into the neural network classification model to obtain the authentication result.

[0076] According to a third aspect of the present application, a computer-readable storage medium is proposed, which stores a computer program. When the computer program is executed by a processor, it implements the face authentication method based on image patch order loss as described in the first aspect of the present application.

[0077] In the embodiments of the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device / system / method embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0078] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0079] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0080] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.

[0081] Obviously, those skilled in the art can make various modifications and changes to the embodiments of the present application without departing from the spirit and scope of the present application. In this way, if these modifications and changes are within the scope of the claims of the present application and their equivalents, the present application is also intended to cover these modifications and changes. The word "comprising" does not exclude the presence of other elements or steps not listed in the claims. The simple fact that certain measures are recorded in mutually different dependent claims does not indicate that the combination of these measures cannot be used to profit. Any figure mark in the claims should not be considered to limit the scope.

Claims

1. A face authentication method based on image patch order loss, characterized in that: The following steps are involved: S1. Obtain a portrait authentication dataset including training images and label information thereof, wherein the training images at least include a face area; S2. Processing the portrait authentication dataset, wherein the data processing includes: Uniformly cutting the training image and disrupting the initial patch sequence set of the training image to obtain a first patch sequence set of the training image; Performing enhancement processing on the first patch sequence set to obtain a second patch sequence set; Randomly exchanging the patch blocks in the first patch sequence set and the second patch sequence set of different training images, so as to obtain a processed portrait counterfeit detection data set; S3, using the processed portrait counterfeit detection data set to train a neural network classification model; and S4, inputting the image to be detected into the neural network classification model to obtain an identification result; The step S3 of using the processed data set to train a neural network classification model specifically includes: Obtaining a first loss function, a second loss function, and a third loss function corresponding to the processed training image and the processed label information, wherein the first loss function is used to characterize the loss generated by the processed training image and the processed label information, the second loss function is used to characterize the loss between the patch sequence prediction and the true sequence in the processed training image, and the third loss function is used to determine the attribution of the patch block in the processed training image; The neural network classification model is jointly trained according to the first loss function, the second loss function and the third loss function to adjust the parameters of the neural network classification model.

2. The method according to claim 1, characterized in that: The step S2 of disrupting the initial patch sequence set of the training image specifically includes: selecting a first patch block from the initial patch sequence set, and finding a second patch block with the largest similarity distance between the remaining patch blocks in the initial patch sequence set and the first patch block, and exchanging the first patch block with the second patch block; repeating the above operation until all patch blocks have completed an exchange.

3. The method according to claim 1, characterized in that The step S2 of enhancing the first patch sequence set specifically includes: Performing random rotation, brightness adjustment, color difference adjustment, compression quality change, and grayscale processing on the patch blocks in the first patch sequence set; The patch blocks in the first patch sequence set are mixed in pairs according to a preset ratio, and the order and main semantics of one of the patch blocks are retained.

4. The method according to claim 1, characterized in that: The portrait authentication dataset in step S1 includes at least five datasets: a full image generation set, a face exchange set, an expression replay set, a beauty set, and a face attribute editing set, wherein the ratio of forged images to natural images is 1:

1.

5. The method according to claim 1, characterized in that The first loss function and the second loss function are both composed of a cross entropy loss function, and the first loss function and the second loss function are jointly defined as: Among them, S is the forged image type index set, N p is the sample index subset capacity involved in each forgery method, L c is the first loss function, h is the convolutional neural network, and is the sample before and after the patch block is shuffled, θ f ,θ c With θ d These are the parameters of each part of the convolutional neural network. is the image content label, K p is the sample index subset capacity generated after the patch block is shuffled, α is the weight, L d is the second loss function, is the true sequence label.

6. The method according to claim 1, characterized in that The third loss function is a binary cross entropy function, and the formal definition of the third loss function is: Among them, y i is the true label, Labels obtained by convolutional neural network inference.

7. The method according to claim 2, characterized in that The similarity distance between two patches is calculated by the image similarity comparison model, which is formally defined as: Among them, G is a deep convolutional neural network, θ is a set of model parameters, and I p For the RGB content of the image, the Euclidean distance is used to compare the CNN linear layer features.

8. A face authentication system based on image patch order loss, characterized in that: include: A sample acquisition module is configured to acquire a portrait authentication dataset containing training images and label information thereof, wherein the training images at least include a face area; A sample processing module is configured to perform data processing on the portrait counterfeit detection data set, wherein the data processing includes: uniformly cutting the training image and disrupting the initial patch sequence set of the training image to obtain a first patch sequence set of the training image; performing enhancement processing on the first patch sequence set to obtain a second patch sequence set; and randomly exchanging patch blocks in the first patch sequence set and the second patch sequence set of different training images to obtain a processed portrait counterfeit detection data set; A training module is configured to use the processed portrait counterfeit detection data set to train a neural network classification model, specifically comprising: obtaining a first loss function, a second loss function, and a third loss function corresponding to the processed training image and the processed label information, wherein the first loss function is used to characterize the loss generated by the processed training image and the processed label information, the second loss function is used to characterize the loss between the patch sequence prediction and the real sequence in the processed training image, and the third loss function is used to judge the attribution of the patch block in the processed training image; jointly training the neural network classification model according to the first loss function, the second loss function, and the third loss function to adjust the parameters of the neural network classification model; The authentication module is configured to input the image to be detected into the neural network classification model to obtain an authentication result.

9. A computer-readable storage medium storing a computer program, wherein the computer program implements the method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • Face-swapping detection methods, devices, equipment and storage media

    CN111353392B

  • Detection methods, apparatus, equipment and computer storage media

    CN111783644B

  • Image detection method, device and equipment and computer readable storage medium

    CN112767303A

  • Recommended product illustration method and device and electronic equipment

    CN107330750A

  • Face image processing method and device, computer equipment and storage medium

    CN111768336A