Face forgery detection method and system based on decentralized federated learning

By adopting an adaptive decentralized federated learning architecture in face forgery detection and combining CNN and Transformer models, the problem of single point failure and client computing power heterogeneity in centralized federated learning structure is solved, and the robustness and accuracy of the detection system are improved.

CN119942615APending Publication Date: 2025-05-06CHONGQING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510036507.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the face forgery detection, the existing technology has the problem of single point failure risk caused by centralized federated learning structure and slow model convergence speed caused by client computing power heterogeneity. At the same time, traditional CNN models lack global perception capabilities, which affects detection accuracy.

Method used

Adaptive decentralized federated learning architecture is adopted to eliminate the dependence of central servers and accelerate the convergence of the model through collaborative learning between distributed nodes and adaptive network topology. Combining the local perception advantages of convolutional neural networks and the global perception ability of Transformer, the fine-grained local features and global dependencies of face images are extracted.

Benefits of technology

It improves the robustness and detection effect of the face forgery detection system, enhances the stability and fault tolerance of the system, and can continue to operate normally in the event of node failure or network problems and speed up the convergence speed of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942615A_ABST
    Figure CN119942615A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision and machine learning, and particularly relates to a face forgery detection method and system for decentralized federated learning. The method comprises the following steps: receiving face image data subjected to data preprocessing, and training and updating parameters of a face counterfeiting detection model by using the face image data; sending parameters of a self face counterfeiting detection model to a neighbor client directly connected with the decentralized network topology in the decentralized network topology; carrying out aggregation operation on the received parameters of the face forgery detection model to obtain updated parameters of the face forgery detection model; according to the completion time of each round of training, updating the decentralized network topology; and repeating the steps until the model reaches a preset termination condition. According to the method, the problems of single-point fault risk caused by dependence of a central server, low model convergence speed caused by heterogeneous computing capability of a client and insufficient local and global feature perception capability of the model in the existing face counterfeit image detection technology are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and machine learning, and specifically relates to a decentralized federated learning method and system for detecting face forgery. Background Art

[0002] With the rapid development of deepfake technology, the generated fake face images and videos are becoming more and more realistic, posing severe challenges to personal privacy, information security and social stability. Traditional face forgery detection methods usually use deep learning models, rely on centralized data training, and require a large amount of real and fake face data to be collected on the central server for model training and updating. However, centralized data processing brings the risk of data privacy leakage. To this end, for example, Chinese patent CN116524557A introduces a federated learning method, which allows face data to be retained on local devices for distributed training of the model to solve the privacy protection problem. However, the traditional centralized federated learning method relies on the central server to aggregate the local model. Once the server fails, the entire model framework will not be able to aggregate and update normally, resulting in system instability.

[0003] In addition, most existing face forgery detection models use the Convolutional Neural Network (CNN) architecture. CNN has advantages in extracting local features of images and can effectively capture detailed features in face images. However, due to the lack of global dependency modeling capabilities, CNN has difficulty in obtaining the overall context of images, which makes the model limited in identifying complex forged images. The Transformer model that has emerged in recent years has strong global perception capabilities and can capture the global dependencies of images through the self-attention mechanism, but it is not as good as CNN in processing fine-grained local features.

[0004] Therefore, there are two major problems in the existing technology when detecting forged face images: First, although the centralized federated learning structure solves the privacy problem of face data, the model update relies on the central server, which is prone to single point failure and affects the stability of the system. At the same time, the heterogeneity of computing power of different clients will affect the convergence speed of the model; second, although the traditional convolutional neural network can extract local features of the face, it lacks global perception ability and is difficult to fully capture the overall information of the image, thus affecting the accuracy of face forgery detection. Therefore, there is an urgent need for a new face forgery detection method that can solve the problems of single point failure of the server and heterogeneous computing power of the client, and combines the local perception of CNN and the global modeling ability of Transformer to improve the robustness and detection effect of the detection system. Summary of the invention

[0005] The present invention aims to overcome the above problems. The present invention proposes a decentralized federated learning face forgery detection method and system, which adopts an adaptive decentralized federated learning architecture, eliminates the dependence of the central server and accelerates the convergence of the model through collaborative learning between distributed nodes and adaptive network topology, thereby improving the fault tolerance and convergence speed of the face forgery detection system while protecting data privacy. At the same time, combining the local perception advantages of convolutional neural networks with the global perception capabilities of Transformer, the convolutional neural network is used to extract fine-grained local features of face images, and then the global dependency is captured through the self-attention mechanism of Transformer, thereby improving the model's discrimination effect on forged and real faces.

[0006] In a first aspect, the present invention provides a decentralized federated learning face forgery detection method, which is applied to a client and includes the following steps:

[0007] The client receives the face image data after data preprocessing, and uses the face image data to train and update the parameters of the face forgery detection model;

[0008] The client sends the parameters of its own face forgery detection model to the neighboring clients directly connected to it in the decentralized network topology;

[0009] The client aggregates the received parameters of the face forgery detection model to obtain updated parameters of the face forgery detection model;

[0010] The client updates the decentralized network topology according to the completion time of each round of training;

[0011] The client repeats the above steps until the model reaches a preset termination condition, wherein the preset termination condition includes the number of communications with neighboring clients and the global loss function of the face forgery detection model reaching a preset threshold.

[0012] In a second aspect, the present invention provides a decentralized federated learning face forgery detection system, which is applied to a client and includes:

[0013] A data receiving module, used to receive the face image data after data preprocessing and receive the parameters of the face forgery detection model from the neighbor client in the decentralized network topology;

[0014] A data sending module, used to send the parameters of the client's own face forgery detection model;

[0015] A model training module, used to train and update the parameters of a face forgery detection model using the face image data;

[0016] A model aggregation module, used to aggregate the received parameters of the face forgery detection model to obtain updated parameters of the face forgery detection model;

[0017] A network update module, used to update the decentralized network topology according to the completion time of each round of training;

[0018] The training control module is used to determine whether the preset termination condition is met. If the preset termination condition is met, the above module is stopped from being called, otherwise the above module is repeatedly called; the preset termination condition includes the number of communications with neighbor clients and the global loss function of the face forgery detection model reaching a preset threshold.

[0019] Beneficial effects of the present invention:

[0020] The present invention overcomes the single point failure risk caused by the reliance on the central server in the existing human face forged image detection technology, the slow model convergence speed caused by the heterogeneous client computing power, and the insufficient model perception ability of local and global features, and provides a human face image forged detection method and system combining adaptive decentralized federated learning, convolutional neural network (ConvolutionalNeural Network) and Transformer. Adopting an adaptive decentralized federated learning architecture, while ensuring data privacy, it avoids the slow model convergence speed caused by the single point failure and the heterogeneous client computing power in the traditional centralized system, and the stability and fault tolerance of the system are improved, and it can continue to maintain normal operation and accelerate the convergence speed of the model in the case of node failure or network problems. In addition, the combination of convolutional neural network and Transformer model makes the system have stronger ability in local and global feature perception, which greatly improves the accuracy of human face forged image detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a schematic diagram of a decentralized federated learning framework according to an embodiment of the present invention;

[0022] Figure 2 is a flow chart of a method for detecting face forgery based on decentralized federated learning according to an embodiment of the present invention;

[0023] Figure 3 is a structural diagram of a face forgery detection model according to an embodiment of the present invention;

[0024] Figure 4 Schematic diagram of the EfficientNet network structure of the face forgery detection model of an embodiment of the present invention;

[0025] Figure 5 4 is a structural diagram of a dual-pooling SE module of a face forgery detection model according to an embodiment of the present invention;

[0026] Figure 6 is a schematic diagram of the ViT network structure of the face forgery detection model of an embodiment of the present invention;

[0027] Figure 7 Schematic diagram of a decentralized federated learning face forgery detection system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0029] Centralized Federated Learning (FL) relies on a central server. Each client uploads the model parameters of the local training to the central server, which aggregates these parameters and then sends the updated model to the client. This architecture makes the central server the bottleneck and single point of failure of the entire system. Decentralized Federated Learning (DFL) is a new distributed machine learning framework that abandons the role of the central server and allows participating clients to communicate with each other directly through the network topology and work together to complete model training. It aims to solve the single point dependency problem in traditional centralized federated learning, while further improving data privacy and system robustness.

[0030] Figure 1 Schematic diagram of a decentralized federated learning framework according to an embodiment of the present invention. Figure 1 As shown in the figure, the decentralized federated learning framework includes a set of working nodes V consisting of n clients, V = {v1, v2…, v n}, each client node v i Use its local dataset D i Train a local modelω i .

[0031] Based on the above decentralized federated learning framework, Figure 2 A flow chart of a face forgery detection method based on decentralized federated learning according to an embodiment of the present invention is given. Figure 2 As shown, applied to a client, the method includes:

[0032] 101. The client receives the face image data after data preprocessing, and uses the face image data to train and update the parameters of the face forgery detection model;

[0033] In an embodiment of the present invention, the client can obtain facial image data from the data center, and the data center can randomly divide the data set into multiple subsets and distribute them to each client, ensuring that each client has different local data to support subsequent decentralized federated learning model training. In addition, the data center can also send the data set to each client's face forgery detection model in a distributed manner to train the model separately to achieve the set local training rounds.

[0034] The facial image data can be a training set using the FaceForensics++ dataset. FaceForensics++ is a facial forgery video dataset that uses four automatic face-changing algorithms, Face2Face, FaceSwap, DeepFakes, and NeuralTextures, to process 1,000 real video sequences. The video is processed for key frames and facial images, and 20,000 processed images are created for each method and the photo size is reshaped to 224*224. Generate real and forged facial images with a balanced number of samples to provide data support for the subsequent training of facial forgery detection models. The data preprocessing operations include normalization, cropping, denoising, etc., which can make the original facial image data more suitable for model training.

[0035] The embodiment of the present invention can also divide the face image data into a training set, a validation set and a test set, for example, in a ratio of 8:1:1. The training set accounts for 80% and is used for model training; the validation set accounts for 10% and is used to monitor the training effect and prevent overfitting; the test set accounts for 10% and is used to evaluate the final performance of the model.

[0036] In the embodiment of the present invention, during the local training phase, each client device receives the face image data (including forged and real face images) that has been pre-processed as input, and trains a face forgery detection model that is based on the combination of EfficientNet and ViT (Vision Transformer), such as Figure 3 As shown, the face forgery detection model includes EfficientNet network, ViT network, global average pooling, multi-layer perceptron and feature classifier. The specific training process of the face forgery detection model is as follows:

[0037] 111. Extract local features of the face image data using an EfficientNet network based on a double-pooled SE module;

[0038] like Figure 4As shown, in this embodiment, the facial image data is subjected to initial feature extraction through a convolutional layer to obtain an initial forged feature image; the initial forged feature image is subjected to deeper feature extraction through multiple MBConv modules including double-pooled SE modules to obtain a local deep-level forged feature image, and the local deep-level forged feature image indicates the local features of the facial image data.

[0039] This embodiment uses the EfficientNet network to extract local feature maps of the input face image and capture the fine-grained features of the face. The model has strong representation capabilities and can extract the slight differences between forged and real face images.

[0040] The input image first passes through a standard convolutional layer (called the Stem layer) to generate a preliminary feature map. This convolutional layer uses a 3×3 convolution kernel to condense the information in the original image and extract preliminary low-level features.

[0041] After the Stem layer, the image features enter a hierarchical structure consisting of multiple MBConv modules stacked together. Each MBConv module gradually captures higher-level features through operations such as expansion, convolution, compression, and residual connection.

[0042] Specifically, Figure 5 As shown in the figure, each MBConv module first expands the number of channels of the feature map through a 1×1 convolution to capture information in a high-dimensional space; the expanded feature map is further extracted through a depthwise separable convolution. The depthwise convolution first performs a convolution operation on each channel independently, and the point-by-point convolution is responsible for integrating information across channels to reduce computational complexity; after the depthwise separable convolution, the double-pooling SE module is used to adjust the channel weights. The double-pooling SE module includes a "Squeeze" operation and an "Excitation" operation to help the network pay more attention to important features; finally, a 1×1 convolution is used to compress the number of channels back to their original size, and a residual connection is applied when the input and output sizes are the same.

[0043] Finally, the convolutional layer changes the dimension of the feature map to make it consistent with the dimension of subsequent ViT processing.

[0044] Through the above process, the network can extract local feature maps containing fine-grained features, providing rich detail information for the difference between forged and real faces.

[0045] In particular, the SE module is used in the MBconv module of the original EfficientNet network. The present invention uses a double-pooled SE module to adjust the weight of the channel. The specific operation is as follows:

[0046] In the Squeeze stage, the feature map U input to the double pooling SE module is subjected to the maximum pooling operation to extract important information in the feature map. Then a two-dimensional global average pooling is performed to compress the feature map U with dimensions W*H*C into a 1*1*C feature vector z. The formula is as follows:

[0047]

[0048] Among them, z c represents the two-dimensional global average pooling feature corresponding to each channel, which is the cth element of z; C represents the number of input channels; H represents the height of the input feature image in the spatial dimension; W represents the width of the input feature image in the spatial dimension; j represents the height index; i represents the width index; x c (i, j) represents the input feature value at the position (i, j) of the spatial dimension on channel C.

[0049] In the Excitation phase, first, a fully connected layer is used to compress C channels into C / r channels, then a RELU nonlinear activation layer is used, and then a fully connected layer is used to restore the number of channels to C channels. After Sigmoid activation, a new weight s is obtained, whose dimension is 1*1*C, which is used to describe the weight of the feature map of C channels in the feature map U. r refers to the compression ratio.

[0050] s=σ(W2δ(W1z)) (2)

[0051] Among them, δ represents the RELU activation function, σ represents the Sigmoid activation function, W1 represents the weight of the first fully connected layer, and W2 represents the weight of the second fully connected layer.

[0052] Finally, each channel in the feature map U is multiplied by the corresponding weight to obtain the final output of the double pooling SE module.

[0053] 112. Extracting global features of the facial image data using a vision-based Transformer network;

[0054] In an embodiment of the present invention, the local feature map is passed to a vision-based Transformer network, i.e., a ViT module. Figure 6 As shown in Figure 1, the ViT module further extracts global features from local feature maps to capture the global dependencies of the image. Through the self-attention mechanism, the ViT module can perform a deeper analysis of local feature maps to construct a global feature representation of the input image.

[0055] The output feature map of the above network is flattened, and each element is used as an input unit of ViT. This flattening operation preserves the spatial structure of local features and provides ViT with a high-dimensional representation of the image.

[0056] Then a learnable classification token is added before each element. This token is designed to interact with other element features through the self-attention mechanism to aggregate the global information of the image.

[0057] Since ViT's self-attention mechanism lacks position awareness, it is necessary to add position encoding to each element to ensure that ViT can capture the relative position of the element in the feature map. The formula for adding position encoding is as follows:

[0058]

[0059]

[0060] Where pos is the position of the element in the image, i is the dimension index, and d model It is the characteristic dimension of the token.

[0061] After multi-head self-attention processing by Transformer Encoder, the features of each element interact with other elements to form a global feature representation class token.

[0062] Among them, the self-attention mechanism captures the relationship between features through the similarity between query, key and value. The input feature vector is linearly transformed to generate Q, K, V matrices, and then the attention weight is calculated by the following formula:

[0063]

[0064] Among them, Q, K, and V are query, key, and value matrices respectively, and d K is the dimension of the key matrix.

[0065] In the multi-head self-attention mechanism, each attention head performs the above process independently. Specifically, the calculation formula for a single attention head is as follows:

[0066]

[0067] in, is the weight matrix that the model needs to learn.

[0068] Finally, the outputs of all attention heads are concatenated and the final feature representation is obtained through a linear transformation, which is calculated as follows:

[0069] MultiHead(Q,K,V)=Concat(head1,head2,...head n )W O (7)

[0070] Among them, W O is the learning parameter matrix used to perform a linear transformation on the concatenation result.

[0071] 113. Use a multi-layer perceptron to fuse local features and global features of the face image data to obtain feature representation of the face image data;

[0072] In the embodiment of the present invention, the local feature map extracted by the convolutional neural network is fused with the global feature representation extracted by ViT after global average pooling by a multilayer perceptron (MLP) to form a comprehensive feature representation. This fused feature contains both local detail information and global dependency, and can be used more accurately for authenticity recognition of face images.

[0073] 114. Classify the feature representation of the facial image data using a feature classifier, and predict whether the facial image data is a real face or a forged face;

[0074] In the embodiment of the present invention, the fused features are input into a classifier to determine the authenticity of the input image.

[0075] 115. The loss function of each client is calculated by the difference between the predicted classification value and the true label value of the face image data;

[0076] The classifier is trained using the cross entropy loss function to minimize the difference between the predicted value and the true label value. i , the local loss function is specifically expressed as follows:

[0077]

[0078] Where ξ is the dataset D i A data sample, L i (w i , ξ) represents the loss corresponding to sample ξ in this face forgery detection model, and E represents the mean. This local loss function represents the loss of the client v i Evaluation of face forgery detection models on local datasets.

[0079] According to the above analysis, for all clients, the optimization goal of decentralized federated learning is to minimize the average loss function of all clients, which can be expressed as:

[0080]

[0081] Where Ω represents the set of model parameters of all nodes.

[0082] 116. Use a stochastic gradient descent algorithm to minimize the loss function of each client, and train and update the parameters of the face forgery detection model.

[0083] In the local model update step, each node uses the Stochastic Gradient Descent (SGD) algorithm to update parameters, which is characterized as follows:

[0084]

[0085] in, Indicates client v i The model parameters of the local iteration t times, η is the learning rate.

[0086] The embodiment of the present invention trains the face forgery detection model locally on the client to avoid the privacy leakage risk that may be faced by face image data during transmission and when stored on a third-party server. Because the data does not need to be uploaded to an external server, all sensitive face data processing is performed locally on the client, which fundamentally reduces the possibility of data theft or abuse, and provides a solid guarantee for the security of user personal information.

[0087] 102. The client sends the parameters of its own face forgery detection model to the neighboring clients directly connected to it in the decentralized network topology;

[0088] In the embodiment of the present invention, an undirected graph G is used. t =(V t , E t ) represents the decentralized network topology at the tth communication iteration, where G t represents the network topology, V t is the set of working nodes, E is the set of communication links, representing the communication connections between nodes. i , define its neighbor node set as This includes the v i There is a direct communication link between the two nodes. i and v j The link between Indicates that there is a node v i and node v j The network topology uses the adjacency matrix Indicates that the elements is defined as follows:

[0089]

[0090] The embodiment of the present invention also defines a weight matrix C in the network topology diagram. The construction of the matrix C is used to describe the network topology structure in the DFL and determines the weighting method of the model parameters during communication between nodes. i,j Represents node v j In the model aggregation, for node v i The contribution of node v i and v j If there is no direct connection between i,j =0.

[0091] In the embodiment of the present invention, the weight matrix C can be defined in a variety of ways, such as using loss, delay, bandwidth, data similarity, etc., without limiting to a specific one.

[0092] Exemplarily, the weight matrix is ​​determined by the normalized value corresponding to the sum of the reciprocal of the loss function value of the client and the reciprocal of the loss function values ​​of each neighboring client. i,j The definition is based on the node v j The loss function value on the local dataset. j The loss function value is f i (ω i ), the inverse of the loss function is used as the basis for the weights:

[0093]

[0094] Among them, the loss value f i (w i ) is smaller, the client node v j The higher the contribution, the higher the contribution. In order to ensure the normalization of the weight matrix, the client node v i The weights of all directly connected neighbor client nodes are normalized.

[0095] 103. The client aggregates the received parameters of the face forgery detection model to obtain updated parameters of the face forgery detection model;

[0096] In an embodiment of the present invention, weighted aggregation is performed on the parameters of the face forgery detection models of neighboring clients received by the current client according to a weight matrix; the weight matrix is ​​used to indicate the contribution of the neighboring clients to the current client during the aggregation process.

[0097] 104. The client updates the decentralized network topology according to the completion time of each round of training;

[0098] In an embodiment of the present invention, the completion time of each round of training is determined by the client with the slowest completion time in the decentralized network topology; the completion time of the client is determined by the local computing time and the communication time; the local computing time is determined by the computing time of all iterations in the current round of training; the communication time is determined by the ratio of the memory occupied by the parameters of the face forgery detection model in the current round of training to the communication bandwidth of the link, and the link is the communication link between the current client and its directly connected neighbor clients in the current decentralized network topology.

[0099] In an embodiment of the present invention, the client updates the decentralized network topology according to the completion time of each round of training, including using a timeout mechanism to determine whether there is a client failure, initializing the current decentralized network topology graph to a complete graph including all available clients and communication links, and initializing the search step to the square root of the number of links in the current decentralized network topology; removing the link with the longest search step under the constraint of topological connectivity, updating the adjacency matrix, the neighbor client set, the weight matrix, and halving the search step and rounding it down; until the search step is 1, stopping the update to obtain an updated decentralized network topology.

[0100] For example, assume that each round of training consists of τ iterations, and the amount of computing resources available to each client is different, so the computing time of different clients is also different. i , the computation time of the kth round t′th iteration is expressed as Then in round k, client v i The computation time can be expressed as:

[0101]

[0102] The topology is rebuilt in each round based on the resources currently available on the client, so the neighbor set of each client is time-varying. i The neighbor set of After local training, each client v i Communicate with its neighbors, client v in round k i Communication time It is expressed as:

[0103]

[0104] Among them, s is the memory size occupied by the model parameters, Indicates link e i,j communication bandwidth.

[0105] The completion time of each link includes local computation time and communication time, which can be expressed as:

[0106]

[0107] The model is trained synchronously, so the completion time of each round is determined by the slowest client. The completion time of each round of training k is t k :

[0108]

[0109] in, Indicates client v i At the local iteration time of round k, Represents node v i In the kth round, its neighbor node v j communication time.

[0110] This embodiment calculates the time required for each link to complete according to the formula, uses the timeout mechanism to determine whether there is a client failure, initializes the current decentralized network topology graph to a complete graph containing all available clients and communication links, initializes the search step length stride to the square root of the number of links in the current topology, calculates the time required for each link to complete according to the formula, removes the stride longest time link under the constraint of topological connectivity, and updates the adjacency matrix A t , neighbor set Weight matrix C t , the search step is halved and rounded down. When stride = 1, stop updating and get the decentralized network topology with time constraints.

[0111] 105. The client repeats the above steps until the model reaches a preset termination condition, where the preset termination condition includes the number of communications with neighboring clients and the global loss function of the face forgery detection model reaching a preset threshold.

[0112] In an embodiment of the present invention, each round of training will update the decentralized network topology, and client nodes with failure risks will be eliminated through a timeout mechanism, so as to further optimize the operating efficiency and stability of the network and significantly improve the robustness and fault tolerance of the system.

[0113] It can be understood that this approach can protect data privacy while avoiding the slow model convergence problem caused by single point failures and client computing power heterogeneity in traditional centralized systems; through the decentralized network topology optimization of the present invention, the stability and fault tolerance of the detection method and system are improved, and it can continue to maintain normal operation and accelerate the convergence speed of the face forgery detection model in the event of node failure or network problems.

[0114] Figure 7 is a schematic diagram of a decentralized federated learning face forgery detection system according to an embodiment of the present invention. Figure 7As shown, the decentralized federated learning face forgery detection system is applied to the client and includes:

[0115] A data receiving module 201 is used to receive the face image data after data preprocessing and receive the parameters of the face forgery detection model from the neighbor client in the decentralized network topology;

[0116] A data sending module 202, used for sending the parameters of the face forgery detection model of the client itself;

[0117] A model training module 203, used to train and update parameters of a face forgery detection model using the face image data;

[0118] A model aggregation module 204 is used to aggregate the received parameters of the face forgery detection model to obtain updated parameters of the face forgery detection model;

[0119] A network update module 205, configured to update the decentralized network topology according to the completion time of each round of training;

[0120] The training control module 206 is used to determine whether the preset termination condition is met. If the preset termination condition is met, the calling of the above module is stopped, otherwise the above module is repeatedly called; the preset termination condition includes the number of communications with neighbor clients and the global loss function of the face forgery detection model reaching a preset threshold.

[0121] The embodiment of the present invention receives the face image data after data preprocessing and receives the parameters of the face forgery detection model from the neighbor client in the decentralized network topology through the data receiving module 201; sends the parameters of the face forgery detection model of the client itself through the data sending module 202; uses the face image data to train and update the parameters of the face forgery detection model through the model training module 203; aggregates the received parameters of the face forgery detection model through the model aggregation module 204 to obtain the updated parameters of the face forgery detection model; updates the decentralized network topology according to the completion time of each round of training through the network update module 205; and the training control module 206 repeatedly calls the above modules to eliminate the problem of slow model convergence caused by single point failures and heterogeneous client computing power as much as possible, thereby improving the parameter update efficiency of the face forgery detection model.

[0122] In a preferred embodiment of the present invention, the data receiving module 201 includes at least two data interfaces, the first data interface is used to receive facial image data after data preprocessing, and the second data interface is used to receive parameters of a face forgery detection model from a neighbor client in a decentralized network topology.

[0123] It should be noted that the client of the embodiment of the present invention may be a terminal or device with data processing function, and the terminal or device may be a backend server belonging to an enterprise, unit or organization. The terminal or device may also be a mobile device, an Internet of Things (IoT) device, a personal computer, etc. Taking mobile devices as an example, in federated learning, mobile devices can use local data to participate in the training of face forgery detection models, such as training personalized face forgery detection models to identify scenes or objects in face images, etc., while protecting the user's privacy data from being leaked.

[0124] In some embodiments, the face forgery detection model obtained by the decentralized federated learning face forgery detection method and system of the present invention can perform forgery detection on the preprocessed face forgery detection image to be detected, thereby obtaining a more accurate forgery detection result, which specifically includes the following steps.

[0125] A1: Receive the input forged image to be detected;

[0126] A2: extracting initial features from the forged image to be detected through a convolutional layer to obtain an initial forged feature image;

[0127] A3: The initial forged feature image is subjected to deeper feature extraction through multiple MBConv modules including double-pooled compression-excitation (SE) attention modules to obtain a local deep-level forged feature image;

[0128] A4: Flatten the deep fake feature image, add a learnable classification token, and then add a position code to obtain a new token for training and learning;

[0129] A5: The token is processed by Transformer Encoder with a multi-head self-attention mechanism to obtain a global feature representation class token;

[0130] A6: Perform global average pooling on the local deep-level forged feature image, and then perform fusion enhancement with the global feature representation classtoken to obtain a comprehensive forged feature;

[0131] A7: Determine a detection result of a forged face image based on the comprehensive forgery features.

[0132] This embodiment can effectively detect the recognition result of the face forgery detection image to be detected. This recognition result will clearly inform the user whether the detected face image is forged, providing a reliable decision-making basis for related security verification, identity recognition and other application scenarios, and effectively ensuring the security and stability of the system.

[0133] A person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware through a program, and the program can be stored in a computer-readable storage medium, which can include: ROM, RAM, disk or CD, etc.

[0134] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A decentralized federated learning method for face forgery detection, characterized in that: Applied to a client, the method comprises: The client receives the face image data after data preprocessing, and uses the face image data to train and update the parameters of the face forgery detection model; The client sends the parameters of its own face forgery detection model to the neighboring clients directly connected to it in the decentralized network topology; The client aggregates the received parameters of the face forgery detection model to obtain updated parameters of the face forgery detection model; The client updates the decentralized network topology according to the completion time of each round of training; The client repeats the above steps until the model reaches a preset termination condition, wherein the preset termination condition includes the number of communications with neighboring clients and the global loss function of the face forgery detection model reaching a preset threshold.

2. According to claim 1, a decentralized federated learning face forgery detection method is characterized in that: The parameters of the face forgery detection model trained and updated using the face image data include: Extracting local features of the face image data using an EfficientNet network based on a double-pooled SE module; Extracting global features of the facial image data using a vision-based Transformer network; Using a multi-layer perceptron to fuse local features and global features of the face image data to obtain a feature representation of the face image data; Using a feature classifier to classify the feature representation of the face image data, and predicting whether the face image data is a real face or a forged face; The loss function of each client is calculated by the difference between the predicted classification value and the true label value of the face image data; The stochastic gradient descent algorithm is used to minimize the loss function of each client, and the parameters of the face forgery detection model are trained and updated.

3. According to claim 2, a decentralized federated learning face forgery detection method is characterized in that: The method of extracting local features of the face image data using an EfficientNet network based on a double-pooled SE module includes: Performing initial feature extraction on the face image data through a convolutional layer to obtain an initial forged feature image; The initial forged feature image is subjected to deeper feature extraction through multiple MBConv modules including double-pooled SE modules to obtain a local deep-level forged feature image, wherein the local deep-level forged feature image indicates the local features of the face image data.

4. The decentralized federated learning face forgery detection method according to claim 1, characterized in that: The client aggregates the received parameters of the face forgery detection model to obtain updated parameters of the face forgery detection model, including weighted aggregation of the parameters of the face forgery detection models of neighboring clients received by the current client according to a weight matrix; the weight matrix is ​​used to indicate the contribution of the neighboring clients to the current client during the aggregation process.

5. According to claim 4, a decentralized federated learning face forgery detection method is characterized in that: The weight matrix is ​​determined by a normalized value corresponding to the reciprocal of the loss function value of the client and the sum of the reciprocals of the loss function values ​​of each neighboring client.

6. The decentralized federated learning face forgery detection method according to claim 1, characterized in that: The completion time of each round of training is determined by the client with the slowest completion time in the decentralized network topology; the completion time of the client is determined by the local computing time and the communication time; the local computing time is determined by the computing time of all iterations in the current round of training; the communication time is determined by the ratio of the memory occupied by the parameters of the face forgery detection model in the current round of training to the communication bandwidth of the link, and the link is the communication link between the current client and its directly connected neighbor clients in the current decentralized network topology.

7. A decentralized federated learning face forgery detection method according to claim 1 or 6, characterized in that: The client updates the decentralized network topology according to the completion time of each round of training, including using a timeout mechanism to determine whether there is a client failure, initializing the current decentralized network topology graph to a complete graph containing all available clients and communication links, and initializing the search step length to the square root of the number of links in the current decentralized network topology; removing the link with the longest search step length under the constraint of topological connectivity, updating the adjacency matrix, the neighbor client set, the weight matrix, and halving the search step length and rounding it down; When the search step is 1, the update stops and the updated decentralized network topology is obtained.

8. The decentralized federated learning face forgery detection method according to claim 1, characterized in that: The global loss function of the face forgery detection model is determined by the local loss function in the process of all clients training and updating the parameters of the face forgery detection model.

9. A decentralized federated learning face forgery detection system, characterized in that: Applied to the client, including: A data receiving module, used to receive the face image data after data preprocessing and receive the parameters of the face forgery detection model from the neighbor client in the decentralized network topology; A data sending module, used to send the parameters of the client's own face forgery detection model; A model training module, used to train and update the parameters of a face forgery detection model using the face image data; A model aggregation module, used to aggregate the received parameters of the face forgery detection model to obtain updated parameters of the face forgery detection model; A network update module, used to update the decentralized network topology according to the completion time of each round of training; The training control module is used to determine whether the preset termination condition is met. If the preset termination condition is met, the above module is stopped from being called, otherwise the above module is repeatedly called; the preset termination condition includes the number of communications with neighbor clients and the global loss function of the face forgery detection model reaching a preset threshold.

10. The decentralized federated learning face forgery detection system according to claim 9, characterized in that: The data receiving module includes at least two data interfaces, the first data interface is used to receive facial image data after data preprocessing, and the second data interface is used to receive parameters of a face forgery detection model from a neighbor client in a decentralized network topology.

Citation Information

Patent Citations

  • Face forgery detection model optimization method, device and system based on federated learning

    CN116524557A