Face forgery recognition method and device based on audio and image, equipment and medium

By combining multimodal information fusion method with audio and image data, the global and local features in face video are extracted and analyzed, and the problem of insufficient recognition accuracy of face forgery in the prior art is solved, and higher recognition accuracy and robustness are achieved.

CN120356074AActive Publication Date: 2025-07-22CHINA UNICOM WO MUSIC & CULTURE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510851017.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify face forgery, especially because the forged content is highly realistic, and it is difficult to distinguish between authenticity through image analysis alone.

Method used

Combining audio and image data, the global feature vector is extracted through the global feature encoder, and the image is cropped into the head, face and lip areas using the global area encoder. The area perception module generates feature weights, and the feature fusion module weights fusion features, which are finally recognized by the multi-layer perceptron classifier.

Benefits of technology

It improves the accuracy and robustness of facial forgery recognition, can capture the cross-modal relationship between images and audio, enhances the global perception of forgery behavior, reduces redundant information interference, and outputs accurate recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356074A_ABST
    Figure CN120356074A_ABST
Patent Text Reader

Abstract

The invention provides a face forgery recognition method, device and equipment based on audios and images and a medium, and relates to the technical field of face forgery recognition, and the method comprises the steps: extracting image data and audio data in to-be-recognized face video data, and constructing a global image; the global image is input into a trained face forgery recognition model, a face forgery recognition result is obtained, and the face forgery recognition model comprises a global feature encoder used for extracting global feature vectors of image data and audio data in the global image; the global region encoder is used for constructing a head region feature set, a face region feature set and a lip close-up feature set; the region sensing module is used for constructing a weight set; the feature fusion module is used for generating fusion features; and the multi-layer perceptron classifier is used for outputting a face forgery recognition result according to the fusion features. Face forgery recognition is performed from the audio angle and the image angle, and the recognition accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face forgery recognition, and in particular, to a method, device, equipment, and medium for face forgery recognition based on audio and images. Background Art

[0002] Face recognition is a biometric technology that automatically identifies or verifies a person's identity by analyzing and comparing digital images or video frames of a face. With the development of technology, face recognition has been widely applied in many fields such as security monitoring, access control, and mobile device unlocking. In various fields, for the purpose of deceiving or training face recognition technology, face forgery technology has emerged, and face forgery technology can use various technical means to create false face images or videos.

[0003] In related technologies, in order to identify whether a face is forged, mainly the face image is analyzed and recognized. However, since face forgery technology can generate very realistic face images, these forged contents are almost indistinguishable from real contents visually, making it difficult to accurately distinguish only through image analysis. Summary of the Invention

[0004] The problem solved by the present invention is how to improve the accuracy of face forgery recognition.

[0005] To solve the above problems, the present invention provides a method, device, equipment, and medium for face forgery recognition based on audio and images.

[0006] In a first aspect, the present invention provides a method for face forgery recognition based on audio and images, including: Extracting image data and audio data from the face video data to be recognized, and constructing a global image; Inputting the global image into a trained face forgery recognition model to obtain a face forgery recognition result, where the face forgery recognition model includes: A global feature encoder for extracting global feature vectors of the image data and the audio data in the global image according to the self-attention mechanism; A global region encoder for cropping the image data in the global image into a head region image, a facial region image, and a lip close-up image, extracting a head region local feature vector of the head region image, a facial region local feature vector of the facial region image, and a lip close-up local feature vector of the lip close-up image according to the residual neural network, and respectively splicing the head region local feature vector, the facial region local feature vector, and the lip close-up local feature vector with the global feature vector to construct a head region feature set, a facial region feature set, and a lip close-up feature set; A region perception module, which is used to extract the head feature weights of the head region feature set, the face feature weights of the face region feature set, and the lip feature weights of the lip close-up feature set respectively through a fully connected layer and a Sigmoid function, and construct a weight set; A feature fusion module, which is used to weightedly fuse the features in the head region feature set, the face region feature set, and the lip close-up feature set according to the weight set to generate fused features; A multi-layer perceptron classifier, which is used to output the face forgery recognition result according to the fused features.

[0007] Optionally, the extracting the image data and audio data in the face video data to be recognized and constructing a global image includes: Intercepting the original face images of a preset number of frames in the face video data to be recognized, and cropping the original face images to generate background-free face images; Extracting the audio data corresponding to the time of the preset number of frames in the face video data to be recognized, and plotting the audio data into a Mel spectrogram; Stitching the face image and the Mel spectrogram to form the global image.

[0008] Optionally, the cropping the original face images to generate background-free face images includes: Based on the face cropping model of dlib, cropping the original face images according to a preset cropping side length to generate the background-free face images.

[0009] Optionally, the trained face forgery recognition model is obtained by training the initial face forgery recognition model based on a binary classification loss function.

[0010] Optionally, the binary classification loss function includes: ; Wherein, is the binary classification loss function value, y is the true label, is the i-th true label, is the prediction result, is the i-th prediction result, and N is the number of samples.

[0011] Optionally, the trained face forgery recognition model is obtained by training the initial face forgery recognition model based on the binary classification loss function and a region perception loss function, and the region perception loss function is used to update the parameters of the fully connected layer of the region perception module.

[0012] Optionally, the region perception loss function includes: ; where loss2 is the value of the region-aware loss function, L is the length of the weight vector, T is the batch size, and K is a hyperparameter, is the maximum weight of the i-th sample at the j-th cropping ratio, is the edge weight of the i-th sample at the j-th cropping ratio.

[0013] In a second aspect, the present invention provides a face forgery recognition device based on audio and image, which applies the face forgery recognition method based on audio and image as described in the first aspect. The face forgery recognition device based on audio and image includes: An extraction module, configured to extract image data and audio data from the face video data to be recognized, and construct a global image; A result module, configured to input the global image into the trained face forgery recognition model to obtain a face forgery recognition result, where the face forgery recognition model includes: A global feature encoder, configured to extract global feature vectors of the image data and the audio data in the global image according to the self-attention mechanism; A global region encoder, configured to crop the image data in the global image into a head region image, a facial region image, and a lip close-up image, extract a head region local feature vector of the head region image, a facial region local feature vector of the facial region image, and a lip close-up local feature vector of the lip close-up image according to the residual neural network, and splice the head region local feature vector, the facial region local feature vector, and the lip close-up local feature vector with the global feature vector respectively to construct a head region feature set, a facial region feature set, and a lip close-up feature set; A region-aware module, configured to extract a head feature weight of the head region feature set, a facial feature weight of the facial region feature set, and a lip feature weight of the lip close-up feature set through a fully connected layer and a Sigmoid function respectively, and construct a weight set; A feature fusion module, configured to weightedly fuse the features in the head region feature set, the facial region feature set, and the lip close-up feature set according to the weight set to generate a fusion feature; A multi-layer perceptron classifier, configured to output the face forgery recognition result according to the fusion feature.

[0014] In a third aspect, the present invention provides an electronic device, including a memory and a processor; The memory is configured to store a computer program; The processor is configured to implement the audio- and image-based face forgery recognition method as described in the first aspect when executing the computer program.

[0015] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the audio- and image-based face forgery recognition method as described in the first aspect is implemented.

[0016] The beneficial effects of the audio- and image-based face forgery recognition method, apparatus, device, and medium of the present invention are as follows: By extracting image data and audio data from the face video data to be recognized, a global image is constructed, which facilitates the unified modeling of images and audio, realizes the effective fusion of multi-modal information, improves the consistency and operability of the input, and inputs the global image into the trained face forgery recognition model. Through the global feature encoder according to the self-attention mechanism, global feature vectors of the image data and the audio data in the global image are extracted, which can capture the cross-modal correlation and overall structure information between the image and the audio, enhance the global perception ability of forgery behaviors, and thus identify face forgery through the image dimension and the audio dimension, improving the accuracy of recognition. Then, the global image is cropped into head, face, and lip regions by the global region encoder, local features are respectively extracted and spliced with the global features to form a multi-scale feature representation with fine-grained semantics, enabling the model to not only perform consistency analysis on the whole but also pay attention to key local tampering traces and conduct detailed analysis on local features; further, the region perception module dynamically generates weights for the features of each region through a fully connected layer and a Sigmoid function, enabling the model to adaptively adjust the focus of attention according to the input content, enhancing the robustness to different forgery methods; after that, the feature fusion module weights and fuses the features of each region according to the weights to generate more representative fusion features, reducing the interference of redundant information; finally, analyzed and predicted by a multi-layer perceptron classifier with strong non-linear discrimination ability and good practicability, accurate face forgery recognition results can be output. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic structural diagram of the trained face forgery recognition model provided by an embodiment of the present invention; Figure 2 It is a schematic diagram of the global image provided by an embodiment of the present invention; Figure 3 It is a schematic diagram of the head region image provided by an embodiment of the present invention; Figure 4 It is a schematic diagram of the face region image provided by an embodiment of the present invention; Figure 5 It is a schematic diagram of the lip close-up image provided by an embodiment of the present invention; Figure 6 This is a schematic structural diagram of the electronic device provided by the embodiment of the present invention. Detailed implementation manners

[0018] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments described herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.

[0019] It should be understood that the various steps recorded in the method embodiments of the present invention can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.

[0020] The term "including" and its variants used herein are open-ended, that is, "including but not limited to"; the term "based on" is "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; the term "optionally" means "optional embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules, or units, and are not used to limit the order of the functions executed by these devices, modules, or units or their interdependent relationships.

[0021] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0023] In view of the problems existing in the above related technologies, this embodiment provides a method, device, equipment, and medium for face forgery recognition based on audio and image.

[0024] A method for face forgery recognition based on audio and image provided by an embodiment of the present invention includes: Extracting the image data and audio data in the face video data to be recognized and constructing a global image.

[0025] Specifically, a multimedia processing tool (such as FFmpeg) or a programming library (such as OpenCV, PyDub) can be used to complete the demultiplexing → decoding → output process, so as to extract the image data from the face video data to be recognized, and extract the audio file in the format of wav from the face video data to be recognized.

[0026] Input the global image into the trained face forgery recognition model to obtain the face forgery recognition result.

[0027] Specifically, the global image is as Figure 2 shown, Figure 2 The upper half of represents audio data, and the lower half represents image data. Input the global image into the trained face forgery recognition model to obtain the face forgery recognition result. The trained face forgery recognition model can be obtained by training the initial recognition model, and the initial recognition model can be constructed by a neural network.

[0028] Among them, as Figure 1 shown, the face forgery recognition model includes: A global feature encoder, which is used to extract the global feature vectors of the image data and the audio data in the global image according to the self-attention mechanism.

[0029] Specifically, the global feature encoder is constructed according to the self-attention mechanism, and can extract the temporal dependence information between the image data and the audio data in the global image. By paying attention to the correspondence between lip movement and audio frequency, the inconsistencies and abnormal patterns existing in the forged video can be recognized. The global feature encoder includes a convolutional layer and a self-attention encoder. After the convolutional layer extracts the preliminary features, the self-attention encoder captures the long-term cross-modal feature correlations and encodes them to obtain the global feature vectors. Exemplarily, the self-attention encoder can adopt the CLIP ViT-L / 14 image encoder.

[0030] A global region encoder, which is used to crop the image data in the global image into a head region image, a facial region image, and a lip close-up image, extract the head region local feature vector of the head region image, the facial region local feature vector of the facial region image, and the lip close-up local feature vector of the lip close-up image respectively according to the residual neural network, and splice the head region local feature vector, the facial region local feature vector, and the lip close-up local feature vector with the global feature vector respectively to construct a head region feature set, a facial region feature set, and a lip close-up feature set.

[0031] Specifically, the global region encoder first crops the head image corresponding to the image data header in the global image through a cropping unit. The size of the head image can be 224*224, that is, 224 pixel points multiplied by 224 pixel points. The cropping unit crops the head image according to the region relationship. The region relationship includes that the head region position can be [0:223,0:223], the face region position can be [28:195,28:195], and the lip close-up region position can be [71+35:152+35,71:152]. The cropping unit divides the head image through the region relationship to obtain the corresponding head region primary image, face region primary image, and lip close-up primary image. Then, the head region primary image with a cropping ratio of 1.00x, the face region primary image with a cropping ratio of 0.65x, and the lip close-up primary image with a cropping ratio of 0.37x are adjusted to a size of 224*224 to obtain the head region image as shown in Figure 3 the face region image as shown in Figure 4 and the lip close-up image as shown in Figure 5 so as to extract the head region local feature vector of the head region image, the face region local feature vector of the face region image, and the lip close-up local feature vector of the lip close-up image respectively according to the residual neural network, and splice the head region local feature vector, the face region local feature vector, and the lip close-up local feature vector with the global feature vector respectively through a splicing unit to construct a head region feature set, a face region feature set, and a lip close-up feature set. Among them, the head region image, the face region image, and the lip close-up image all include images of a preset number of frames. The preset number of frames can be set according to the actual situation, such as 5 frames, that is, the head region image, the face region image, and the lip close-up image all include the corresponding 5 images.

[0032] Exemplarily, since the lips are at the bottom of the picture, an additional parameter is added to adjust the center position to place the lips at the center position of the lip close-up image. Since the facial structures of different people do not differ much and the regional ratios are often very fixed, the regions can be cropped through fixed parameters.

[0033] The region perception module is used to extract the head feature weight of the head region feature set, the face feature weight of the face region feature set, and the lip feature weight of the lip close-up feature set respectively through a fully connected layer and a Sigmoid function, and construct a weight set.

[0034] Specifically, the region perception module extracts the head feature weights of the head region feature set, the face feature weights of the face region feature set, and the lip feature weights of the lip close-up feature set through a fully connected layer and a Sigmoid function, and constructs a weight set. The parameters of the fully connected layer can learn and assign weights according to the contribution degree of different region features in forgery recognition, so as to achieve the purpose of highlighting the expression ability of key region features, weakening redundant or noise information, and effectively improving the overall detection performance. Then, through normalization by the Sigmoid function, the head feature weights, face feature weights, and lip feature weights can be obtained, and the head feature weights, face feature weights, and lip feature weights are used to construct a weight set.

[0035] A feature fusion module is configured to fuse the features in the head region feature set, the face region feature set, and the lip close-up feature set according to the weight set to generate a fused feature.

[0036] Specifically, the feature fusion module fuses the features in the head region feature set, the face region feature set, and the lip close-up feature set according to the weight set by using a weighted fusion formula to generate a fused feature. The weighted fusion formula includes ; where F is the fused feature, is the i-th feature at the j-th cropping ratio in the head region feature set, the face region feature set, and the lip close-up feature set, is the feature weight corresponding to the i-th feature at the j-th cropping ratio. Exemplarily, there are 3 cropping ratios, namely 1.00x, 0.65x, and 0.37x. The number of features i corresponds to the preset number of frames. When the preset number of frames is 5 frames, there are 5 features at each cropping ratio.

[0037] A multi-layer perceptron classifier is configured to output the face forgery recognition result according to the fused feature.

[0038] Specifically, the multi-layer perceptron classifier has a multi-layer perceptron structure. First, the fused feature is input into a fully connected layer to obtain a scalar, and after sigmoid processing of this scalar, the face forgery recognition result is obtained.

[0039] In this embodiment, by extracting image data and audio data from the face video data to be recognized, a global image is constructed to facilitate the unified modeling of images and audio, realize the effective fusion of multi-modal information, improve the consistency and operability of the input, and input the global image into the trained face forgery recognition model. According to the self-attention mechanism, the global feature encoder extracts the global feature vectors of the image data and the audio data in the global image, which can capture the cross-modal association and overall structure information between the image and the audio, enhance the global perception ability of forgery behavior, and thus recognize face forgery through the image dimension and the audio dimension to improve the recognition accuracy. Then, the global region encoder crops the image into head, face, and lip regions, extracts local features respectively and splices them with the global features to form a multi-scale feature representation with fine-grained semantics, enabling the model to not only perform consistency analysis on the whole but also pay attention to key local tampering traces and perform detailed analysis on local features. Further, the region perception module dynamically generates the weights of the features of each region through a fully connected layer and the Sigmoid function, enabling the model to adaptively adjust the focus of attention according to the input content and improve the robustness to different forgery methods. After that, the feature fusion module weights and fuses the features of each region according to the weights to generate more representative fusion features and reduce the interference of redundant information. Finally, the multi-layer perceptron classifier with strong non-linear discrimination ability and good practicality is used for analysis and prediction, and the accurate face forgery recognition result can be output.

[0040] Optionally, the extracting the image data and audio data from the face video data to be recognized and constructing a global image includes: Intercepting the original face images of a preset number of frames in the face video data to be recognized, and cropping the original face images to generate background-free face images; Extracting the audio data corresponding to the time of the preset number of frames in the face video data to be recognized, and plotting the audio data into a Mel spectrogram; Splicing the face images and the Mel spectrograms to form the global image.

[0041] Specifically, intercept the original face images of a preset number of frames in the face video data to be recognized, and crop the original face images to remove the irrelevant background, generating background-free face images with only faces. The preset number of frames can be set according to the actual situation, for example, 5 frames; extract the audio data corresponding to the time of the preset number of frames in the face video data to be recognized. The audio data format can be wav, and plot the audio data into a Mel spectrogram. The size of the Mel spectrogram is 500*2500, and the size of a single face image is 500*500. When the preset number of frames is 5 frames, 5 face images are obtained. Stitch the 5 face images with a size of 500*500 and the Mel spectrogram with a size of 500*2500 to form a composite image with a size of 1000*2500. Adjust the dimensions and size of the composite image to form the global image as shown in Figure 2 shown, where Figure 2 the upper part represents the Mel spectrogram, and the lower part represents 5 face images.

[0042] Optionally, the cropping of the original face image to generate a background-free face image includes: Based on the face cropping model of dlib, crop the original face image according to the preset cropping side length to generate the background-free face image.

[0043] Specifically, the face cropping model of dlib is used to detect the frontal face in the image. The preset cropping side length is set according to the actual situation. For example, increase the side length of the cropping area to 1.3 times that of the conventional dlib face cropping model to more comprehensively obtain face information and avoid the problem of information loss caused by some face images being cropped off.

[0044] Optionally, the trained face forgery recognition model is obtained by training the initial face forgery recognition model based on the binary classification loss function.

[0045] Specifically, the trained face forgery recognition model is obtained by training the initial face forgery recognition model based on the binary classification loss function. The binary classification loss function can solve the problem that the conventional loss function, such as the binary cross-entropy loss function, makes the model unable to fit. When training the initial face forgery recognition model, the global feature encoder crops the global image into a Mel spectrogram and face images, shuffles the order, and then performs feature extraction to avoid overfitting of the model to specific data. In addition, the global image formed by stitching the Mel spectrogram and face images is input into the face forgery recognition model, and the global image is cropped into a Mel spectrogram and face images by the global feature encoder, which is convenient for data storage and calling.

[0046] Optionally, the binary classification loss function includes: ; Wherein, is the binary classification loss function value, y is the true label, is the i-th said true label, is the prediction result, is the i-th said prediction result, and N is the number of samples.

[0047] Optionally, the trained face forgery recognition model is obtained by training the initial face forgery recognition model based on the binary classification loss function and the region-aware loss function, and the region-aware loss function is used to update the parameters of the fully connected layer of the region-aware module.

[0048] Specifically, the trained face forgery recognition model is obtained by training the initial face forgery recognition model based on the binary classification loss function and the region-aware loss function. The sum of the loss results of the binary classification loss function and the region-aware loss function is used as the training basis for backpropagation and gradient calculation to update the model parameters. In the weight set obtained by the region-aware module through the fully connected layer and the Sigmoid function, there are also the maximum weight and the marginal weight. The maximum weight is the maximum item of all items, and the marginal weight is the first item. The maximum weight and the marginal weight are used to determine the result of the region-aware loss function, so that the region-aware loss function updates the parameters of the fully connected layer of the region-aware module through the result of the region-aware loss function.

[0049] Optionally, the region-aware loss function includes: ; where loss2 is the region-aware loss function value, L is the length of the weight vector, T is the batch size, K is a hyperparameter that can adjust the steepness of the loss change and is used to update the parameters of the fully connected layer, is the maximum weight of the i-th sample at the j-th cropping ratio, is the marginal weight of the i-th sample at the j-th cropping ratio.

[0050] An apparatus for face forgery recognition based on audio and image provided by an embodiment of the present invention applies the method for face forgery recognition based on audio and image as described above. The apparatus for face forgery recognition based on audio and image includes: An extraction module, configured to extract image data and audio data from the face video data to be recognized and construct a global image; A result module, configured to input the global image into the trained face forgery recognition model to obtain a face forgery recognition result, wherein the face forgery recognition model includes: A global feature encoder, configured to extract global feature vectors of the image data and the audio data in the global image according to the self-attention mechanism; A global region encoder, configured to crop the image data in the global image into a head region image, a face region image, and a lip close-up image, extract a head region local feature vector of the head region image, a face region local feature vector of the face region image, and a lip close-up local feature vector of the lip close-up image respectively according to a residual neural network, and splice the head region local feature vector, the face region local feature vector, and the lip close-up local feature vector with the global feature vector respectively to construct a head region feature set, a face region feature set, and a lip close-up feature set; A region perception module, configured to extract a head feature weight of the head region feature set, a face feature weight of the face region feature set, and a lip feature weight of the lip close-up feature set respectively through a fully connected layer and a Sigmoid function, and construct a weight set; A feature fusion module, configured to fuse the features in the head region feature set, the face region feature set, and the lip close-up feature set according to the weight set to generate a fused feature; A multi-layer perceptron classifier, configured to output the face forgery recognition result according to the fused feature.

[0051] As Figure 6 shown, an electronic device 600 provided in an embodiment of the present invention includes a memory 610 and a processor 620; the memory 610 is configured to store a computer program; the processor 620 is configured to implement the above-mentioned face forgery recognition method based on audio and image when executing the computer program.

[0052] Or, an electronic device 600 includes a memory 610 and a processor 620 coupled to the memory 610; the memory 610 is configured to store a computer program; the processor 620 is configured to perform the following operations when executing the computer program: Extract image data and audio data in the face video data to be recognized, and construct a global image; Input the global image into a trained face forgery recognition model to obtain a face forgery recognition result, where the face forgery recognition model includes: A global feature encoder, configured to extract a global feature vector of the image data and the audio data in the global image according to a self-attention mechanism; A global region encoder is used to crop the image data in the global image into a head region image, a face region image, and a lip close-up image. According to a residual neural network, a head region local feature vector of the head region image, a face region local feature vector of the face region image, and a lip close-up local feature vector of the lip close-up image are respectively extracted, and the head region local feature vector, the face region local feature vector, and the lip close-up local feature vector are respectively concatenated with the global feature vector to construct a head region feature set, a face region feature set, and a lip close-up feature set; A region perception module is used to respectively extract a head feature weight of the head region feature set, a face feature weight of the face region feature set, and a lip feature weight of the lip close-up feature set through a fully connected layer and a Sigmoid function, and construct a weight set; A feature fusion module is used to weightedly fuse the features in the head region feature set, the face region feature set, and the lip close-up feature set according to the weight set to generate a fused feature; A multi-layer perceptron classifier is used to output the face forgery recognition result according to the fused feature.

[0053] A computer-readable storage medium provided by an embodiment of the present invention, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned face forgery recognition method based on audio and image is implemented.

[0054] Or, a non-volatile computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor performs the following operations: Extract the image data and audio data in the face video data to be recognized, and construct a global image; Input the global image into a trained face forgery recognition model to obtain a face forgery recognition result, where the face forgery recognition model includes: A global feature encoder is used to extract a global feature vector of the image data and the audio data in the global image according to a self-attention mechanism; A global region encoder is used to crop the image data in the global image into a head region image, a face region image, and a lip close-up image, extract a head region local feature vector of the head region image, a face region local feature vector of the face region image, and a lip close-up local feature vector of the lip close-up image respectively according to a residual neural network, and splice the head region local feature vector, the face region local feature vector, and the lip close-up local feature vector with the global feature vector respectively to construct a head region feature set, a face region feature set, and a lip close-up feature set; A region perception module is used to extract a head feature weight of the head region feature set, a face feature weight of the face region feature set, and a lip feature weight of the lip close-up feature set respectively through a fully connected layer and a Sigmoid function, and construct a weight set; A feature fusion module is used to weightedly fuse the features in the head region feature set, the face region feature set, and the lip close-up feature set according to the weight set to generate a fused feature; A multi-layer perceptron classifier is used to output the face forgery recognition result according to the fused feature.

[0055] Now, an electronic device 600 that can be a server or a client of the present invention will be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device 600 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 600 can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0056] The electronic device 600 includes a computing unit that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) or a computer program loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The computing unit, the ROM, and the RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0057] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above various methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc. In this application, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. One can select some or all of the units according to actual needs to achieve the purpose of the solution of the embodiments of the present invention. In addition, the functional units in the various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0058] Although the present invention is disclosed as above, the scope of protection of the present invention is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will all fall within the scope of protection of the present invention.

Claims

1. A face forgery recognition method based on audio and images, characterized in that, Including: Extract the image data and audio data in the face video data to be recognized, and construct a global image; Input the global image into a trained face forgery recognition model to obtain a face forgery recognition result, where the face forgery recognition model includes: A global feature encoder for extracting global feature vectors of the image data and the audio data in the global image according to the self-attention mechanism; A global region encoder for cropping the image data in the global image into a head region image, a facial region image, and a lip close-up image, and respectively extracting a head region local feature vector of the head region image, a facial region local feature vector of the facial region image, and a lip close-up local feature vector of the lip close-up image according to a residual neural network, and respectively splicing the head region local feature vector, the facial region local feature vector, and the lip close-up local feature vector with the global feature vector to construct a head region feature set, a facial region feature set, and a lip close-up feature set; A region perception module for respectively extracting a head feature weight of the head region feature set, a facial feature weight of the facial region feature set, and a lip feature weight of the lip close-up feature set through a fully connected layer and a Sigmoid function, and constructing a weight set; A feature fusion module for weighted fusion of the features in the head region feature set, the facial region feature set, and the lip close-up feature set according to the weight set to generate a fusion feature; A multi-layer perceptron classifier for outputting the face forgery recognition result according to the fusion feature.

2. The method for identifying face forgery based on audio and image according to claim 1, wherein The extracting the image data and audio data in the face video data to be recognized and constructing a global image includes: Intercept the original face images of a preset number of frames in the face video data to be recognized, and crop the original face images to generate background-free face images; Extract the audio data corresponding to the time of the preset number of frames in the face video data to be recognized, and plot the audio data into a Mel spectrogram; Splice the face image and the Mel spectrogram to form the global image.

3. The face forgery recognition method based on audio and image according to claim 2, wherein The cropping the original face images to generate background-free face images includes: Based on the face cropping model of dlib, crop the original face images according to a preset cropping side length to generate the background-free face images.

4. The face forgery recognition method based on audio and image according to claim 1, wherein The trained face forgery recognition model is obtained by training an initial face forgery recognition model based on a binary classification loss function.

5. The face forgery recognition method based on audio and image according to claim 4, characterized in that, The binary classification loss function includes: ; Among them, is the binary classification loss function value, y is the true label, is the i-th said true label, is the prediction result, is the i-th said prediction result, and N is the number of samples.

6. The face forgery recognition method based on audio and image according to claim 4, characterized in that The trained face forgery recognition model is obtained by training the initial face forgery recognition model based on the binary classification loss function and a region perception loss function, and the region perception loss function is used to update the parameters of the fully connected layer of the region perception module.

7. The face forgery recognition method based on audio and image according to claim 6, characterized in that, The region perception loss function includes: ; Among them, loss2 is the value of the region perception loss function, L is the length of the weight vector, T is the batch size, and K is a hyperparameter. is the maximum weight of the i-th sample at the j-th cropping ratio. is the edge weight of the i-th sample at the j-th cropping ratio.

8. A face forgery recognition device based on audio and images, characterized in that, Applying the audio- and image-based face forgery recognition method according to any one of claims 1-7, the audio- and image-based face forgery recognition device includes: An extraction module, configured to extract image data and audio data from the face video data to be recognized, and construct a global image; A result module, configured to input the global image into a trained face forgery recognition model to obtain a face forgery recognition result, wherein the face forgery recognition model includes: A global feature encoder, configured to extract global feature vectors of the image data and the audio data in the global image according to a self-attention mechanism; A global region encoder, configured to crop the image data in the global image into a head region image, a facial region image, and a lip close-up image, and respectively extract a head region local feature vector of the head region image, a facial region local feature vector of the facial region image, and a lip close-up local feature vector of the lip close-up image according to a residual neural network, and respectively splice the head region local feature vector, the facial region local feature vector, and the lip close-up local feature vector with the global feature vector to construct a head region feature set, a facial region feature set, and a lip close-up feature set; A region perception module, configured to respectively extract a head feature weight of the head region feature set, a facial feature weight of the facial region feature set, and a lip feature weight of the lip close-up feature set through a fully connected layer and a Sigmoid function, and construct a weight set; A feature fusion module, configured to weightedly fuse features in the head region feature set, the facial region feature set, and the lip close-up feature set according to the weight set to generate a fusion feature; A multi-layer perceptron classifier, configured to output the face forgery recognition result according to the fusion feature.

9. An electronic device, characterized in that, Including a memory and a processor; The memory is configured to store a computer program; The processor is configured to, when executing the computer program, implement the audio- and image-based face forgery recognition method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the storage medium, and when the computer program is executed by the processor, the audio- and image-based face forgery recognition method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Living body detection method, device, electronic equipment and storage medium

    CN113505652A

  • Multi-mode face tampering video detection method and detector training method

    CN118097798A

  • Deep forgery active evidence obtaining method based on separable perceptual hash enhancement

    CN118587568A

  • Risk content identification method based on multi-modal large model

    CN119339419A