Methods, models, devices, and equipment for recognizing synthetic mouth-shaped faces.

By combining image and audio feature data, a classification model is trained to recognize synthetic mouth-shaped faces, solving the problem of difficulty in recognizing synthetic mouth-shaped faces in existing technologies and achieving higher recognition accuracy and security.

CN115223214BActive Publication Date: 2026-05-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2021-04-15
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively identify synthetically created faces with lip shapes, which could allow criminals to use virtual face synthesis technology to spread rumors and commit fraud.

Method used

By detecting facial feature points in sample images and combining them with audio feature data, a pre-trained speech-to-lip synthesis model is used to generate second facial feature data, and a classification model is trained to distinguish between real faces and synthesized lip faces.

Benefits of technology

It improves the accuracy of recognizing synthetic mouth-shaped faces, effectively identifies synthetic mouth-shaped faces and issues warnings, reducing the probability of rumors and fraud.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115223214B_ABST
    Figure CN115223214B_ABST
Patent Text Reader

Abstract

This disclosure relates to the field of artificial intelligence technology, and provides a method, model acquisition method, apparatus, and device for recognizing synthetic lip-sync faces. The method includes: detecting facial feature points in a sample image to obtain first facial feature data; inputting the first facial feature data and audio feature data to be synthesized into a pre-trained speech-to-lip synthesis model, and determining second facial feature data based on the output of the speech-to-lip synthesis model; training a classification model based on the first and second facial feature data to obtain a synthetic lip-sync face recognition model; inputting an image to be recognized into the synthetic lip-sync face recognition model, and determining whether the image to be recognized is a synthetic lip-sync face image based on the output of the synthetic lip-sync face recognition model. This technical solution can effectively identify synthetic lip-sync faces in images or videos, thereby reducing the probability of rumors, fraud, and other behaviors caused by synthetic lip-sync faces.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, specifically to a method and apparatus for recognizing a synthetic mouth-shaped face, a method and apparatus for acquiring a synthetic mouth-shaped face recognition model, and an electronic device for implementing the above method. Background Technology

[0002] With the development of artificial intelligence technology, virtual face synthesis technology has become increasingly sophisticated. For example, AI technology can realistically combine video footage with sound to generate a face that produces the aforementioned lifelike voice effect. However, this technology could potentially be exploited by criminals to spread rumors, commit fraud, and other illegal activities.

[0003] To minimize the aforementioned risks, there is an urgent need for a synthetic mouth-shaped face recognition solution.

[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure. Summary of the Invention

[0005] The purpose of this disclosure is to provide a method and apparatus for recognizing synthetic mouth-shaped faces, a method and apparatus for acquiring synthetic mouth-shaped face recognition models, and a computer-readable storage medium and electronic device for implementing the above methods. The synthetic mouth-shaped face recognition scheme provided by this scheme can effectively recognize synthetic mouth-shaped faces in images or videos, thereby improving the recognition accuracy of synthetic mouth-shaped faces.

[0006] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.

[0007] According to one aspect of this disclosure, a method for recognizing a synthetic lip-sync face is provided, comprising: performing facial feature point detection on a sample image to obtain first facial feature data; inputting the first facial feature data and audio feature data to be synthesized into a pre-trained speech-to-lip-sync synthesis model, and determining second facial feature data based on the output of the speech-to-lip-sync synthesis model; training a classification model based on the first facial feature data and the second facial feature data to obtain a synthetic lip-sync face recognition model; inputting an image to be recognized into the synthetic lip-sync face recognition model, and determining whether the image to be recognized is a synthetic lip-sync face image based on the output of the synthetic lip-sync face recognition model.

[0008] According to one aspect of this disclosure, a method for obtaining a synthetic lip-sync face recognition model is provided, comprising: performing facial feature point detection on a sample image to obtain first facial feature data; inputting the first facial feature data and audio feature data to be synthesized into a pre-trained speech-to-lip-sync synthesis model, and determining second facial feature data based on the output of the speech-to-lip-sync synthesis model; and training a classification model based on the first facial feature data and the second facial feature data to obtain a synthetic lip-sync face recognition model.

[0009] In some embodiments of this disclosure, based on the foregoing scheme, before performing facial feature point detection on the sample image to obtain the first facial feature data, the method further includes: acquiring a sample video, extracting frames from the video frames of the sample video to obtain the sample image; and performing face detection on the sample image to determine the face region.

[0010] The method of obtaining first facial feature data by detecting facial feature points in a sample image includes: detecting facial feature points in the facial region of the sample image to obtain feature point data; performing face alignment processing on the feature point data; and determining the feature data in the face-aligned image as the first facial feature data.

[0011] In some embodiments of this disclosure, based on the foregoing scheme, the method further includes: sampling the speech to obtain a speech sample; performing a short-time Fourier transform on the speech sample to convert the speech sample into a speech spectrogram; and filtering the speech spectrogram using a Mel filter bank to obtain the audio feature data to be synthesized.

[0012] In some embodiments of this disclosure, based on the foregoing scheme, before inputting the first facial feature data and the audio feature data to be synthesized into the pre-trained speech-to-lip synthesis model, the method further includes: performing image augmentation processing on the sample image to determine the first facial feature data through the augmented sample image; and performing audio augmentation processing on the collected speech to determine the audio feature data to be synthesized through the augmented speech.

[0013] In some embodiments of this disclosure, based on the foregoing scheme, the first facial feature data and the audio feature data to be synthesized are input into a pre-trained speech-to-lip synthesis model, and the second facial feature data is determined according to the output of the speech-to-lip synthesis model, including: converting the audio feature data to be synthesized into target lip shape data according to the pre-trained speech-to-lip synthesis model; fusing the target lip shape data with the first facial feature data according to the pre-trained speech-to-lip synthesis model to obtain a virtual synthesized image; and determining the second facial feature data according to the virtual synthesized image.

[0014] In some embodiments of this disclosure, based on the foregoing scheme, a classification model is trained according to the first facial feature data and the second facial feature data to obtain a synthetic mouth-shaped face recognition model, including: determining the i-th set of sample data according to the i-th first facial feature data and the i-th second facial feature data, wherein the i-th second facial data is determined according to the i-th first facial feature data; and training the classification model through N sets of sample data to obtain the synthetic mouth-shaped face recognition model, where i is a positive integer not greater than N.

[0015] In some embodiments of this disclosure, based on the foregoing scheme, a classification model is trained according to the first facial feature data and the second facial feature data to obtain a synthetic mouth-shaped face recognition model, including: determining a first sample according to the first facial feature data, and determining a second sample according to the second facial feature data; determining a first objective function according to the first sample and the label of the first sample, and determining a second objective function according to the second sample and the label of the second sample; determining a model objective function according to the first objective function and the second objective function; and training a classification model by minimizing the function value of the model objective function to obtain a synthetic mouth-shaped face recognition model.

[0016] In some embodiments of this disclosure, based on the foregoing scheme, determining the first sample according to the first facial feature data includes: rotating the sample image within a preset angle range to obtain a first set of images; scaling the sample image within a preset range to obtain a second set of images; and determining the first sample according to the facial feature data corresponding to the images in the first set of images, the facial feature data corresponding to the images in the second set of images, and the first facial feature data.

[0017] In some embodiments of this disclosure, based on the foregoing scheme, determining a first objective function according to the first sample and the label of the first sample includes: inputting the first sample into the synthetic mouth-shaped face recognition model to obtain a first discrimination result; and determining the first objective function based on the logarithmic loss of the first discrimination result and the label of the first sample.

[0018] In some embodiments of this disclosure, based on the foregoing scheme, determining the model objective function according to the first objective function and the second objective function includes: determining a first weight of the first objective function and determining a second weight of the second objective function; and determining the model objective function according to the first weight, the first objective function, the second weight, and the second objective function.

[0019] According to one aspect of this disclosure, a device for recognizing a synthetic mouth-shaped face is provided, comprising: a first real data acquisition module, a first synthetic data acquisition module, a first recognition model training module, and a face recognition module.

[0020] The first real data acquisition module is configured to: perform facial feature point detection on the sample image to obtain first facial feature data; the first synthetic data acquisition module is configured to: input the first facial feature data and the audio feature data to be synthesized into a pre-trained speech-to-lip synthesis model, and determine the second facial feature data based on the output of the speech-to-lip synthesis model; the first recognition model training module is configured to: train a classification model based on the first facial feature data and the second facial feature data to obtain a synthetic lip-shaped face recognition model; and the face recognition module is configured to: input the image to be recognized into the synthetic lip-shaped face recognition model, and determine whether the image to be recognized is a synthetic lip-shaped face image based on the output of the synthetic lip-shaped face recognition model.

[0021] In some embodiments of this disclosure, based on the foregoing scheme, the above-mentioned device further includes: a first face region determination module.

[0022] The first face region determination module is configured to: acquire a sample video, extract frames from the video frames of the sample video to obtain the sample image; and perform face detection on the sample image to determine the face region.

[0023] The aforementioned first real data acquisition module is specifically configured to: perform facial feature point detection on the aforementioned face region in the aforementioned sample image to obtain feature point data; perform face alignment processing on the aforementioned feature point data, and determine the feature data in the face-aligned image as the aforementioned first facial feature data.

[0024] In some embodiments of this disclosure, based on the foregoing scheme, the above-mentioned device further includes: a first audio feature data acquisition module.

[0025] The first audio feature data acquisition module is configured to: sample the speech to obtain a speech sample; perform a short-time Fourier transform on the speech sample to convert it into a speech spectrogram; and filter the speech spectrogram using a Mel filter bank to obtain the audio feature data to be synthesized.

[0026] In some embodiments of this disclosure, based on the foregoing scheme, the above-mentioned device further includes: a first image augmentation module and a first audio augmentation module.

[0027] The first image augmentation module is configured to perform image augmentation processing on the sample image to determine the first facial feature data through the augmented sample image; the first audio augmentation module is configured to perform audio augmentation processing on the acquired speech to determine the audio feature data to be synthesized through the augmented speech.

[0028] In some embodiments of this disclosure, based on the foregoing scheme, the first synthetic data acquisition module is specifically configured to: convert the audio feature data to be synthesized into target lip shape data according to the pre-trained speech-to-lip synthesis model; fuse the target lip shape data with the first facial feature data according to the pre-trained speech-to-lip synthesis model to obtain a virtual synthetic image; and determine the second facial feature data according to the virtual synthetic image.

[0029] In some embodiments of this disclosure, based on the foregoing scheme, the first recognition model training module is specifically configured to: determine the i-th set of sample data based on the i-th first face feature data and the i-th second face feature data, wherein the i-th second face data is determined based on the i-th first face feature data; train the classification model through N sets of sample data to obtain the above-mentioned synthetic mouth shape face recognition model, where i is a positive integer not greater than N.

[0030] In some embodiments of this disclosure, based on the foregoing scheme, the first recognition model training module includes: a sample determination unit, an objective function determination unit, and a training unit.

[0031] The sample determination unit is configured to: determine a first sample based on the first facial feature data, and determine a second sample based on the second facial feature data; the objective function determination unit is configured to: determine a first objective function based on the first sample and its label, and determine a second objective function based on the second sample and its label; determine a model objective function based on the first objective function and the second objective function; and the training unit is configured to: train a classification model by minimizing the function value of the model objective function to obtain a synthetic mouth shape face recognition model.

[0032] In some embodiments of this disclosure, based on the foregoing scheme, the sample determination unit is specifically configured to: rotate the sample image within a preset angle range to obtain a first set of images; scale the sample image within a preset range to obtain a second set of images; and determine the first sample based on the facial feature data corresponding to the images in the first set of images, the facial feature data corresponding to the images in the second set of images, and the first facial feature data.

[0033] In some embodiments of this disclosure, based on the foregoing scheme, the objective function determination unit is specifically configured to: input the first sample into the synthetic mouth-shaped face recognition model to obtain a first discrimination result; and determine the first objective function based on the first discrimination result and the logarithmic loss of the label of the first sample.

[0034] In some embodiments of this disclosure, based on the foregoing scheme, the objective function determination unit is further specifically configured to: determine the first weight of the first objective function and determine the second weight of the second objective function; and determine the model objective function based on the first weight, the first objective function, the second weight, and the second objective function.

[0035] In some embodiments of this disclosure, based on the foregoing scheme, the face recognition module is specifically configured to: acquire a video to be recognized; extract frames from the video frames of the video to be recognized to obtain the image to be recognized; perform face detection processing on the image to be recognized to determine the face region; perform face feature point detection processing on the face region in the image to be recognized to obtain feature point data; perform face alignment processing on the feature point data in the image to be recognized; and input the target image after face alignment processing, after image scaling, into the synthetic mouth-shaped face recognition model; and determine whether the image to be recognized is a synthetic mouth-shaped face image based on the recognition result output by the synthetic mouth-shaped face recognition model.

[0036] According to one aspect of this disclosure, a device for acquiring a synthetic mouth-shaped face recognition model is provided, comprising: a second real data acquisition module, a second synthetic data acquisition module, and a second recognition model training module.

[0037] The second real data acquisition module is configured to: perform facial feature point detection on the sample image to obtain first facial feature data; the second synthetic data acquisition module is configured to: input the first facial feature data and the audio feature data to be synthesized into a pre-trained speech-to-lip synthesis model, and determine the second facial feature data based on the output of the speech-to-lip synthesis model; the second recognition model training module is configured to: train a classification model based on the first facial feature data and the second facial feature data to obtain a synthetic lip-sync face recognition model.

[0038] According to one aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method for recognizing a synthetic mouth-shaped face as described in the first aspect and the method for acquiring a synthetic mouth-shaped face recognition model as described in the second aspect.

[0039] According to one aspect of this disclosure, an electronic device is provided, comprising: one or more processors; and a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the synthetic mouth-shaped face recognition method described in the first aspect above, and to implement the synthetic mouth-shaped face recognition model acquisition method described in the second aspect above.

[0040] According to one aspect of this disclosure, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the synthetic mouth-shaped face recognition method and the synthetic mouth-shaped face recognition model acquisition method provided in the various embodiments described above.

[0041] As can be seen from the above technical solutions, the synthetic mouth-shaped face recognition method, synthetic mouth-shaped face recognition device, computer-readable storage medium, and electronic device in the exemplary embodiments of this disclosure have at least the following advantages and positive effects:

[0042] In some embodiments of this disclosure, a synthetic mouth-shaped face recognition model is used to identify whether an image to be recognized belongs to a synthetic mouth-shaped face or a real face. For example, if a synthetic mouth-shaped face is detected in the image to be recognized, a warning can be issued to the user to increase the user's vigilance, thereby reducing the probability of rumors, fraud, and other behaviors caused by artificial intelligence technology that can synthesize video images and audio.

[0043] The training samples for the aforementioned face recognition model include first face (real face) feature data and second face (synthetic mouth-shaped face) feature data synthesized from the first face feature data. The second face (synthetic mouth-shaped face) feature data is synthesized by combining the first face feature data with audio feature data to be synthesized. Therefore, by training the classification model with multiple sets of training data, a face recognition model for distinguishing between real faces and synthetic mouth-shaped faces can be obtained. Furthermore, this technical solution utilizes big data to recognize synthetic mouth-shaped faces, which helps ensure the accuracy of the recognition results.

[0044] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0045] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0046] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this disclosure can be applied is shown.

[0047] Figure 2 This diagram illustrates a flowchart of a method for recognizing a synthesized mouth-shaped face in an exemplary embodiment of this disclosure.

[0048] Figure 3 This diagram illustrates a flowchart of a method for recognizing a synthetic mouth-shaped face in another exemplary embodiment of this disclosure.

[0049] Figure 4 This diagram illustrates a flowchart of a method for determining first facial feature data in an exemplary embodiment of this disclosure.

[0050] Figure 5 This diagram illustrates a flowchart of a method for determining audio feature data in an exemplary embodiment of this disclosure.

[0051] Figure 6 This diagram illustrates a flowchart of a method for determining second facial feature data in an exemplary embodiment of this disclosure.

[0052] Figure 7 This diagram illustrates a flowchart of a training method for a synthetic mouth-shaped face recognition model in an exemplary embodiment of this disclosure.

[0053] Figure 8 This diagram illustrates a flowchart of a training method for a synthetic mouth-shaped face recognition model in another exemplary embodiment of this disclosure.

[0054] Figure 9 This diagram illustrates a flowchart of a method for recognizing a synthesized mouth-shaped face in an exemplary embodiment of this disclosure.

[0055] Figure 10 This diagram illustrates a flowchart of a method for recognizing a synthetic mouth-shaped face in another exemplary embodiment of this disclosure.

[0056] Figure 11 This diagram illustrates the structure of a device for recognizing a synthesized mouth-shaped face in an exemplary embodiment of this disclosure.

[0057] Figure 12 This diagram illustrates the structure of a device for acquiring a synthetic mouth-shaped face recognition model in an exemplary embodiment of this disclosure.

[0058] Figure 13 A schematic diagram of the structure of an electronic device in an exemplary embodiment of this disclosure is shown. Detailed Implementation

[0059] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art.

[0060] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0061] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0062] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0063] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0064] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0065] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, transfer learning, inductive learning, and instructional learning.

[0066] The solutions provided in this disclosure involve technologies such as machine learning in artificial intelligence, and are specifically illustrated through the following embodiments:

[0067] In an exemplary embodiment, the application scenario of this technical solution can be that a well-known person is narrating something in the current video. The current video can be used as the target to be identified, and this technical solution can be used to identify whether the face and voice in the video are synthesized in post-production or whether the voice is actually spoken by the well-known person.

[0068] Compared to replacing the entire face of person A in an image with the face of person B, this solution only alters the mouth features of the original image in the case of synthesized mouth-shaped faces. However, other facial features (such as eyes and nose) remain realistic. Therefore, the synthesized mouth-shaped faces are more deceptive and difficult to identify. This high level of deception makes users more susceptible to being fooled, seriously endangering public safety.

[0069] To address the aforementioned technical problems, this technical solution provides a method, apparatus, medium, and device for recognizing synthetic mouth-shaped faces.

[0070] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this disclosure can be applied is shown.

[0071] like Figure 1 As shown, the system architecture 100 may include a terminal 110, a network 120, and a server 130. The terminal 110 and the server 130 are connected through the network 120.

[0072] Terminal 110 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. Network 120 can be a communication medium of various connection types that can provide a communication link between terminal 110 and server 130, such as a wired communication link, a wireless communication link, or a fiber optic cable, etc., which is not limited in this application. Server 130 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The number of servers 130 and the number of terminals 110 are not limited.

[0073] Specifically, server 130 can train the classification model in this scheme to obtain a synthetic lip-sync face recognition model, thereby realizing the recognition of the image to be recognized. For example, server 130 obtains the items containing user browsing behavior to obtain an item list. Then, server 130 performs the following steps: detecting facial feature points on the sample image to obtain first facial feature data; inputting the first facial feature data and the audio feature data to be synthesized into a pre-trained speech-to-lip synthesis model, determining second facial feature data based on the output of the speech-to-lip synthesis model; training a classification model based on the first and second facial feature data to obtain a synthetic lip-sync face recognition model.

[0074] For example, server 130 may also provide pre-training of the speech-to-lip synthesis model in this solution to synthesize the first face (real face) feature data and the audio feature data to be synthesized into a model, thereby determining the second face (synthesized lip-shaped face) feature data. Server 130 may also store the recognition device of the synthesized lip-shaped face, send the recognition result to client 110, and store the recognition result on server 130.

[0075] Alternatively, users can provide an image or video to be recognized via terminal 120 and send it to server 130 for synthetic mouth-shaped face recognition. For example, the video to be recognized could be a scene of a person speaking. Further, the image or video to be recognized is sent to server 130, and for example, server 130 performs the following steps: inputting the image to be recognized into the aforementioned synthetic mouth-shaped face recognition model, and determining whether the image to be recognized is a synthetic mouth-shaped face image based on the output of the synthetic mouth-shaped face recognition model.

[0076] The method for recognizing synthesized mouth-shaped faces in this disclosure can also be applied to terminals. This disclosure does not impose any special limitations on this method. This disclosure primarily uses the application of the synthesized mouth-shaped face recognition method to server 130 as an example for illustration.

[0077] The following section introduces the method for recognizing synthetic mouth shapes provided in this technical solution. Specifically, Figure 2 This diagram illustrates a flowchart of a method for recognizing a synthesized mouth-shaped face in an exemplary embodiment of this disclosure. (See reference...) Figure 2 The method for recognizing synthetic mouth-shaped faces provided in this embodiment includes:

[0078] Step S210: Perform facial feature point detection on the sample image to obtain the first facial feature data;

[0079] Step S220: Input the first facial feature data and the audio feature data to be synthesized into the pre-trained speech-to-lip synthesis model, and determine the second facial feature data according to the output of the speech-to-lip synthesis model.

[0080] Step S230: Train a classification model based on the first facial feature data and the second facial feature data to obtain a synthetic mouth shape facial recognition model; and,

[0081] Step S240: Input the image to be recognized into the synthetic mouth-shaped face recognition model, and determine whether the image to be recognized is a synthetic mouth-shaped face image based on the output of the synthetic mouth-shaped face recognition model.

[0082] In some embodiments of this disclosure, a synthetic mouth-shaped face recognition model is used to identify whether an image to be recognized belongs to a synthetic mouth-shaped face or a real face. For example, if the image to be recognized is found to be a synthetic mouth-shaped face, a warning can be issued to the user to increase the user's vigilance, thereby reducing the probability of rumors, fraud, and other behaviors caused by artificial intelligence technology that can synthesize video images and audio. The training samples of the aforementioned face recognition model include first face (real face) feature data and second face (synthetic mouth-shaped face) feature data synthesized from the first face feature data. The second face (synthetic mouth-shaped face) feature data is synthesized by combining the first face feature data with the audio feature data to be synthesized. Therefore, by training the classification model with multiple sets of the above training data, a face recognition model for distinguishing between real faces and synthetic mouth-shaped faces is obtained. Compared with related solutions that recognize the entire face, this technical solution can obtain more refined face recognition results and has stronger recognition targeting. Meanwhile, this technical solution uses big data to recognize synthetic mouth-shaped faces, which helps to ensure the accuracy of the recognition results.

[0083] In this technical solution, the so-called "real face" image specifically refers to an image in which the mouth shape of the face is not synthesized in post-processing.

[0084] The following examples illustrate... Figure 2 The specific implementation methods of each step in the illustrated embodiment are described in detail below:

[0085] In an exemplary embodiment, Figure 3 This diagram illustrates the flowchart of a method for obtaining a synthetic mouth-shape facial recognition model. (Reference) Figure 3 The process of determining the sample data for the aforementioned synthetic mouth-shaped face model 400 includes: a process 310 for determining the feature data of the first face (real face) and a process 330 for determining the feature data of the second face (synthetic mouth-shaped face). To determine the feature data of the second face, a process 320 for determining the audio feature data to be synthesized is also required. Further, a classification model is trained based on the aforementioned first face feature data and the aforementioned second face feature data to obtain the synthetic mouth-shaped face recognition model 400.

[0086] It should be noted that the method for obtaining the synthetic mouth shape face recognition model provided in this technical solution is as follows: Figure 2 As shown in steps S210, S220 and S230, the method for obtaining the synthetic mouth shape face recognition model is described by the specific implementation of steps S210, S220 and S230.

[0087] For example, this technical solution determines N sets (N is a positive integer) of sample data to train a classification model to obtain the aforementioned synthetic mouth-shaped face recognition model 400. Each set of sample data includes first face (real face) feature data and second face (synthetic mouth-shaped face) feature data synthesized from the first face feature data. In other words, this set of samples is determined based on the same first face (real face) feature data. Therefore, the difference between the two types of face feature data in the same set of sample data lies in the mouth area. Training the classification model using this feature data helps to obtain a synthetic mouth-shaped face recognition model 400 with higher recognition accuracy.

[0088] In an exemplary embodiment, the process of determining the first face (real face) feature data is first described 310:

[0089] refer to Figure 3 The aforementioned first facial feature image can be derived from a video (denoted as "Sample Video 31") or an image (denoted as "Sample Image 31"). The following uses... Figure 4 The illustrated embodiment describes the process 310 of determining the first facial feature data 34 based on sample video 31'. (See reference...) Figure 4 This includes steps S410-S440.

[0090] In step S410, a sample video is acquired, and the video frames of the sample video are extracted to obtain the sample image.

[0091] In an exemplary embodiment, reference is made to Figure 3 The sample image 31 is obtained by separating video frames from the sample video 31' through video frame extraction processing (step S35) and then extracting several frames at certain intervals.

[0092] It should be noted that if the first facial feature data is determined based on an image containing a face (such as sample image 31), then step S420 is executed directly.

[0093] In step S420, face detection is performed on the sample image to determine the face region.

[0094] In an exemplary embodiment, a face region can be outlined in the sample image 31 using a face detection algorithm (step S31). For example, to improve image processing efficiency, MTCNN (Multi-task convolutional neural network) can be used to process the sample data. MTCNN implements a multi-task learning network through cascaded CNN models. For example, a shallow CNN network quickly generates a series of candidate windows, and a more powerful CNN network further filters out most non-face candidate windows, thus obtaining the face region in the sample image. Simultaneously, multiple sample images can be processed in parallel to quickly determine the face regions in multiple sample images.

[0095] In step S430, facial feature point detection is performed on the face region in the sample image to obtain feature point data. In step S440, face alignment processing is performed on the feature point data, and the feature data in the face-aligned image is determined as the first facial feature data.

[0096] In this embodiment, landmark localization, also known as face alignment, is used to improve the accuracy of synthesized mouth-shaped face recognition by utilizing the facial feature data of the image after face alignment processing. For example, a Face Alignment Network (FAN) can be used to detect facial landmarks (step S32), such as detecting 68 facial landmarks. Further, in an exemplary embodiment, face alignment processing is performed based on the detected landmarks (step S33), such as mapping the eyes and mouth in each image to preset positions using affine transformation processing. Further, the feature data obtained from the aligned image is determined to obtain the aforementioned first face (real face) feature data 34.

[0097] It should be noted that the aforementioned first facial feature data 34 will be used, on the one hand, to synthesize a mouth-shaped face with the audio feature data 33 to obtain second facial (synthesized mouth-shaped face) feature data 36; on the other hand, the aforementioned first facial feature data 34 will also be used with the corresponding second facial feature data 36 to construct sample data for the synthesized mouth-shaped face recognition model 400. Next, relevant embodiments for synthesizing a mouth-shaped face based on the first facial feature data and the audio feature data will be introduced.

[0098] The following is passed Figure 5 The illustrated embodiment describes the process 320 for determining audio feature data 33. (See reference...) Figure 5 ,include:

[0099] Step S510: Sampling the speech to obtain speech samples. Step S520: Performing a short-time Fourier transform on the speech samples to convert them into a speech spectrogram. Step S530: Filtering the speech spectrogram using a Mel filter bank to obtain the audio feature data to be synthesized.

[0100] In an exemplary embodiment, reference is made to Figure 3 First, the speech 32 is processed by frame segmentation. Speech frames are sampled with a length of 0.02s and a step size of 0.01s to obtain speech samples. A short-time Fourier transform is performed on each obtained speech sample, and for example, the first 257 coefficients are retained, thus converting the speech sample into a speech spectrogram. For example, the above audio is then filtered by S (S is a positive integer) Mel filters. The speech spectrogram corresponding to each speech sample, after passing through the preset number of Mel filters, is converted into a vector of length S, thus obtaining the audio feature data 33 to be synthesized. In other words, the number of Mel filters can be set according to the required vector length.

[0101] It should be noted that the method of obtaining audio feature data is not limited to this; other methods of obtaining audio feature data within the field can also be used, and no limitation is made here.

[0102] Furthermore, a synthesized lip-shaped face can be determined based on the first face (real face) feature data 34 and the audio feature data 33 to be synthesized, to obtain the second face feature data 36. To further enrich the sample data, this technical solution can perform image augmentation processing (step S34) on the sample image or the image after face alignment processing, and further employ... Figure 3 The illustrated embodiment obtains the first facial feature data 34 corresponding to the augmented image. Additionally, audio augmentation processing (step S36) can be performed on the aforementioned speech, and further... Figure 4 The embodiment shown obtains audio feature data 33 corresponding to the augmented speech.

[0103] For example, the image augmentation process (step S34) may include: rotating the sample image or the image after face alignment within a preset angle range (e.g., -5° to 5°), or scaling the sample image within a preset range (e.g., 0.98x to 1.02x). Alternatively, the sample image may be rotated within a preset angle range and then scaled within a preset range, thereby obtaining richer image resources. It should be noted that since the second face feature data in the same set of samples is determined based on the first face feature data, the images corresponding to the two face feature data in the same set of sample data are rotated by the same angle and scaled by the same ratio.

[0104] For example, the audio augmentation process (step S36) may include stretching the audio length of the above speech within a preset range (e.g., 0.8 times to 1.2 times) or adjusting the pitch of the above speech within a preset range (e.g., -1 semitone to 1 semitone) to obtain richer audio resources, so as to further synthesize richer second facial feature data.

[0105] In an exemplary embodiment, after determining the first facial feature data and obtaining the audio feature data to be synthesized in step S210, in step S220, the first facial feature data and the audio feature data to be synthesized are input into a pre-trained speech-to-lip synthesis model, and the second facial feature data is determined based on the output of the speech-to-lip synthesis model. The following is an example... Figure 6 The illustrated embodiment describes the process 330 for determining the second facial feature data 36.

[0106] In an exemplary embodiment, the following is by Figure 6The illustrated embodiment describes a scheme for determining the aforementioned second facial feature data. (See reference...) Figure 6 This includes steps S610 to S630.

[0107] In step S610, the audio feature data to be synthesized is converted into target lip shape data according to the pre-trained speech-to-lip synthesis model. In step S620, the target lip shape data is fused with the first facial feature data according to the pre-trained speech-to-lip synthesis model to obtain a virtual synthesized image.

[0108] In an exemplary embodiment, a pre-trained speech-to-lip synthesis model 300 (e.g., wav2lip) can be used. wav2lip can combine a video of a person and a target audio segment, converting the target audio into the lip movements of the person in the video, thus demonstrating the effect of the person in the video speaking the target audio. The pre-trained speech-to-lip synthesis model can convert the audio feature data to be synthesized into target lip movement data. Further, the pre-trained speech-to-lip synthesis model then fuses the target lip movement data with the first facial feature data through a fusion process to obtain a virtual synthesized image 37.

[0109] In step S630, the second facial feature data is determined based on the virtual synthetic image.

[0110] In an exemplary embodiment, the following can be employed: Figure 3 The method embodiment shown for determining the first facial feature data obtains the second facial feature data corresponding to the above-mentioned synthesized image, which will not be described in detail here.

[0111] In an exemplary embodiment, in step S230: a classification model is trained based on the first facial feature data and the second facial feature data to obtain a synthetic mouth shape facial recognition model.

[0112] In this embodiment, the i-th set of sample data is determined based on the i-th first face feature data and the i-th second face feature data, wherein the i-th second face data is determined based on the i-th first face feature data. Further, a classification model is trained using N sets of sample data to obtain the synthetic mouth-shape face recognition model, where i is a positive integer not greater than N.

[0113] As can be seen, the same set of sample data includes first face (real face) feature data and second face (synthetic mouth-shaped face) feature data synthesized from the first face feature data. In other words, the sample set is determined based on the same first face (real face) feature data. Therefore, the difference between the two face feature data in the same set of sample data lies in the mouth. Using this to train the classification model is beneficial for obtaining a synthetic mouth-shaped face recognition model with higher recognition accuracy.

[0114] Specifically, Figure 7 An example of training a classification model using multiple sets of sample data is shown. (Reference) Figure 7 This includes steps S710-S740.

[0115] In step S710, a first sample is determined based on the first facial feature data, and a second sample is determined based on the second facial feature data.

[0116] In an exemplary embodiment, the sample image can be rotated within a preset angle range (e.g., -5° to 5°) to obtain a first set of images. The sample image can also be scaled within a preset range (e.g., 0.98x to 1.02x) to obtain a second set of images. Alternatively, the sample image can be rotated within a preset angle range and then scaled within a preset range to obtain a third set of images.

[0117] The first, second, and third sets of images mentioned above, along with the first face feature data that has not been scaled / selected, can all be used to determine, for example... Figure 3 / Figure 8 The first sample 35 shown is shown.

[0118] In an exemplary embodiment, similarly, refer to Figure 3 The virtual composite image 37 and the image obtained by image augmentation processing (step S37) can be rotated within a preset angle range (e.g., -5° to 5°) to obtain a fourth set of images. The virtual composite image can also be scaled within a preset range (e.g., 0.98x to 1.02x) to obtain a fifth set of images. Furthermore, the virtual composite image can be rotated within a preset angle range and then scaled within a preset range to obtain a sixth set of images.

[0119] The fourth, fifth, and sixth sets of images mentioned above, along with the second set of facial feature data that has not been scaled or selected, can all be used to determine, for example... Figure 3 / Figure 8 The second sample 38 shown is illustrated.

[0120] In summary, the first sample 35 represents the feature data of a real human face, and the second sample 38 represents the feature data of a synthetic mouth-shaped face. That is, each set of sample data includes the first sample 35 and the second sample 38 determined based on the first sample 35.

[0121] In step S720, a first objective function is determined based on the first sample and its label, and a second objective function is determined based on the second sample and its label.

[0122] In an exemplary embodiment, reference is made to Figure 8 Set the label of the first sample 35 in each group of sample data to "0" and set the label of the second sample in each group of sample data to "1".

[0123] In an exemplary embodiment, the synthetic mouth-shape face recognition model 400 provided by this solution is a binary classification model, and therefore the model loss can be calculated using the logarithmic loss method (binary_crossentropy). It should be noted that the method for calculating the model loss is not limited to this; other methods within the art can also be used, and are not limited here.

[0124] Specifically, refer to Figure 8 For the first sample 35 in each set of sample data, the first sample 35 is image scaled (step S81) and then input into the synthetic mouth shape face recognition model 400 to obtain the first discrimination result 82. Then, based on the first discrimination result 83 and the log loss of the label "0" of the first sample 35, the first objective function 83 is determined.

[0125] In an exemplary embodiment, reference is made again. Figure 8 For the second sample 38 in each set of sample data, the second sample 38 is image scaled (step S81) and then input into the synthetic mouth shape face recognition model 400 to obtain the second discrimination result 82'. Then, based on the second discrimination result 82' and the log loss of the label "1" of the second sample, the second objective function 83' is determined.

[0126] In step S730, the model objective function is determined based on the first objective function and the second objective function.

[0127] In an exemplary embodiment, the model objective function 84 can be determined by weighting the first objective function 83 and the second objective function 83'. Specifically, a first weight of the first objective function and a second weight of the second objective function are determined, where the first and second weights are normalized weights. For example, when the impact of the first objective function on the total model loss is the same as the impact of the second objective function on the total model loss, both the first and second weights can be set to 0.5. Further, the product of the first weight and the first objective function, and the product of the second weight and the second objective function are calculated, and the model objective function is determined based on the sum of the two products. This model objective function encompasses both aspects of loss and allows for flexible adjustment of the impact of different losses on the total loss through weights.

[0128] For example, the objective function of the model can also be determined by directly adding the first model loss and the second model loss.

[0129] It should be noted that the specific implementation methods for determining the model objective function based on the first objective function and the second objective function are not limited to the above-mentioned methods. That is, the model objective function in this scheme must cover the above two aspects of loss, and the method of combining the first loss and the second loss is not limited.

[0130] In step S740, a classification model is trained by minimizing the function value of the model objective function to obtain a synthetic mouth shape face recognition model.

[0131] In an exemplary embodiment, gradient descent or Adam optimization algorithm can be used to optimize the objective function of the above model so that the synthetic mouth-shaped face model has a high recognition accuracy.

[0132] In an exemplary embodiment, Table 1 shows the network structure of the above-described synthetic mouth-shaped face recognition model.

[0133] Table 1

[0134]

[0135]

[0136] Referring to Table 1, since the network structure of this synthetic mouth-shaped face recognition model includes fully connected layers, the image size input to the model must be uniform. For example, the image input size is set to "224×224×3" in this network structure. Furthermore, the network structure also includes multiple convolutional and pooling layers for image processing. Table 1 records the number of channels, kernel size, stride, and other information for each convolutional, pooling, and regularization layer. Each convolutional layer uses the Rectified Luminaire (ReLU) as its activation function.

[0137] Meanwhile, the model incorporates multiple fully connected layers, which effectively increases the number of neurons, thus improving model complexity. Furthermore, deepening the number of fully connected layers enhances the model's nonlinear expressive ability and learning capacity.

[0138] Furthermore, to avoid overfitting, this embodiment incorporates three fully connected layers in its improved network structure. To further mitigate overfitting, dropout layers are added after the first two fully connected layers to enhance the model's generalization ability. The final fully connected layer uses the sigmoid function as the activation function to classify the image to be recognized.

[0139] The above embodiments provide a scheme for obtaining a synthetic mouth-shaped face recognition model. Further, the following describes a scheme for recognition based on the obtained synthetic mouth-shaped face recognition model. Specifically, the synthetic mouth-shaped face recognition model can be used to identify whether an image in a video is a synthetic mouth-shaped face image or a real image, and it can also be used to identify a single image as either a synthetic mouth-shaped face image or a real image.

[0140] For example, Figure 9 This diagram illustrates the flowchart of a method for recognizing synthetic mouth shapes. (Combined with...) Figure 10 refer to Figure 9 The method includes:

[0141] Step S910: Acquire the video 100 to be recognized, and extract frames from the video 100 to obtain the image 101 to be recognized. Step S920: Perform face detection processing on the image 101 to be recognized to determine the face region 102. Step S930: Perform face feature point detection processing on the face region 102 in the image to be recognized to obtain feature point data 103. Step S940: Perform face alignment processing on the feature point data 103 in the image to be recognized. And, Step S950: Scale the target image 104 after face alignment processing and input it to the synthetic mouth-shaped face recognition model 105. Determine whether the image to be recognized is a synthetic mouth-shaped face image based on the recognition result 106 output by the synthetic mouth-shaped face recognition model 105.

[0142] The specific implementation methods of steps S910-S940 are as follows: Figure 4 The embodiments shown are similar and will not be described again here.

[0143] In step S950, the target image after face alignment processing is scaled to the image size required by the model (e.g., 224×224×3) and input into the aforementioned synthetic mouth-shaped face recognition model. Further, after convolution processing, pooling processing, regularization processing, and fully connected processing of the network structure shown in Table 1, the recognition result of the target image can be output. An example is the recognition probability value. If the probability value is greater than 0.5, it can be considered that the current video contains a synthetic mouth-shaped face image; conversely, if the probability value is less than 0.5, it can be considered that the face image contained in the current video is a real face.

[0144] This technical solution uses a synthetic mouth-shaped face recognition model to identify whether an image to be identified belongs to a synthetic mouth-shaped face or a real face. For example, if the image to be identified is a synthetic mouth-shaped face, a warning can be issued to the user to increase their vigilance and reduce the probability of rumors, fraud, and other behaviors caused by artificial intelligence technology that can synthesize video images and audio. In addition, this technical solution trains a classification model using the aforementioned multiple sets of training data to obtain a face recognition model for distinguishing between real faces and synthetic mouth-shaped faces. Compared with related solutions that identify the entire face, this technical solution can obtain more refined face recognition results and has stronger recognition targeting. At the same time, this technical solution uses big data to achieve the recognition of synthetic mouth-shaped faces, which helps to ensure the accuracy of the recognition results.

[0145] Those skilled in the art will understand that all or part of the steps of the above embodiments are implemented as a computer program executed by a processor (including a GPU / CPU). When the computer program is executed by the GPU / CPU, it performs the functions defined by the methods provided in this disclosure. The program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disk.

[0146] Furthermore, it should be noted that the above figures are merely illustrative representations of the processes included in the methods according to exemplary embodiments of this disclosure, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0147] The following is passed Figure 11 This invention introduces an embodiment of a synthetic mouth-shaped face recognition device, which can be used to perform the synthetic mouth-shaped face recognition method described above.

[0148] Figure 11 This diagram illustrates the structure of a device for recognizing a synthesized mouth-shaped face in an exemplary embodiment of this disclosure. Figure 11As shown, the above-mentioned synthetic mouth-shaped face recognition device 1100 includes: a first real data acquisition module 1101, a first synthetic data acquisition module 1102, a first recognition model training module 1103, and a face recognition module 1104.

[0149] The first real data acquisition module 1101 is configured to: perform facial feature point detection on the sample image to obtain first facial feature data; the first synthetic data acquisition module 1102 is configured to: input the first facial feature data and the audio feature data to be synthesized into a pre-trained speech-to-lip synthesis model, and determine the second facial feature data based on the output of the speech-to-lip synthesis model; the first recognition model training module 1103 is configured to: train a classification model based on the first facial feature data and the second facial feature data to obtain a synthetic lip-shaped face recognition model; and the face recognition module 1104 is configured to: input the image to be recognized into the synthetic lip-shaped face recognition model, and determine whether the image to be recognized is a synthetic lip-shaped face image based on the output of the synthetic lip-shaped face recognition model.

[0150] In some embodiments of this disclosure, based on the foregoing scheme, the above-mentioned device further includes: a first face region determination module.

[0151] The first face region determination module is configured to: acquire a sample video, extract frames from the video frames of the sample video to obtain the sample image; and perform face detection on the sample image to determine the face region.

[0152] The aforementioned first real data acquisition module 1101 is specifically configured to: perform facial feature point detection on the aforementioned face region in the aforementioned sample image to obtain feature point data; perform face alignment processing on the aforementioned feature point data, and determine the feature data in the face-aligned image as the aforementioned first facial feature data.

[0153] In some embodiments of this disclosure, based on the foregoing scheme, the above-mentioned device further includes: a first audio feature data acquisition module.

[0154] The first audio feature data acquisition module is configured to: sample the speech to obtain a speech sample; perform a short-time Fourier transform on the speech sample to convert it into a speech spectrogram; and filter the speech spectrogram using a Mel filter bank to obtain the audio feature data to be synthesized.

[0155] In some embodiments of this disclosure, based on the foregoing scheme, the above-mentioned device further includes: a first image augmentation module and a first audio augmentation module.

[0156] The first image augmentation module is configured to perform image augmentation processing on the sample image to determine the first facial feature data through the augmented sample image; the first audio augmentation module is configured to perform audio augmentation processing on the acquired speech to determine the audio feature data to be synthesized through the augmented speech.

[0157] In some embodiments of this disclosure, based on the foregoing scheme, the first synthetic data acquisition module 1102 is specifically configured to: convert the audio feature data to be synthesized into target lip shape data according to the pre-trained speech-to-lip synthesis model; fuse the target lip shape data with the first facial feature data according to the pre-trained speech-to-lip synthesis model to obtain a virtual synthetic image; and determine the second facial feature data according to the virtual synthetic image.

[0158] In some embodiments of this disclosure, based on the foregoing scheme, the first recognition model training module 1103 is specifically configured to: determine the i-th set of sample data based on the i-th first face feature data and the i-th second face feature data, wherein the i-th second face data is determined based on the i-th first face feature data; train the classification model through N sets of sample data to obtain the above-mentioned synthetic mouth shape face recognition model, where i is a positive integer not greater than N.

[0159] In some embodiments of this disclosure, based on the foregoing scheme, the first recognition model training module 1103 includes: a sample determination unit, an objective function determination unit, and a training unit.

[0160] The sample determination unit is configured to: determine a first sample based on the first facial feature data, and determine a second sample based on the second facial feature data; the objective function determination unit is configured to: determine a first objective function based on the first sample and its label, and determine a second objective function based on the second sample and its label; determine a model objective function based on the first objective function and the second objective function; and the training unit is configured to: train a classification model by minimizing the function value of the model objective function to obtain a synthetic mouth shape face recognition model.

[0161] In some embodiments of this disclosure, based on the foregoing scheme, the sample determination unit is specifically configured to: rotate the sample image within a preset angle range to obtain a first set of images; scale the sample image within a preset range to obtain a second set of images; and determine the first sample based on the facial feature data corresponding to the images in the first set of images, the facial feature data corresponding to the images in the second set of images, and the first facial feature data.

[0162] In some embodiments of this disclosure, based on the foregoing scheme, the objective function determination unit is specifically configured to: input the first sample into the synthetic mouth-shaped face recognition model to obtain a first discrimination result; and determine the first objective function based on the first discrimination result and the logarithmic loss of the label of the first sample.

[0163] In some embodiments of this disclosure, based on the foregoing scheme, the objective function determination unit is further specifically configured to: determine the first weight of the first objective function and determine the second weight of the second objective function; and determine the model objective function based on the first weight, the first objective function, the second weight, and the second objective function.

[0164] In some embodiments of this disclosure, based on the foregoing scheme, the face recognition module 1104 is specifically configured to: acquire a video to be recognized; extract frames from the video frames of the video to be recognized to obtain the image to be recognized; perform face detection processing on the image to be recognized to determine the face region; perform face feature point detection processing on the face region in the image to be recognized to obtain feature point data; perform face alignment processing on the feature point data in the image to be recognized, and input the target image after face alignment processing after image scaling into the synthetic mouth-shaped face recognition model, and determine whether the image to be recognized is a synthetic mouth-shaped face image based on the recognition result output by the synthetic mouth-shaped face recognition model.

[0165] The specific details of each unit in the above-mentioned synthetic mouth-shaped face recognition device have been described in detail in the synthetic mouth-shaped face recognition method, so they will not be repeated here.

[0166] The following is passed Figure 12 This invention introduces an embodiment of a device for acquiring a synthetic mouth-shaped face recognition model, which can be used to execute the synthetic mouth-shaped face recognition method described above.

[0167] Figure 12 This diagram illustrates the structure of a device for acquiring a synthetic mouth-shape facial recognition model in an exemplary embodiment of this disclosure. Figure 12 As shown, the above-mentioned synthetic mouth shape face recognition model acquisition device 1200 includes: a second real data acquisition module 1201, a second synthetic data acquisition module 1202, and a second recognition model training module 1203.

[0168] The second real data acquisition module 1201 is configured to: perform facial feature point detection on the sample image to obtain first facial feature data; the second synthetic data acquisition module 1202 is configured to: input the first facial feature data and the audio feature data to be synthesized into a pre-trained speech-to-lip synthesis model, and determine the second facial feature data based on the output of the speech-to-lip synthesis model; the second recognition model training module 1203 is configured to: train a classification model based on the first facial feature data and the second facial feature data to obtain a synthetic lip-sync face recognition model.

[0169] In some embodiments of this disclosure, based on the foregoing scheme, the above-mentioned device further includes: a second face region determination module.

[0170] The second face region determination module is configured to: acquire a sample video, extract frames from the video frames of the sample video to obtain the sample image; and perform face detection on the sample image to determine the face region.

[0171] The aforementioned second real data acquisition module 1201 is specifically configured to: perform facial feature point detection on the aforementioned face region in the aforementioned sample image to obtain feature point data; perform face alignment processing on the aforementioned feature point data, and determine the feature data in the face-aligned image as the aforementioned first facial feature data.

[0172] In some embodiments of this disclosure, based on the foregoing scheme, the above-mentioned device further includes: a second audio feature data acquisition module.

[0173] The second audio feature data acquisition module is configured to: sample the speech to obtain a speech sample; perform a short-time Fourier transform on the speech sample to convert it into a speech spectrogram; and filter the speech spectrogram using a Mel filter bank to obtain the audio feature data to be synthesized.

[0174] In some embodiments of this disclosure, based on the foregoing scheme, the above-mentioned device further includes: a second image augmentation module and a second audio augmentation module.

[0175] The second image augmentation module is configured to perform image augmentation processing on the sample image to determine the first facial feature data through the augmented sample image; the second audio augmentation module is configured to perform audio augmentation processing on the acquired speech to determine the audio feature data to be synthesized through the augmented speech.

[0176] In some embodiments of this disclosure, based on the foregoing scheme, the second synthetic data acquisition module 1202 is specifically configured to: convert the audio feature data to be synthesized into target lip shape data according to the pre-trained speech-to-lip synthesis model; fuse the target lip shape data with the first facial feature data according to the pre-trained speech-to-lip synthesis model to obtain a virtual synthetic image; and determine the second facial feature data according to the virtual synthetic image.

[0177] In some embodiments of this disclosure, based on the foregoing scheme, the second recognition model training module 1203 is specifically configured to: determine the i-th set of sample data based on the i-th first face feature data and the i-th second face feature data, wherein the i-th second face data is determined based on the i-th first face feature data; train the classification model through N sets of sample data to obtain the above-mentioned synthetic mouth shape face recognition model, where i is a positive integer not greater than N.

[0178] In some embodiments of this disclosure, based on the foregoing scheme, the second recognition model training module 1203 includes: a sample determination unit, an objective function determination unit, and a training unit.

[0179] The sample determination unit is configured to: determine a first sample based on the first facial feature data, and determine a second sample based on the second facial feature data; the objective function determination unit is configured to: determine a first objective function based on the first sample and its label, and determine a second objective function based on the second sample and its label; determine a model objective function based on the first objective function and the second objective function; and the training unit is configured to: train a classification model by minimizing the function value of the model objective function to obtain a synthetic mouth shape face recognition model.

[0180] In some embodiments of this disclosure, based on the foregoing scheme, the sample determination unit is specifically configured to: rotate the sample image within a preset angle range to obtain a first set of images; scale the sample image within a preset range to obtain a second set of images; and determine the first sample based on the facial feature data corresponding to the images in the first set of images, the facial feature data corresponding to the images in the second set of images, and the first facial feature data.

[0181] In some embodiments of this disclosure, based on the foregoing scheme, the objective function determination unit is specifically configured to: input the first sample into the synthetic mouth-shaped face recognition model to obtain a first discrimination result; and determine the first objective function based on the first discrimination result and the logarithmic loss of the label of the first sample.

[0182] In some embodiments of this disclosure, based on the foregoing scheme, the objective function determination unit is further specifically configured to: determine the first weight of the first objective function and determine the second weight of the second objective function; and determine the model objective function based on the first weight, the first objective function, the second weight, and the second objective function.

[0183] This technical solution uses a synthetic mouth-shaped face recognition model to identify whether an image to be identified belongs to a synthetic mouth-shaped face or a real face. For example, if the image to be identified is a synthetic mouth-shaped face, a warning can be issued to the user to increase their vigilance and reduce the probability of rumors, fraud, and other behaviors caused by the synthesis of video images and audio using artificial intelligence technology. In addition, this technical solution trains a classification model to distinguish between real faces and synthetic mouth-shaped faces using the aforementioned multiple sets of training data (each set of training data includes first face feature data and second face feature data synthesized from the first face feature data and the audio feature data to be synthesized). Compared with related solutions that identify the entire face, this technical solution can obtain more refined face recognition results and has stronger recognition targeting. At the same time, this technical solution uses big data to achieve the recognition of synthetic mouth-shaped faces, which helps to ensure the accuracy of the recognition results.

[0184] The specific details of each unit in the above-mentioned device for acquiring the synthetic mouth shape face recognition model have been described in detail in the method for acquiring the synthetic mouth shape face recognition model, so they will not be repeated here.

[0185] Figure 13 A schematic diagram of a computer system suitable for implementing embodiments of the present disclosure is shown. The electronic device described herein can be... Figure 1 Terminal 110 or server 130 in the middle.

[0186] It should be noted that, Figure 13 The computer system 1300 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.

[0187] like Figure 13As shown, the computer system 1300 includes a processor 1301, which specifically includes a Graphics Processing Unit (GPU) and a Central Processing Unit (CPU). The processor 1301 can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 1302 or programs loaded from storage portion 1308 into Random Access Memory (RAM) 1303. RAM 1303 also stores various programs and data required for system operation. The processor 1301, ROM 1302, and RAM 1303 are interconnected via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.

[0188] In some embodiments, the following components are connected to the I / O interface 1305: an input section 1306 including a keyboard, mouse, etc.; an output section 1307 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a local area network (LAN) card, modem, etc. The communication section 1309 performs communication processing via a network such as the Internet. A drive 1310 is also connected to the I / O interface 1305 as needed. A removable medium 1311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1310 as needed so that computer programs read from it can be installed into the storage section 1308 as needed.

[0189] In particular, according to embodiments of this disclosure, the processes described below with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1309, and / or installed from removable medium 1311. When the computer program is executed by processor 1301, it performs various functions defined in the system of this application.

[0190] It should be noted that the computer-readable medium shown in the embodiments of this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such transmitted data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0191] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0192] The units described in the embodiments of this disclosure can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the unit itself.

[0193] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods described in the above embodiments.

[0194] For example, the electronic device described above can achieve the following: Figure 2 As shown: Step S210, perform facial feature point detection on the sample image to obtain first facial feature data; Step S220, input the first facial feature data and the audio feature data to be synthesized into a pre-trained speech-to-lip synthesis model, and determine the second facial feature data based on the output of the speech-to-lip synthesis model; Step S230, train a classification model based on the first facial feature data and the second facial feature data to obtain a synthesized lip-shaped face recognition model; and Step S240, input the image to be recognized into the synthesized lip-shaped face recognition model, and determine whether the image to be recognized is a synthesized lip-shaped face image based on the output of the synthesized lip-shaped face recognition model.

[0195] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0196] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this disclosure.

[0197] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the preceding claims.

Claims

1. A method for recognizing synthetic mouth-shaped faces, characterized in that, The method includes: Face detection is performed on the sample image to determine the face region; facial feature point detection is performed on the face region in the sample image to obtain feature point data; The feature point data is subjected to face alignment processing, and the feature data in the image after face alignment processing is determined as the first face feature data; The audio feature data to be synthesized is converted into target lip shape data based on the pre-trained speech-to-lip synthesis model; The target lip-sync data is fused with the first facial feature data according to the pre-trained speech-to-lip synthesis model to obtain a virtual synthesized image; The second facial feature data is determined based on the virtual synthesized image; A classification model is trained based on the first facial feature data and the second facial feature data to obtain a synthetic mouth shape facial recognition model; The image to be identified is input into the synthetic mouth-shaped face recognition model, and the output of the synthetic mouth-shaped face recognition model is used to determine whether the image to be identified is a synthetic mouth-shaped face image.

2. The method according to claim 1, characterized in that, Before performing facial feature point detection on the sample image to obtain the first facial feature data, the method further includes: A sample video is acquired, and the video frames of the sample video are extracted to obtain the sample image.

3. The method according to claim 1, characterized in that, The method further includes: Speech samples are obtained by sampling the speech. The speech samples are subjected to a short-time Fourier transform to convert the speech samples into a speech spectrogram; The speech spectrogram is filtered using a Mel filter bank to obtain the audio feature data to be synthesized.

4. The method according to claim 1, characterized in that, Before inputting the first facial feature data and the audio feature data to be synthesized into the pre-trained speech-to-lip synthesis model, the method further includes: The sample image is subjected to image augmentation processing in order to determine the first facial feature data through the augmented sample image; The acquired speech is subjected to audio augmentation processing in order to determine the audio feature data to be synthesized from the augmented speech.

5. The method according to claim 1, characterized in that, A classification model is trained based on the first facial feature data and the second facial feature data to obtain a synthetic mouth-shape facial recognition model, including: The i-th group of sample data is determined based on the i-th first face feature data and the i-th second face feature data, wherein the i-th second face feature data is determined based on the i-th first face feature data; The synthetic mouth shape face recognition model is obtained by training a classification model with N sets of sample data, where i is a positive integer not greater than N.

6. The method according to claim 1, characterized in that, A classification model is trained based on the first facial feature data and the second facial feature data to obtain a synthetic mouth-shape facial recognition model, including: A first sample is determined based on the first facial feature data, and a second sample is determined based on the second facial feature data; A first objective function is determined based on the first sample and its label, and a second objective function is determined based on the second sample and its label. Determine the model objective function based on the first objective function and the second objective function; A synthetic mouth-shaped face recognition model is obtained by training a classification model by minimizing the function value of the objective function of the model.

7. The method according to claim 6, characterized in that, Determining the first sample based on the first facial feature data includes: The sample image is rotated within a preset angle range to obtain the first set of images; The sample images are scaled within a preset range to obtain a second set of images; The first sample is determined based on the facial feature data corresponding to the images in the first group of images, the facial feature data corresponding to the images in the second group of images, and the first facial feature data.

8. The method according to claim 6, characterized in that, Determining a first objective function based on the first sample and its label includes: The first sample is input into the synthetic mouth shape face recognition model to obtain the first discrimination result; The first objective function is determined based on the first discrimination result and the log loss of the label of the first sample.

9. The method according to claim 6, characterized in that, The model objective function is determined based on the first objective function and the second objective function, including: Determine the first weight of the first objective function, and determine the second weight of the second objective function; The model objective function is determined based on the first weight, the first objective function, the second weight, and the second objective function.

10. The method according to any one of claims 1 to 9, characterized in that, The process involves inputting the image to be recognized into the synthetic mouth-shaped face recognition model, and determining whether the image to be recognized is a synthetic mouth-shaped face image based on the output of the synthetic mouth-shaped face recognition model, including: The video to be identified is obtained, and the video frames of the video to be identified are extracted to obtain the image to be identified; The image to be identified is subjected to face detection processing to determine the face region; Facial feature point detection processing is performed on the face region in the image to be identified to obtain feature point data; The feature point data in the image to be identified is subjected to face alignment processing, and the target image after face alignment processing is scaled and input into the synthetic mouth shape face recognition model. The recognition result output by the synthetic mouth shape face recognition model is used to determine whether the image to be identified is a synthetic mouth shape face image.

11. A method for obtaining a synthetic mouth-shaped face recognition model, characterized in that, The method includes: Face detection is performed on the sample image to determine the face region; facial feature point detection is performed on the face region in the sample image to obtain feature point data; The feature point data is subjected to face alignment processing, and the feature data in the image after face alignment processing is determined as the first face feature data; The audio feature data to be synthesized is converted into target lip shape data based on the pre-trained speech-to-lip synthesis model; The target lip-sync data is fused with the first facial feature data according to the pre-trained speech-to-lip synthesis model to obtain a virtual synthesized image; The second facial feature data is determined based on the virtual synthesized image; A classification model is trained based on the first facial feature data and the second facial feature data to obtain a synthetic mouth shape facial recognition model.

12. A device for recognizing a synthetic mouth-shaped face, characterized in that, The device includes: The first real data acquisition module is configured to: perform face detection on the sample image to determine the face region; perform face feature point detection on the face region in the sample image to obtain feature point data; perform face alignment processing on the feature point data, and determine the feature data in the face-aligned image as the first face feature data; The first synthetic data acquisition module is configured to: convert the audio feature data to be synthesized into target lip shape data according to a pre-trained speech-to-lip shape synthesis model; fuse the target lip shape data with the first facial feature data according to the pre-trained speech-to-lip shape synthesis model to obtain a virtual synthetic image; and determine the second facial feature data according to the virtual synthetic image. The first recognition model training module is configured to: train a classification model based on the first face feature data and the second face feature data to obtain a synthetic mouth shape face recognition model; The face recognition module is configured to: input the image to be recognized into the synthetic mouth-shaped face recognition model, and determine whether the image to be recognized is a synthetic mouth-shaped face image based on the output of the synthetic mouth-shaped face recognition model.

13. A device for acquiring a synthetic mouth-shaped facial recognition model, characterized in that, The device includes: The second real data acquisition module is configured to: perform face detection on the sample image to determine the face region; perform face feature point detection on the face region in the sample image to obtain feature point data; perform face alignment processing on the feature point data, and determine the feature data in the face-aligned image as the first face feature data; The second synthetic data acquisition module is configured to: convert the audio feature data to be synthesized into target lip shape data according to a pre-trained speech-to-lip shape synthesis model; fuse the target lip shape data with the first facial feature data according to the pre-trained speech-to-lip shape synthesis model to obtain a virtual synthetic image; and determine the second facial feature data according to the virtual synthetic image. The second recognition model training module is configured to: train a classification model based on the first facial feature data and the second facial feature data to obtain a synthetic mouth shape facial recognition model.

14. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the synthetic mouth-shaped face recognition method as described in any one of claims 1 to 10, or to implement the synthetic mouth-shaped face recognition model acquisition method as described in claim 11.