Face-forgery detection by face feature watermarking

Face feature watermarking techniques encode unique face features into images for independent detection, addressing the limitations of existing methods by enabling accurate identification of manipulated images without source watermark reliance.

WO2026044437A1PCT designated stage Publication Date: 2026-03-05QUALCOMM INC +5
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/114416
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing face-forgery detection methods, including passive and proactive approaches, struggle to accurately identify images manipulated by advanced artificial intelligence content generation (AICG) technologies like deepfakes and face swaps, and require access to source watermarks that can be compromised by nefarious actors.

Method used

A method using face feature watermarking that encodes unique face features into images, allowing independent detection without reliance on source watermarks, utilizing invertible mapping models for encoding and decoding, and comparing face features to determine authenticity.

Benefits of technology

Enables effective detection of manipulated images without requiring shared source watermarks, enhancing usability and flexibility by allowing independent verification on various devices, and overcoming limitations of current methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024114416_05032026_PF_FP_ABST
    Figure CN2024114416_05032026_PF_FP_ABST
Patent Text Reader

Abstract

Certain aspects of the present disclosure provide techniques for face-forgery detection. Aspects include a method for face-forgery detection by an apparatus including obtaining an image comprising a face; extracting, with a face feature extraction process, a face feature from the image; generating a watermark based on the face feature; encoding the watermark into the image to generate a watermarked image; and storing the watermarked image in the one or more memories.
Need to check novelty before this filing date? Find Prior Art

Description

FACE-FORGERY DETECTION BY FACE FEATURE WATERMARKING

[0001] INTRODUCTION

[0002] Field of the Disclosure

[0003] Aspects of the present disclosure relate to techniques for face-forgery detection.

[0004] Description of Related Art

[0005] Artificial intelligence content generation (AICG) technology enable users to easily generate content in text, image, video, and other formats. The AICG technology may generate original content or may enable a user to perform content manipulation on existing content. For example, a user may utilize AICG technology to edit images or videos. The use of AICG technology can reduce editing time, improve the quality of the edited content, and even enable users who may not be skilled in content generation or content editing to perform tasks that would otherwise be reserved to those highly-skilled or talented, such as graphic designers and artists.

[0006] Unfortunately, AICG technology provides the same advantages to nefarious actors who may desire to generate or manipulate content for malicious activities. For example, AICG technology, such as deep-learning-based technologies, which include deepfakes and face swaps are attracting attention in both society and academia, particularly ones used to synthesize face-forged images.

[0007] Accordingly, there is a need to advance the detection and identification of original content from manipulated content.SUMMARY

[0008] One aspect provides a method for face-forgery detection by an apparatus. The method includes obtaining an image comprising a face; extracting, with a face feature extraction process, a face feature from the image; generating a watermark based on the face feature; encoding the watermark into the image to generate a watermarked image; and storing the watermarked image in the one or more memories.

[0009] Another aspect provides a method for face-forgery detection by an apparatus. The method includes obtaining a watermarked image comprising a face; decoding the watermarked image to obtain a watermark, wherein the watermark is based on a first face  feature corresponding to the face in the watermarked image; resolving the first face feature from the watermark; extracting, with a face feature extraction process, a second face feature from the watermarked image; comparing the first face feature to the second face feature to determine whether the first face feature and the second face feature match within a degree of similarity; and one of: determining that the watermarked image is a real image when the match of the first face feature and the second face feature is within the degree of similarity, or determining that the watermarked image is an altered image when the match of the first face feature and the second face feature is less than the degree of similarity.

[0010] Other aspects provide: one or more apparatuses operable, configured, or otherwise adapted to perform any portion of any method described herein (e.g., such that performance may be by only one apparatus or in a distributed fashion across multiple apparatuses) ; one or more non-transitory, computer-readable media comprising instructions that, when executed by one or more processors of one or more apparatuses, cause the one or more apparatuses to perform any portion of any method described herein (e.g., such that instructions may be included in only one computer-readable medium or in a distributed fashion across multiple computer-readable media, such that instructions may be executed by only one processor or by multiple processors in a distributed fashion, such that each apparatus of the one or more apparatuses may include one processor or multiple processors, and / or such that performance may be by only one apparatus or in a distributed fashion across multiple apparatuses) ; one or more computer program products embodied on one or more computer-readable storage media comprising code for performing any portion of any method described herein (e.g., such that code may be stored in only one computer-readable medium or across computer-readable media in a distributed fashion) ; and / or one or more apparatuses comprising one or more means for performing any portion of any method described herein (e.g., such that performance would be by only one apparatus or by multiple apparatuses in a distributed fashion) . By way of example, an apparatus may comprise a processing system, a device with a processing system, or processing systems cooperating over one or more networks.

[0011] The following description and the appended figures set forth certain features for purposes of illustration.BRIEF DESCRIPTION OF DRAWINGS

[0012] The appended figures depict certain features of the various aspects described herein and are not to be considered limiting of the scope of this disclosure.

[0013] FIG. 1 depicts an example progression of the creation, manipulation, and manipulation detection of an image.

[0014] FIG. 2 depicts illustrative examples of a passive detection framework and a proactive detection framework.

[0015] FIG. 3 depicts an illustrative face-forgery detection framework.

[0016] FIG. 4 depicts a method for face-forgery detection.

[0017] FIG. 5 depicts another method for face-forgery detection.

[0018] FIG. 6 depicts aspects of an example user equipment.

[0019] FIG. 7 depicts aspects of an example user equipment.DETAILED DESCRIPTION

[0020] Aspects of the present disclosure provide apparatuses, methods, processing systems, and computer-readable mediums for face-forgery detection. In certain aspects, the techniques for face-forgery detection include processes for encoding an image with a (e.g., unique) watermark based on face features extracted from the image. In some aspects, the techniques for face-forgery detection include processes for decoding a face feature from a watermark encoded in the image and comparing the face feature with another face feature extracted from the image to determine whether the image is real or fake based on a similarity match between the decoded feature and the extracted feature.

[0021] Certain aspects of the face-forgery detection techniques described herein may provide several advantages and improvements over exiting processes for attempting to determine whether an image has been maliciously transformed. For example, as discussed in more detail herein, in certain aspects, the face-forgery detection techniques may not require that a source watermark be communicated between a device that creates the image and those devices that view the image. Accordingly, this may reduce the possibility that nefarious actors obtain the source watermark and integrate the source watermark into the maliciously transformed image, thereby defeating the beneficial effect of utilizing the source watermark. Additionally, in certain aspects, devices do not need  to generate, share, and store private and public keys associated with source watermarks, for example, in order to carry out face-forgery detection processes. Accordingly, in certain aspects, the face-forgery detection techniques described herein may provide the technical benefit of being able to independently protect (e.g., encode) an image and independently determine whether an image has been transformed, for example by AICG manipulation.

[0022] In certain aspects, another advantage of the face-forgery detection techniques described herein may be that the techniques are capable of detecting face-forged images that were generated by AICG technology, which is a shortcoming of both current passive and proactive methods, which are described generally with reference to FIG. 2.

[0023] FIG. 1 depicts an example progression of the creation, manipulation, and manipulation detection of an image. A computing device, such as a mobile phone, tablet, server, desktop, user equipment, or other computer, may capture an image of a person. The original image 102 that is captured by the device may be protected through one of a variety of protection processes depicted by block 104. It may be desired that the quality and / or data size of a watermarked image 106 that is protected not be adversely effected by the protection process. Some existing protection processes can inject vast amounts of information into the image thereby reducing the image quality and / or undesirably increasing the data size of the image. Additionally, some existing protection processes require a private watermark for forgery detection, thereby limiting their usability as devices that are not equipped with the corresponding public watermark cannot make a determination as to the legitimacy of the image. In certain aspects, the face-forgery detection techniques discussed herein may provide solutions to these issues.

[0024] The watermarked image 106 that is generated by the protection process at block 104 may then be transmitted or otherwise communicated from the device that originally captured and / or created it. For example, the watermarked image 106 may be posted to a social media platform, stored in a cloud based server or other data storage device, transmitted to one or more other user equipments, or the like. Now that the watermarked image 106 is no longer secured by data protections afforded by the user’s device, the watermarked image 106 may be readily subjected to malicious and / or benign transformations. Malicious transformations, for example, at block 108, refer to manipulations such as the generation of deepfakes and / or face swaps using AICG technology such as deep-learning-based technologies. Deepfakes refer to digital  manipulation of an image, for example, manipulating an image that includes replacing the likeness of a person with the likeness of another such that the manipulated image appears real. Face swaps refer more specifically to the swapping of a person’s face in an image with the face of another.

[0025] For example, a nefarious actor may provide an AICG process, for example at block 108, with the watermarked image 106 containing a face of a first person and a replacement image 110 containing a face of a second person to replace with the face of the first person in the watermarked image 106. The AICG process may effectively and with high-levels of reality generate a forged image 112. Current face-forgery detection processes, which include both passive and proactive detection processes may struggle to accurately detect whether the forged image 112 is manipulated. Since AICG technology can generate face forged content with high-levels of reality, passive detection processes, which focus on detecting local artifacts, may fail to determine that an image is forged because there are typically undetectable artifacts as a result of the AICG technology based manipulation. Proactive detection processes may be able to detect face-forged images, but may require that the device configured to perform the determination have access to the source watermark that was used. However, in some instances, the source watermark may be known to the nefarious actor and subsequently incorporated into the manipulated image thereby spoofing the proactive detection process from accurately determining that the image was manipulated.

[0026] Benign transformations performed, for example, at block 116, may include manipulations such as cropping the image, adding noise to the image, compressing the image, or the like. These transformations may be determined with or without face-forgery detection processes and are not necessarily ones arising from nefarious activity, such as manipulation of the content of the image.

[0027] FIG. 2 depicts illustrative examples of a passive detection framework and a proactive detection framework. The passive detection framework 210 includes a device such as a first user equipment 202 that captures or otherwise obtains an image 102 and then shares the image 102, for example, onto an internet based platform 204 or with another device, such as a second user equipment 206. The first user equipment 202 and the second user equipment 206, collectively referred to as user equipments, may be one of a number of devices such as a mobile phone, tablet, laptop, or other computing device.

[0028] The user equipments may be any device or combination of components comprising one or more processors and non-transitory computer readable memory, referred to herein as one or more memories. The one or more processors may be any device capable of executing the processor-executable instructions stored in the one or more memories. Accordingly, the one or more processors may be an electric controller, an integrated circuit, a microchip, a computer, or any other computing device. The one or more processors are communicatively coupled to the other components, such as one or more cameras and / or network interface hardware by the communication path. Accordingly, the communication path may communicatively couple any number of processors with one another, and allow the components coupled to the communication path to operate in a distributed computing environment. Specifically, each of the components may operate as a node that may send and / or receive data.

[0029] The one or more memories may comprise RAM, ROM, flash memories, hard drives, or any non-transitory memory device capable of storing processor-executable instructions such that the processor-executable instructions can be accessed and executed by the one or more processors. The machine-readable instruction set may comprise logic or algorithm (s) written in any programming language of any generation (e.g., 1GL, 2GL, 3GL, 4GL, or 5GL) such as, for example, machine language that may be directly executed by the one or more processors, or assembly language, object-oriented programming (OOP) , scripting languages, microcode, etc., that may be compiled or assembled into processor-executable instructions and stored in the one or more memories. Alternatively, the processor-executable instructions may be written in a hardware description language (HDL) , such as logic implemented via either a field-programmable gate array (FPGA) configuration or an application-specific integrated circuit (ASIC) , or their equivalents. Accordingly, the functionality described herein may be implemented in any conventional computer programming language, as pre-programmed hardware elements, or as a combination of hardware and software components.

[0030] The internet based platform 204 may include social media, blogs, data storage devices, cloud storage devices, or the like. Nefarious actors can obtain images (or videos) from the internet based platform 204 and perform benign or malicious transformations, for example, as discussed with reference to blocks 108 and 116 depicted in FIG. 1.

[0031] A second user equipment 206 may retrieve or view a perceived image 122 (e.g., a potentially manipulated image) provided through the internet based platform 204.  Under a passive detection framework, the second user equipment 206 may implement a classifier at block 124. The classifier may include a face recognition algorithm that is configured to perform geometry-based or a template-based analysis of the image content being viewed to make a determination as to whether the image is real or fake at block 126. Passive methods focus on detecting local artifacts, which are not effective for modern generative models, such as AICG technology.

[0032] An illustrative example of a proactive detection framework 220 is also depicted in FIG. 2. The proactive detection framework 220 includes a first user equipment 202 that captures or otherwise obtains an image 102. The captured image 102 may then be encoded at block 130 by an encoder with a source watermark 132. The source watermark 132 may be a predetermined private or public key based watermark or the source watermark 132 may be randomly generated data. In either case, the source watermark 132 is not based on or derived from content in the image 102. A watermarked image 134 is output by the encoder. The first user equipment 202 may then share the image 102, for example, onto an internet based platform 204 or with another device, such as a second user equipment 206. The internet based platform 204 may include social media, blogs, data storage devices, cloud storage devices, or the like. Nefarious actors can obtain images (or videos) from the internet based platform 204 and perform benign or malicious transformations, for example, as discussed with reference to blocks 108 and 116 depicted in FIG. 1.

[0033] A second user equipment 206 may retrieve or view a perceived image 136 (e.g., a potentially manipulated image) provided through the internet based platform 204. Under the proactive detection framework, the second user equipment 206 is configured to decode the perceived image 136 at block 138 with a decoder to obtain a decoded watermark 140. The second user equipment 206 additionally needs to be capable of retrieving the source watermark 132 that was encoded in the image 102 to compare to the decoded watermark 140 at block 142 in order to make a determination as to whether the perceived image 136 is real or fake.

[0034] While the proactive detection framework 220 may add visually imperceptible perturbation to images for forgery detection, which are more effective than passive methods, proactive detection frameworks rely on potentially interceptable source watermarks or source watermarks that can be backward engineered and integrated into a manipulated image 136 because they are not based on unique features of the image 102.

[0035] In certain aspects, the face-forgery detection techniques, which will now be described in detail, do not require that each user equipment that is generating a protected image or determining whether a protected image has been maliciously manipulated have knowledge of the source watermark. Instead, in certain aspects, the process of encoding a watermark in an image and decoding the watermark from the image can be independently performed by each user equipment (e.g., the first user equipment 202 and the second user equipment 206) based solely on the image obtained by each user equipment.

[0036] FIG. 3 depicts an illustrative face-forgery detection framework 300. The face-forgery detection framework 300 includes a first user equipment 202 that captures an image 102. The first user equipment 202 proceeds with extracting one or more face features 304 from the image 102 at block 302, using a face feature extraction model, such as FaceNet. FaceNet is a facial recognition system that uses a deep convolutional neural network to learn a mapping (also referred to as an embedding) from a set of face images to a 128-dimensional Euclidean space, and assesses the similarity between faces based on the square of the Euclidean distance between the images’ corresponding normalized vectors in the 128-dimensional Euclidean space. FaceNet uses the triplet loss function as its cost function and an online triplet mining method. The face feature extraction model may be or include one or more algorithms and / or trained models configured to extract one or more face features such as eyes, nose, mouth, etc. from a human face. The face feature extraction model may also be configured to detect faces in an image, align the detected faces based on a face shape, for example, determined based on a location of the detected face in the image and a size of the detected face in the image. In aspects, face alignment methods may be implemented. For example, face alignment methods may normalize facial images by correcting for variations caused by lighting conditions, head poses, and facial expressions. Such methods may use facial landmarks like the tip of the nose, eye centers, and mouth corners to determine the face’s shape and components, like the eyes and nose, based on translation, scale, and rotation. Some face alignment methods may use deformable models that encode prior knowledge of face shape or appearance. These models are then adjusted iteratively to account for low-level image evidence and find the face in the image. For example, if a person's eyes in an image are at an angle to the image frame, the technique will rotate the image so that the angle is 180 degrees. The one or more face features 304 extracted from the image 102 may be defined by values  derived from the type of feature, the location of the feature, and / or other quantifiable characteristic generated from the feature extraction process. For example, aspects such as the width between eyes, the curvature of a nose, one or proportions between one or more of the eyes, nose, cheeks, mouth, and / or chin may be quantified into one or more values that correspond to the one or more face features 304 used for watermarking the image. In certain aspects, face features are extracted by deep learning models from the entire face region. The face features may be a unified mathematical representation of what a face looks like. Accordingly, in certain aspects, the face features do not need to have explicit correspondences to nose, eyes, or the like.

[0037] At block 306, the first user equipment 202 proceeds with generating a watermark 308 based on the one or more face features 304. The process of generating the watermark may include implementing an invertible mapping model that maps the one or more face features 304 to the watermark 308. The invertible mapping model (e.g., feature-to-watermark, F2W) may define a map between vector spaces of the one or more features to vector spaces of the image. Since the map is linear in nature, an inverse of the map be determined through a linear transformation, such that the mapping process implemented in block 306 may be reversed, for example and utilized in block 330 discussed herein. The invertible mapping model may include two or more fully connected layers of a neural network in which each neuron applies a linear transformation to the input vector (e.g., the vector space corresponding to the one or more features) through a weights matrix. In certain aspects, the one or more face features 304 are sized and mapped to a watermark 308 that is the same size (e.g., pixel height x pixel width) as the image 102.

[0038] At block 310, the first user equipment 202 proceeds with encoding the image 102 with the watermark 308 to generate a watermarked image 312 that may not be visually distinct from the image 102 that was originally obtained by the first user equipment 202. For example, the encoder implemented for encoding the image 102 may be the encoder of implemented in the U-Net architecture.

[0039] The first user equipment 202 may then share or otherwise communicate the watermarked image 312, for example, onto an internet based platform 204 or with another device, such as a second user equipment 206. The internet based platform 204 may include social media, blogs, data storage devices, cloud storage devices, or the like. Nefarious actors may obtain images (or videos) from the internet based platform 204 and  perform benign or malicious transformations, for example, as discussed with reference to blocks 108 and 116 depicted in FIG. 1.

[0040] A second user equipment 206 may obtain, such as retrieve or view, a perceived image 320 (e.g., a potentially manipulated image) , such as provided through the internet based platform 204 or obtained elsewhere. At block 322, the second user equipment 206 proceeds with extracting one or more first face features 324 from the perceived image 122 using a face feature extraction model. The first face features 324 may also be referred to as explicit features, which refers to features on the surface of a face. The explicit features may change if a face swap is performed on the image. The face feature extraction model may be the same face feature extraction model as used by the first user equipment 202 at block 302. Consistency between the models that are used by the first user equipment 202 and the second user equipment 206 can improve the determination as to whether the perceived image 320 was manipulated.

[0041] At block 326, the second user equipment 206 is configured to decode the perceived image 320 with a decoder, which is a counterpart to the encoder used by the first user equipment 202, to generate a decoded watermark 328. For example, the decoder implemented for encoding the image 102 may be the decoder of implemented in the U-Net architecture. In certain aspects, the decoder generates a vector space corresponding to the decoded watermark 328 from the perceived image 320. In certain aspects, the decoded watermark 328 is extracted by decoding the entire perceived image 320 and not just the spaces corresponding to face features. In certain aspects, the decoder may reliably recover the embedded watermark no matter whether the image has gone through benign or malicious transformations.

[0042] In certain aspects, the one or more second face features 332 that the decoded watermark 328 is based on need to be resolved (e.g., determined) so that they can be compared to the one or more first face features 324. The second face features 332 may also be referred to as implicit features. Implicit features refer to face features that are hidden in the whole image. These are unknown and not visible to observation of an image, thus difficult to change. At block 330, the second user equipment 206 is configured to resolve the one or more second face features 332 (e.g., watermark-to-feature, W2F) from the decoded watermark 328, for example, using the invertible mapping model. Implementation of the invertible mapping model at block 330 may be  the inverted form of the same invertible mapping model used by the first user equipment 202 at block 306.

[0043] At block 334, the second user equipment 206 is configured to compare the one or more first face features 324 to the one or more second face features 332 to determine whether the features match. The determination as to whether the features match may be made based on a degree of similarity (e.g., threshold) . The degree of similarity may be set based on potential benign differences in the such as noise introduced from either the decoding and resolving processes at blocks 326 and 330 and / or the face feature extraction process corresponding to block 322. The process of comparing the one or more first face features 324 to the one or more second face features 332 may include determining a cosine similarity. For example, cosine similarity may be represented by the Cosine similarity function:

[0044] where Ai and Bi are the ith component of n-dimensional vectors A and B, and θ represents the degree of similarity as an angle between vectors A and B, which represent vector embeddings of one or more first face features 324 and the one or more second face features 332.

[0045] In certain aspects, the second user equipment 206 is configured to determine that the watermarked image is a real image when the match of the one or more first face features 324 (e.g., first face feature (s) ) to the one or more second face features 332 (e.g., second face feature (s) ) is within the degree of similarity at block 336. In certain aspects, the second user equipment 206 is configured to determine that the watermarked image is an altered image when the match of the one or more first face features 324 (e.g., first face feature (s) ) to the one or more second face features 332 (e.g., second face feature (s)) is less than the degree of similarity at block 336. The degree of similarity may be set at or about a predefined threshold, for example, 0.9 or greater.

[0046] Aspects of the face-forgery detection framework 300 described herein use face features as unique watermarks for image protection. This enables face-forgery detection from image itself without the need for knowledge of a source watermark thereby significantly expanding its usability. That is, a decoding device needs to only be  configured to implement the verification stage, for example, blocks 322, 326, 330, and 334 depicted and described with reference to FIG. 3 to make a real or fake determination.

[0047] Additionally, by using face features as the basis for the watermark in the face-forgery detection process, the user equipments (e.g., the first user equipment 202 and the second user equipment 206) do not have to maintain a watermark database for forgery detection purpose. The simple framework and feature design enable many different user equipments to implement the framework. Furthermore, the decoupled nature of the components between the encoding and decoding operations enable the face-forgery detection framework 300 to be flexible and independently upgradable. That is, in certain aspects, the face detection model, face recognition model, image watermarking models may be independent components. As such, in certain aspects, there is no correlation between them. Accordingly, in certain aspects, each model can be replaced with a newer or different version of the respective model without affected the other models. While aspects of the face-forgery detection framework 300 described herein are directed to utilizing features related to human faces, the face-forgery detection framework 300 may be configured to utilize features other than those of a human face. For example, some aspects may utilize features extracted from a human form, or proportions corresponding to other objects or animals.

[0048] Example Methods for Face-Forgery Detection

[0049] FIG. 4 shows a method 400 for face-forgery detection by an apparatus, such as the first user equipment 202 of FIG. 3.

[0050] Method 400 begins at block 405 with obtaining an image comprising a face. For example, block 405 may correspond to obtaining image 102 discussed with reference to FIG. 3.

[0051] Method 400 then proceeds to block 410 with extracting, with a face feature extraction process, a face feature from the image. For example, block 410 may include utilizing the face feature extraction model discussed with reference to block 302 of the face-forgery detection framework 300 of FIG. 3. The face feature extraction process of block may include extracting the one or more face features 304 from the image 102.

[0052] Method 400 then proceeds to block 415 with generating a watermark based on the face feature. For example, block 415 may include the first user equipment 202 implementing an invertible mapping model that maps the one or more face features 304  to the watermark 308 as discussed with reference to block 306 of the face-forgery detection framework 300 of FIG. 3.

[0053] Method 400 then proceeds to block 420 with encoding the watermark into the image to generate a watermarked image. For example, block 420 may include the first user equipment 202 encoding the image 102 with the watermark 308 to generate a watermarked image 312 as discussed with reference to block 310 of the face-forgery detection framework 300 of FIG. 3.

[0054] Method 400 then proceeds to block 425 with storing the watermarked image in the one or more memories. In some aspects, the watermarked image 312, for example, may be stored in the one or more memories of the first user equipment 202 or other device.

[0055] In one aspect, block 410 includes: detecting the face in the image; aligning the detected face based on a face shape determined based on a location of the detected face in the image and a size of the detected face in the image; and identifying the face feature to extract from the image from the detected and aligned face.

[0056] In one aspect, the watermark comprises a deterministic value based on the face feature.

[0057] In one aspect, the watermark comprises a floating-point value based on the face feature.

[0058] In one aspect, block 415 includes utilizing an invertible mapping model.

[0059] In one aspect, the invertible mapping model is configured to transform the face feature into the watermark and scale the watermark to correspond to a pixel size of the image.

[0060] In one aspect, method 400 further includes transmitting the watermarked image.

[0061] In one aspect, method 400, or any aspect related to it, may be performed by an apparatus, such as apparatus 600 of FIG. 6, which includes various components operable, configured, or adapted to perform the method 400. Apparatus 600 is described below in further detail.

[0062] Note that FIG. 4 is just one example of a method, and other methods including fewer, additional, or alternative operations are possible consistent with this disclosure.

[0063] FIG. 5 shows a method 500 for face-forgery detection by an apparatus, such as the second user equipment 206 of FIG. 3.

[0064] Method 500 begins at block 505 with obtaining a watermarked image comprising a face. For example, block 505 may correspond to obtaining the perceived image 320 discussed with reference to FIG. 3.

[0065] Method 500 then proceeds to block 510 with decoding the watermarked image to obtain a watermark, wherein the watermark is based on a first face feature corresponding to the face in the watermarked image. For example, block 510 may include the second user equipment 206 decoding the perceived image 320 to obtain the watermark 328 as discussed with reference to block 326 of the face-forgery detection framework 300 of FIG. 3.

[0066] Method 500 then proceeds to block 515 with resolving the first face feature from the watermark. For example, block 515 may include the second user equipment 206 being configured to resolve the one or more second face features 332 (e.g., watermark-to-feature, W2F) from the decoded watermark 328, for example, using the invertible mapping model as discussed with reference to block 330 of the face-forgery detection framework 300 of FIG. 3.

[0067] Method 500 then proceeds to block 520 with extracting, with a face feature extraction process, a second face feature from the watermarked image. For example, block 520 may include utilizing the face feature extraction model discussed with reference to block 322 of the face-forgery detection framework 300 of FIG. 3. The face feature extraction process of block may include extracting the one or more first features 324 from the perceived image 320.

[0068] Method 500 then proceeds to block 525 with comparing the first face feature to the second face feature to determine whether the first face feature and the second face feature match within a degree of similarity. For example, at block 525, the second user equipment 206 may be configured to compare the one or more first face features 324 to the one or more second face features 332 to determine whether the features match as described with reference to block 334 of the face-forgery detection framework 300 of FIG. 3.

[0069] Method 500 then proceeds to block 530. At block 530 the second user equipment 206 may be configured for either one of: determining that the watermarked  image is a real image when the match of the first face feature and the second face feature is within the degree of similarity; or determining that the watermarked image is an altered image when the match of the first face feature and the second face feature is less than the degree of similarity. Block 530 may correspond to block 334 of the face-forgery detection framework 300 of FIG. 3.

[0070] In one aspect, block 520 includes: detecting the face in the watermarked image; aligning the detected face based on a face shape determined based on a location of the detected face in the watermarked image and a size of the detected face in the watermarked image; and identifying the first face feature to extract from the watermarked image from the detected and aligned face.

[0071] In one aspect, the watermark comprises a deterministic value based on the first face feature.

[0072] In one aspect, the watermark comprises a floating-point value based on the first face feature.

[0073] In one aspect, block 515 includes utilizing an invertible mapping model.

[0074] In one aspect, the invertible mapping model is configured to transform the watermark into the first face feature and scale the first face feature to correspond to a pixel size of the face in the watermarked image.

[0075] In one aspect, block 525 includes determining cosine similarity between the first face feature and the second face feature.

[0076] In one aspect, block 505 includes retrieving the watermarked image from a data storage.

[0077] In one aspect, method 500, or any aspect related to it, may be performed by an apparatus, such as apparatus 700 of FIG. 7, which includes various components operable, configured, or adapted to perform the method 500. Apparatus 700 is described below in further detail.

[0078] Note that FIG. 5 is just one example of a method, and other methods including fewer, additional, or alternative operations are possible consistent with this disclosure.

[0079] Example Devices

[0080] FIG. 6 depicts aspects of an example apparatus 600. In some aspects, apparatus 600 is a user equipment, such as first user equipment 202 described above with respect to FIG. 3.

[0081] The apparatus 600 includes a processing system 602 coupled to a transceiver 650 (e.g., a transmitter and / or a receiver) . The transceiver 650 is configured to transmit and receive signals for the apparatus 600 via an antenna 652, such as the various signals as described herein. The processing system 602 may be configured to perform processing functions for the apparatus 600, including processing signals received and / or to be transmitted by the apparatus 600.

[0082] The processing system 602 includes one or more processors 604. The one or more processors 604 are coupled to a computer-readable medium / memory 626 via a bus 648. In certain aspects, the computer-readable medium / memory 626 is configured to store instructions (e.g., computer-executable code) , including code 628-646, that when executed by the one or more processors 604, enable and cause the one or more processors 604 to perform the method 400 described with respect to FIG. 4, or any aspect related to it, including any operations described in relation to FIG. 4. Note that reference to a processor performing a function of apparatus 600 may include one or more processors performing that function of apparatus 600, such as in a distributed fashion.

[0083] In the depicted example, computer-readable medium / memory 626 stores code for obtaining 628, code for extracting 630, code for generating 632, code for encoding 634, code for storing 636, code for detecting 638, code for aligning 640, code for identifying 642, code for utilizing 644, and code for transmitting 646. Processing of the code 628-646 may enable and cause the apparatus 600 to perform the method 400 described with respect to FIG. 4, or any aspect related to it.

[0084] The one or more processors 604 include circuitry configured to implement (e.g., execute) the code (e.g., executable instructions) stored in the computer-readable medium / memory 626, including circuitry for obtaining 606, circuitry for extracting 608, circuitry for generating 610, circuitry for encoding 612, circuitry for storing 614, circuitry for detecting 616, circuitry for aligning 618, circuitry for identifying 620, circuitry for utilizing 622, and circuitry for transmitting 624. Processing with circuitry 606-624 may  enable and cause the apparatus 600 to perform the method 400 described with respect to FIG. 4, or any aspect related to it.

[0085] More generally, means for communicating, transmitting, sending or outputting for transmission may include the transceiver 650 and / or antenna 652 of the apparatus 600 in FIG. 6, and / or one or more processors 604 of the apparatus 600 in FIG. 6. Means for communicating, receiving or obtaining may include the transceiver 650 and / or antenna 652 of the apparatus 600 in FIG. 6, and / or one or more processors 604 of the apparatus 600 in FIG. 6.

[0086] FIG. 7 depicts aspects of an example apparatus 700. In some aspects, apparatus 700 is a user equipment, such as the second user equipment 206 described above with respect to FIG. 3.

[0087] The apparatus 700 includes a processing system 702 coupled to a transceiver 754 (e.g., a transmitter and / or a receiver) . The transceiver 754 is configured to transmit and receive signals for the apparatus 700 via an antenna 756, such as the various signals as described herein. The processing system 702 may be configured to perform processing functions for the apparatus 700, including processing signals received and / or to be transmitted by the apparatus 700.

[0088] The processing system 702 includes one or more processors 704. The one or more processors 704 are coupled to a computer-readable medium / memory 728 via a bus 752. In certain aspects, the computer-readable medium / memory 728 is configured to store instructions (e.g., computer-executable code) , including code 730-750, that when executed by the one or more processors 704, enable and cause the one or more processors 704 to perform the method 500 described with respect to FIG. 5, or any aspect related to it, including any operations described in relation to FIG. 5. Note that reference to a processor performing a function of apparatus 700 may include one or more processors performing that function of apparatus 700, such as in a distributed fashion.

[0089] In the depicted example, computer-readable medium / memory 728 stores code for obtaining 730, code for decoding 732, code for resolving 734, code for extracting 736, code for comparing 738, code for determining 740, code for detecting 742, code for aligning 744, code for identifying 746, code for utilizing 748, and code for retrieving 750. Processing of the code 730-750 may enable and cause the apparatus 700 to perform the method 500 described with respect to FIG. 5, or any aspect related to it.

[0090] The one or more processors 704 include circuitry configured to implement (e.g., execute) the code (e.g., executable instructions) stored in the computer-readable medium / memory 728, including circuitry for obtaining 706, circuitry for decoding 708, circuitry for resolving 710, circuitry for extracting 712, circuitry for comparing 714, circuitry for determining 716, circuitry for detecting 718, circuitry for aligning 720, circuitry for identifying 722, circuitry for utilizing 724, and circuitry for retrieving 726. Processing with circuitry 706-726 may enable and cause the apparatus 700 to perform the method 500 described with respect to FIG. 5, or any aspect related to it.

[0091] More generally, means for communicating, transmitting, sending or outputting for transmission may include the transceiver 754 and / or antenna 756 of the apparatus 700 in FIG. 7, and / or one or more processors 704 of the apparatus 700 in FIG. 7. Means for communicating, receiving or obtaining may include the transceiver 754 and / or antenna 756 of the apparatus 700 in FIG. 7, and / or one or more processors 704 of the apparatus 700 in FIG. 7.

[0092] Example Clauses

[0093] Implementation examples are described in the following numbered clauses:

[0094] Clause 1: A method for face-forgery detection by an apparatus comprising: obtaining an image comprising a face; extracting, with a face feature extraction process, a face feature from the image; generating a watermark based on the face feature; encoding the watermark into the image to generate a watermarked image; and storing the watermarked image in the one or more memories.

[0095] Clause 2: The method of Clause 1, wherein extracting, with the face feature extraction process, the face feature from the image, comprises: detecting the face in the image; aligning the detected face based on a face shape determined based on a location of the detected face in the image and a size of the detected face in the image; and identifying the face feature to extract from the image from the detected and aligned face.

[0096] Clause 3: The method of any one of Clauses 1-2, wherein the watermark comprises a deterministic value based on the face feature.

[0097] Clause 4: The method of any one of Clauses 1-3, wherein the watermark comprises a floating-point value based on the face feature.

[0098] Clause 5: The method of any one of Clauses 1-4, wherein generating the watermark based on the face feature comprises utilizing an invertible mapping model.

[0099] Clause 6: The method of Clause 5, wherein the invertible mapping model is configured to transform the face feature into the watermark and scale the watermark to correspond to a pixel size of the image.

[0100] Clause 7: The method of any one of Clauses 1-6, further comprising transmitting the watermarked image.

[0101] Clause 8: A method for face-forgery detection by an apparatus comprising: obtaining a watermarked image comprising a face; decoding the watermarked image to obtain a watermark, wherein the watermark is based on a first face feature corresponding to the face in the watermarked image; resolving the first face feature from the watermark; extracting, with a face feature extraction process, a second face feature from the watermarked image; comparing the first face feature to the second face feature to determine whether the first face feature and the second face feature match within a degree of similarity; and one of determining that the watermarked image is a real image when the match of the first face feature and the second face feature is within the degree of similarity; or determining that the watermarked image is an altered image when the match of the first face feature and the second face feature is less than the degree of similarity.

[0102] Clause 9: The method of Clause 8, wherein extracting, with the face feature extraction process, the second face feature from the watermarked image comprises: detecting the face in the watermarked image; aligning the detected face based on a face shape determined based on a location of the detected face in the watermarked image and a size of the detected face in the watermarked image; and identifying the first face feature to extract from the watermarked image from the detected and aligned face.

[0103] Clause 10: The method of any one of Clauses 8-9, wherein the watermark comprises a deterministic value based on the first face feature.

[0104] Clause 11: The method of any one of Clauses 8-10, wherein the watermark comprises a floating-point value based on the first face feature.

[0105] Clause 12: The method of any one of Clauses 8-11, wherein resolving the first face feature from the watermark comprises utilizing an invertible mapping model.

[0106] Clause 13: The method of Clause 12, wherein the invertible mapping model is configured to transform the watermark into the first face feature and scale the first face feature to correspond to a pixel size of the face in the watermarked image.

[0107] Clause 14: The method of any one of Clauses 8-13, wherein comparing the first face feature to the second face feature to determine whether the first face feature and the second face feature match within the degree of similarity comprises determining cosine similarity between the first face feature and the second face feature.

[0108] Clause 15: The method of any one of Clauses 8-14, wherein obtaining the watermarked image comprises retrieving the watermarked image from a data storage.

[0109] Clause 16: One or more apparatuses, comprising: one or more memories comprising executable instructions; and one or more processors configured to execute the executable instructions and cause the one or more apparatuses to perform a method in accordance with any one of Clauses 1-15.

[0110] Clause 17: One or more apparatuses, comprising: one or more memories; and one or more processors, coupled to the one or more memories, configured to cause the one or more apparatuses to perform a method in accordance with any one of Clauses 1-15.

[0111] Clause 18: One or more apparatuses, comprising: one or more memories; and one or more processors, coupled to the one or more memories, configured to perform a method in accordance with any one of Clauses 1-15.

[0112] Clause 19: One or more apparatuses, comprising means for performing a method in accordance with any one of Clauses 1-15.

[0113] Clause 20: One or more non-transitory computer-readable media comprising executable instructions that, when executed by one or more processors of one or more apparatuses, cause the one or more apparatuses to perform a method in accordance with any one of Clauses 1-15.

[0114] Clause 21: One or more computer program products embodied on one or more computer-readable storage media comprising code for performing a method in accordance with any one of Clauses 1-15.

[0115] Additional Considerations

[0116] The preceding description is provided to enable any person skilled in the art to practice the various aspects described herein. The examples discussed herein are not limiting of the scope, applicability, or aspects set forth in the claims. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. For example, changes may be made in the function and arrangement of elements discussed without departing from the scope of the disclosure. Various examples may omit, substitute, or add various procedures or components as appropriate. For instance, the methods described may be performed in an order different from that described, and various actions may be added, omitted, or combined. Also, features described with respect to some examples may be combined in some other examples. For example, an apparatus may be implemented or a method may be practiced using any number of the aspects set forth herein. In addition, the scope of the disclosure is intended to cover such an apparatus or method that is practiced using other structure, functionality, or structure and functionality in addition to, or other than, the various aspects of the disclosure set forth herein. It should be understood that any aspect of the disclosure disclosed herein may be embodied by one or more elements of a claim.

[0117] The various illustrative logical blocks, modules and circuits described in connection with the present disclosure may be implemented or performed with a general purpose processor, an AI processor, a digital signal processor (DSP) , an ASIC, a field programmable gate array (FPGA) or other programmable logic device (PLD) , discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any commercially available processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, a system on a chip (SoC) , or any other such configuration.

[0118] As used herein, a phrase referring to “at least one of” a list of items refers to any combination of those items, including single members. As an example, “at least one of: a, b, or c” is intended to cover a, b, c, a-b, a-c, b-c, and a-b-c, as well as any  combination with multiples of the same element (e.g., a-a, a-a-a, a-a-b, a-a-c, a-b-b, a-c-c, b-b, b-b-b, b-b-c, c-c, and c-c-c or any other ordering of a, b, and c) .

[0119] As used herein, the term “determining” encompasses a wide variety of actions. For example, “determining” may include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, a database or another data structure) , ascertaining and the like. Also, “determining” may include receiving (e.g., receiving information) , accessing (e.g., accessing data in a memory) and the like. Also, “determining” may include resolving, selecting, choosing, establishing and the like.

[0120] As used herein, “coupled to” and “coupled with” generally encompass direct coupling and indirect coupling (e.g., including intermediary coupled aspects) unless stated otherwise. For example, stating that a processor is coupled to a memory allows for a direct coupling or a coupling via an intermediary aspect, such as a bus.

[0121] The methods disclosed herein comprise one or more actions for achieving the methods. The method actions may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of actions is specified, the order and / or use of specific actions may be modified without departing from the scope of the claims. Further, the various operations of methods described above may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software component (s) and / or module (s) , including, but not limited to a circuit, an application specific integrated circuit (ASIC) , or processor.

[0122] The following claims are not intended to be limited to the aspects shown herein, but are to be accorded the full scope consistent with the language of the claims. Reference to an element in the singular is not intended to mean only one unless specifically so stated, but rather “one or more. ” The subsequent use of a definite article (e.g., “the” or “said” ) with an element (e.g., “the processor” ) is not intended to invoke a singular meaning (e.g., “only one” ) on the element unless otherwise specifically stated. For example, reference to an element (e.g., “aprocessor, ” “acontroller, ” “amemory, ” “atransceiver, ” “an antenna, ” “the processor, ” “the controller, ” “the memory, ” “the transceiver, ” “the antenna, ” etc. ) , unless otherwise specifically stated, should be understood to refer to one or more elements (e.g., “one or more processors, ” “one or more controllers, ” “one or more memories, ” “one more transceivers, ” etc. ) . The terms “set”  and “group” are intended to include one or more elements, and may be used interchangeably with “one or more. ” Where reference is made to one or more elements performing functions (e.g., steps of a method) , one element may perform all functions, or more than one element may collectively perform the functions. When more than one element collectively performs the functions, each function need not be performed by each of those elements (e.g., different functions may be performed by different elements) and / or each function need not be performed in whole by only one element (e.g., different elements may perform different sub-functions of a function) . Similarly, where reference is made to one or more elements configured to cause another element (e.g., an apparatus) to perform functions, one element may be configured to cause the other element to perform all functions, or more than one element may collectively be configured to cause the other element to perform the functions. Unless specifically stated otherwise, the term “some” refers to one or more. All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims.

Claims

1.An apparatus, comprising: one or more memories; and one or more processors, coupled to the one or more memories, and configured to cause the apparatus to:obtain an image comprising a face;extract, with a face feature extraction process, a face feature from the image;generate a watermark based on the face feature;encode the watermark into the image to generate a watermarked image; andstore the watermarked image in the one or more memories.2.The apparatus of claim 1, wherein to extract, with the face feature extraction process, the face feature from the image, the one or more processors are configured to cause the apparatus to:detect the face in the image;align the detected face based on a face shape determined based on a location of the detected face in the image and a size of the detected face in the image; andidentify the face feature to extract from the image from the detected and aligned face.3.The apparatus of claim 1, wherein the watermark comprises a deterministic value based on the face feature.4.The apparatus of claim 1, wherein the watermark comprises a floating-point value based on the face feature.5.The apparatus of claim 1, wherein to generate the watermark based on the face feature, the one or more processors are configured to cause the apparatus to utilize an invertible mapping model.6.The apparatus of claim 5, wherein the invertible mapping model is configured to transform the face feature into the watermark and scale the watermark to correspond to a pixel size of the image.7.The apparatus of claim 1, wherein the one or more processors are configured to further cause the apparatus to transmit the watermarked image.8.An apparatus, comprising: one or more memories; and one or more processors, coupled to the one or more memories, and configured to cause the apparatus to:obtain a watermarked image comprising a face;decode the watermarked image to obtain a watermark, wherein the watermark is based on a first face feature corresponding to the face in the watermarked image;resolve the first face feature from the watermark;extract, with a face feature extraction process, a second face feature from the watermarked image;compare the first face feature to the second face feature to determine whether the first face feature and the second face feature match within a degree of similarity; andone of:determine that the watermarked image is a real image when the match of the first face feature and the second face feature is within the degree of similarity; ordetermine that the watermarked image is an altered image when the match of the first face feature and the second face feature is less than the degree of similarity.9.The apparatus of claim 8, wherein to extract, with the face feature extraction process, the second face feature from the watermarked image, the one or more processors are configured to cause the apparatus to:detect the face in the watermarked image;align the detected face based on a face shape determined based on a location of the detected face in the watermarked image and a size of the detected face in the watermarked image; andidentify the first face feature to extract from the watermarked image from the detected and aligned face.10.The apparatus of claim 8, wherein the watermark comprises a deterministic value based on the first face feature.11.The apparatus of claim 8, wherein the watermark comprises a floating-point value based on the first face feature.12.The apparatus of claim 8, wherein to resolve the first face feature from the watermark, the one or more processors are configured to cause the apparatus to utilize an invertible mapping model.13.The apparatus of claim 12, wherein the invertible mapping model is configured to transform the watermark into the first face feature and scale the first face feature to correspond to a pixel size of the face in the watermarked image.14.The apparatus of claim 8, wherein to compare the first face feature to the second face feature to determine whether the first face feature and the second face feature match within the degree of similarity, the one or more processors are configured to cause the apparatus to determine cosine similarity between the first face feature and the second face feature.15.The apparatus of claim 8, wherein to obtain the watermarked image, the one or more processors are configured to cause the apparatus to retrieve the watermarked image from a data storage.16.A method for face-forgery detection by an apparatus comprising:obtaining a watermarked image comprising a face;decoding the watermarked image to obtain a watermark, wherein the watermark is based on a first face feature corresponding to the face in the watermarked image;resolving the first face feature from the watermark;extracting, with a face feature extraction process, a second face feature from the watermarked image;comparing the first face feature to the second face feature to determine whether the first face feature and the second face feature match within a degree of similarity; andone of:determining that the watermarked image is a real image when the match of the first face feature and the second face feature is within the degree of similarity; ordetermining that the watermarked image is an altered image when the match of the first face feature and the second face feature is less than the degree of similarity.17.The method of claim 16, wherein extracting, with the face feature extraction process, the second face feature from the watermarked image further comprises:detecting the face in the watermarked image;aligning the detected face based on a face shape determined based on a location of the detected face in the watermarked image and a size of the detected face in the watermarked image; andidentifying the first face feature to extract from the watermarked image from the detected and aligned face.18.The method of claim 16, wherein the step of resolving the first face feature from the watermark comprises utilizing an invertible mapping model.19.The method of claim 18, wherein the invertible mapping model is configured to transform the watermark into the first face feature and scale the first face feature to correspond to a pixel size of the face in the watermarked image.20.The method of claim 16, wherein the step of comparing the first face feature to the second face feature to determine whether the first face feature and the second face feature match within the degree of similarity comprises determining a cosine similarity between the first face feature and the second face feature.

Citation Information

Patent Citations

  • Deep forgery active detection method based on face features and associated watermarks thereof

    CN117975578A

  • Digital watermarking of picture identity documents using facial features

    US20070237354A1

  • Methods and devices for enrollment and verification of biometric information in identification documents

    US20100052852A1

  • Method for embedding and extracting multi-scale space based watermark

    US20160012564A1