Agent-based methods, devices, equipment, storage media, and products for detecting fake faces.
By generating fake operation commands and training a fake detection model using intelligent agents and multimodal samples, the problem of poor training performance of fake face detection models is solved, and the model's recognition ability and accuracy are improved.
Patent Information
- Application Number
- CN202510956277.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-07-11
AI Technical Summary
Existing fake face detection models suffer from poor training performance and difficulty in accurately identifying fake faces due to the lack of high-quality fake face datasets.
By generating forgery operation instructions based on the agent's profile and memory information, the original image is processed to generate candidate forgery images and ordinary text, constructing multimodal samples, and training a forgery detection model based on these samples.
It improves the generalization ability and accuracy of the forgery detection model, enhances the ability to identify new types of forged faces, and solves the problem of insufficient high-quality forged face datasets.
Smart Images

Figure CN120452075B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of face detection technology, and in particular to methods, devices, equipment, storage media and products for detecting fake faces based on intelligent agents. Background Technology
[0002] In the current field of artificial intelligence, fake face detection technology is crucial for maintaining cybersecurity and social stability. However, due to the lack of large-scale, high-quality fake face datasets, the training performance of existing fake face detection models is often unsatisfactory. This affects the accuracy of the fake face detection models, making it difficult to identify fake faces.
[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of this application is to provide a method, apparatus, device, storage medium, and product for detecting fake faces based on intelligent agents, aiming to solve the technical problem that the lack of high-quality fake face datasets affects the training effect of fake face detection models, making it difficult to identify fake faces.
[0005] To achieve the above objectives, this application proposes a method for detecting fake faces based on intelligent agents, the method comprising:
[0006] Based on the agent's profile and memory information, forged operation instructions are generated;
[0007] The original image is processed according to the forgery operation instructions to obtain a candidate forgery image, and the plain text of the candidate forgery image is generated.
[0008] Generate multimodal samples based on the candidate forged images and the ordinary text;
[0009] Based on the multimodal samples, the forgery detection model is trained to obtain the target forgery detection model, and forged faces are detected according to the target forgery detection model.
[0010] In one embodiment, the step of generating forged operation instructions based on the agent's profile information and memory information includes:
[0011] Determine inherent forgery preferences based on the portrait information;
[0012] Based on the memory information, the inherent forgery preferences are adjusted according to the forgery tree to generate forgery operation instructions, wherein the forgery tree is used to simulate the decision-making process of an intelligent agent when performing face forgery.
[0013] In one embodiment, the step of generating multimodal samples based on the candidate forged image and the ordinary text includes:
[0014] The candidate forged image is input into an external face forgery detector to obtain a first evaluation score;
[0015] The candidate forged image is input into the large language model to obtain a second evaluation score;
[0016] Based on the first evaluation score and the second evaluation score, the target evaluation score is obtained;
[0017] Determine whether the target evaluation score is greater than or equal to a preset threshold;
[0018] When the target evaluation score is greater than or equal to the preset threshold, multimodal samples are generated based on the candidate forged image and the ordinary text.
[0019] In one embodiment, after the step of determining whether the target evaluation score is greater than or equal to a preset threshold, the method further includes:
[0020] When the target evaluation score is greater than or equal to the preset threshold, the forgery attempt result of the candidate forgery image is determined as a successful forgery attempt;
[0021] When the target evaluation score is less than the preset threshold, the forgery attempt result of the candidate forgery image is determined as a forgery attempt failure;
[0022] The first evaluation score, the second evaluation score, and the forgery attempt result of the candidate forged image are used as emotional memory information in the target attempt information, and the forgery operation type and forgery operation area of the candidate forged image are used as factual memory information in the target attempt information.
[0023] The target attempt information is analyzed using a large language model to obtain common information about success and failure, and the memory information is updated based on this common information.
[0024] In one embodiment, the step of analyzing the target attempt information using a large language model to obtain common success and failure information, and updating the memory information based on the common success and failure information, includes:
[0025] Obtain a set of target attempt information within a short-term reflection period, wherein the set of target attempt information includes multiple target attempt information;
[0026] Obtain a long-term attempt information set, wherein the long-term attempt information set includes all target attempt information;
[0027] The target attempt information set and the long-term attempt information set are analyzed by the large language model to obtain common information on success and failure, and the memory information is updated according to the common information on success and failure.
[0028] In one embodiment, the step of generating multimodal samples based on the candidate forged image and the ordinary text includes:
[0029] The candidate forged image is published in a simulated social environment, where multiple agents with different roles interact with the candidate forged image to generate social text, and a text description is determined based on the social text and the ordinary text.
[0030] Determine the forged label based on the original image and the candidate forged image;
[0031] Multimodal samples are generated based on the original image, the candidate forged image, the text description, and the forged label.
[0032] Furthermore, to achieve the above objectives, this application also proposes an agent-based fake face detection device, which includes:
[0033] The generation module is used to generate fake operation instructions based on the intelligent agent's profile information and memory information;
[0034] The processing module is used to process the original image according to the forgery operation instruction to obtain a candidate forgery image and generate plain text of the candidate forgery image;
[0035] The generation module is further configured to generate multimodal samples based on the candidate forged image and the ordinary text;
[0036] The training module is used to train the forgery detection model based on the multimodal samples to obtain the target forgery detection model, and to detect forged faces based on the target forgery detection model.
[0037] Furthermore, to achieve the above objectives, this application also proposes an agent-based fake face detection device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the agent-based fake face detection method described above.
[0038] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the agent-based fake face detection method described above.
[0039] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the agent-based fake face detection method described above.
[0040] One or more technical solutions proposed in this application have at least the following technical effects:
[0041] This application proposes a method, apparatus, device, storage medium, and product for detecting forged faces based on intelligent agents. It generates forgery operation instructions based on the agent's profile and memory information; processes the original image according to the forgery operation instructions to obtain candidate forged images and generates plain text for the candidate forged images; generates multimodal samples based on the candidate forged images and the plain text; trains a forgery detection model based on the multimodal samples to obtain a target forgery detection model; and detects forged faces based on the target forgery detection model. This solves the technical problem that the lack of high-quality forged face datasets affects the training effect of forgery detection models, making it difficult to identify forged faces. Compared with existing technologies, this application simulates intelligent agents with different profiles and memories to perform face forgery, thereby generating more realistic and diverse high-quality multimodal samples. This method effectively solves the problem of insufficient high-quality forged face datasets, enhances the generalization ability and accuracy of the forgery detection model, and thus improves the model's ability to recognize new types of forged attack faces. Attached Figure Description
[0042] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 This is a flowchart illustrating an embodiment of the agent-based fake face detection method of this application.
[0045] Figure 2 This is a schematic diagram of the forgery tree structure provided in Embodiment 1 of the agent-based forgery face detection method of this application;
[0046] Figure 3 This is a flowchart illustrating Embodiment 2 of the agent-based fake face detection method of this application;
[0047] Figure 4 This is a schematic diagram of image quality assessment provided in Embodiment 2 of the agent-based fake face detection method of this application;
[0048] Figure 5 This is a schematic diagram of data construction provided in Embodiment 2 of the agent-based fake face detection method of this application;
[0049] Figure 6 This is a schematic diagram of the module structure of the agent-based fake face detection device according to an embodiment of this application;
[0050] Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the agent-based fake face detection method in this application embodiment.
[0051] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0052] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0053] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0054] The main solution of this application embodiment is as follows: generating forgery operation instructions based on the portrait information and memory information of the intelligent agent; processing the original image according to the forgery operation instructions to obtain candidate forgery images and generating ordinary text of the candidate forgery images; generating multimodal samples based on the candidate forgery images and the ordinary text; training the forgery detection model based on the multimodal samples to obtain a target forgery detection model, and detecting forged faces according to the target forgery detection model.
[0055] As can be seen from the above embodiments, this application generates forgery operation instructions based on the agent's portrait information and memory information; processes the original image according to the forgery operation instructions to obtain candidate forgery images, and generates ordinary text of the candidate forgery images; generates multimodal samples based on the candidate forgery images and the ordinary text; trains a forgery detection model based on the multimodal samples to obtain a target forgery detection model, and detects forged faces based on the target forgery detection model. This solves the technical problem that the lack of high-quality forgery face datasets affects the training effect of the forgery detection model, making it difficult to identify forged faces. Compared with existing technologies, this application simulates agents with different portraits and memories to perform face forgery, thereby generating more realistic and diverse high-quality multimodal samples. This method effectively solves the problem of insufficient high-quality forgery face datasets, enhances the generalization ability and accuracy of the forgery detection model, and thus improves the model's ability to recognize new types of forged attack faces.
[0056] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, such as an agent-based fake face detection device. The following description uses an agent-based fake face detection device as an example to illustrate this embodiment and the subsequent embodiments.
[0057] Based on this, embodiments of this application provide a method for detecting fake faces based on intelligent agents, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the agent-based fake face detection method of this application.
[0058] In this embodiment, the agent-based method for detecting fake faces includes steps S10 to S40:
[0059] Step S10: Based on the agent's profile information and memory information, generate fake operation instructions;
[0060] It should be noted that before constructing multimodal samples, the Agent4FaceForgery framework needs to be initialized. This includes defining a set of available forgery tools (e.g., Stable Diffusion, SBI, etc.), defining a set of participant roles in a social interaction environment (e.g., a student named Jack, aged 23), and defining and initializing multiple LLM-driven agents, assigning each agent an initial profile representing a specific social role. Specifically: the Watcher agent is configured to tend to browse content and has weak judgment; the Explorer agent is configured to tend to identify inconsistencies by comparing relevant information; the Critic agent is configured to proactively identify and point out forgery traces; and the Gemini Auditor agent is specifically configured to issue deceptive statements contrary to the facts regarding forged content, its role being to proactively generate highly adversarial text-image inconsistency samples.
[0061] It should be noted that each agent's profile information can consist of a forgery preference and a priority attribute. The pseudo-preference vector defines the agent's preference for different forgery types (e.g., local adjustments, full-face adjustments, or background replacement), while the priority attribute defines the specific facial regions the agent focuses on when performing forgery (e.g., left eye, right eye, nose, or mouth). Memory information refers to the agent's historical experience, including factual memory information of past forgery attempts (e.g., edits already performed) and emotional memory information. Finally, the agent's profile information and memory information can be fused to obtain a final forgery operation instruction (i.e., specifying the specific visual editing operation to be performed on the original image). For example, a replacement operation can be performed on the eye region of the original image.
[0062] In one feasible implementation, the step of generating forgery operation instructions based on the agent's profile information and memory information includes: determining inherent forgery preferences based on the profile information; adjusting the inherent forgery preferences based on the memory information and a forgery tree to generate forgery operation instructions, wherein the forgery tree is used to simulate the decision-making process of the agent when performing face forgery.
[0063] In practical implementation, the agent's inherent forgery preferences can be determined based on its profile information (such as the agent's preference for different forgery types and the specific facial regions the agent focuses on when performing forgery). This might involve preferences for specific facial features, such as face shape or glasses type. Then, based on the agent's memory information, which may include the results of past forgery attempts, these preferences are adjusted using a forgery tree. The structure of the forgery tree is as follows: Figure 2 As shown, the root node is divided into first-level branches based on local and whole-face features, and second-level nodes correspond to manipulable facial organs (such as the glasses node which includes nodes for frame type and reflectivity). For example, starting from the root node, it is divided into two main branches: "local" and "whole-face," further subdivided into specific features such as "face," "glasses," "nose," and "mouth." Each node represents a decision point. The agent selects different paths in the tree, such as choosing "ellipse" or "circle" under the "face" node, or "left" or "right" under the "glasses" node, to generate specific forgery operation instructions. This tree-structured decision-making process allows the agent to systematically explore different forgery strategies and optimize its choices based on memory information, thereby generating more accurate and effective forgery operation instructions. This method not only improves the quality and success rate of forged images but also, by simulating the agent's decision-making process, makes the generated forgery operation instructions more consistent with the agent's experience and preferences.
[0064] In practical implementation, the agent's propensity to forge specific facial regions (such as glasses and mouth shape) can be analyzed based on its profile information to form an initial preference weight vector (i.e., the attention given to each facial feature). Then, a forgery tree is constructed, with the root node divided into first-level branches based on "local-whole face," and second-level nodes corresponding to manipulable facial organs (e.g., the glasses node includes sub-nodes for frame type and reflectivity). Each decision node is associated with a success rate matrix in the memory bank, where the matrix elements S... ij This represents the historical success rate of the agent performing forgery operations using the j-th parameter combination under the i-th type of lighting conditions. Path scores need to be calculated in real-time when traversing the forgery tree.
[0065]
[0066] In the formula, Wp represents the agent's inherent preference weight vector (e.g., the preference for modifying glasses is 0.7, and the preference for modifying mouth shape is 0.3), and X represents the feature vector of the original image (e.g., 512-dimensional features extracted by CNN). This indicates the degree of matching between the feature vector of the original image and the inherent preferences of the agent (if the glasses region features are significant in X and the glasses weight is high in Wp, then this item scores highly).
[0067] It should be noted that when the path scores of multiple child nodes are similar (e.g., |Pscore1−Pscore2|<0.05), it is necessary to trace back to the nearest parent node and check whether its decision conditions are outdated (e.g., the associated memory item has not been updated for more than 6 months). If so, the parent node's path score needs to be recalculated to select a new child node branch. Finally, the final path is converted into an executable operation instruction set (i.e., fake operation instructions), including: region positioning coordinates (e.g., the glasses region BBox); feature modification parameters (e.g., lens reflectivity ±15%); and anti-detection injection points (e.g., adding specific frequency noise to the edge of the glasses).
[0068] Step S20: Process the original image according to the forgery operation instruction to obtain a candidate forgery image, and generate plain text of the candidate forgery image;
[0069] It should be noted that the original image is randomly selected from the queue of real images to be processed; the queue of real images to be processed is a queue of all images selected from the initial real face image library (a portion of which is selected from public datasets such as FF++ as the initial real face image library). According to the forgery operation instructions, specific modifications or synthesis operations are performed on the original image. These operations may include facial feature replacement, expression modification, lighting adjustment, background replacement, etc., to simulate a forgery scene in the real world, thereby generating candidate forgery images.
[0070] In its implementation, the agent receives forgery instructions and executes one or more forgery operations (using integrated forgery tools for face replacement, facial generation, expression modification, etc.) to generate a candidate forged image. Simultaneously, during visual editing, the agent's action module initiates a memory retrieval to retrieve historical experience related to the currently executed forgery operation. If the retrieved emotional memory indicates that past use of certain fusion parameters or techniques (e.g., a specific edge feathering radius) resulted in significant visual artifacts (e.g., overly blurry evaluations of edge areas), the agent's action module will adjust the parameters or forgery synthesis method in this operation (e.g., changing image lighting conditions or replacing the image background) to avoid repeating past errors. This adjustment aims to dynamically optimize the execution details of the forgery operation based on historical experience, thereby improving the realism of the final generated candidate forged image. Specifically, the parameters of the forgery tools (such as feathering radius, illumination intensity, and GAN mixing coefficients) are constructed into a high-dimensional parameter space (that is, the discrete and isolated parameters of the forgery tools are transformed into continuous mathematical representations, thereby mapping the tool parameters and results of past operations to this high-dimensional space to form a data point distribution). A lightweight prediction model (such as a Bayesian optimized surrogate model) is trained using the forgery parameter tools recorded in the memory module. When the characteristics of the current operation object (such as low-resolution faces and complex backgrounds) are detected to be similar to the forgery failure cases in memory, the model is invoked to generate parameter adjustment suggestions in real time.
[0071] In practical implementation, plain text can be generated based on the profile information to represent the forged image. The generated plain text has two possibilities: 1. If the agent's profile information tends towards objective statements, a more accurate description of the facial editing performed on the original image in the steps will be generated. Example: If the visual editing performed was replacing the eyes, the generated plain text would be: "The eyes of the person in the image have been replaced." 2. If the agent's profile information tends towards deceptive behavior, a misleading statement that contradicts the true state of the image will be generated. Example: Even if the original image has undergone significant forgery processing, the generated plain text might be: "This is an original image that has not been modified in any way."
[0072] Step S30: Generate multimodal samples based on the candidate forged image and the ordinary text;
[0073] In one feasible implementation, the step of generating multimodal samples based on the candidate forged image and the ordinary text includes: publishing the candidate forged image to a simulated social environment, having multiple agents with different roles interact with the candidate forged image to generate social text, and determining a text description based on the social text and the ordinary text; determining a forgery label based on the original image and the candidate forged image; and generating multimodal samples based on the original image, the candidate forged image, the text description, and the forgery label.
[0074] In the implementation, candidate forged images selected for training the forgery detection model are published in a simulated social environment. Multiple simulated user agents, driven by an LLM (Local Management Model) with different roles (e.g., Watcher, Critic, Explorer, GeminiAuditor), interact with these candidate images, generating comments, judgments, and sharing, thus producing a series of social texts. Finally, the forgery label (y=1), the social text, and the deceptive statement of the "Gemini Auditor" are combined to construct multimodal samples with complex image-text consistency challenges. For example, an image pointed out as forgery by the Critic, even if its original text claims to be genuine, can still be labeled as requiring the model to identify this inconsistency.
[0075] It should be noted that the interaction between the AI agents with different roles and the candidate forged images is as follows: Specifically, an AI agent assigned the profile of a critic tends to generate cautious, questioning social text, potentially pointing out specific visual inconsistencies. For example: "Wait a minute, the lighting directions for the left and right eyes of the person in this image seem inconsistent, making it look like it was composited later," or "The edges where the neck and face connect are a bit blurry; something doesn't seem right." An AI agent assigned the profile of an explorer may generate social text reflecting judgments made through association or information comparison. For example: "I think I've seen this background elsewhere, but the person in it isn't him. Is this image photoshopped?" An AI agent assigned the profile of a gemini auditor is specifically designed to generate highly deceptive, leading statements that contradict the facts.
[0076] Step S40: Based on the multimodal samples, train the forgery detection model to obtain the target forgery detection model, and detect forged faces according to the target forgery detection model.
[0077] It's important to note that multimodal samples combine image and text information, providing the model with more comprehensive forgery features, enabling it to learn and recognize forged images from different dimensions. During training, the model analyzes these samples and continuously optimizes its parameters to reduce prediction errors, thereby improving its ability to recognize forged images. After training, the resulting forgery detection model can accurately classify and recognize new forged face images. The advantage of this method lies in its ability to improve detection accuracy, enhance the model's adaptability to different forgery techniques, and improve the model's generalization ability.
[0078] This embodiment generates forgery operation instructions based on the agent's profile and memory information; processes the original image according to the forgery operation instructions to obtain candidate forgery images, and generates plain text for the candidate forgery images; generates multimodal samples based on the candidate forgery images and the plain text; trains a forgery detection model based on the multimodal samples to obtain a target forgery detection model, and detects forged faces based on the target forgery detection model. This solves the technical problem that the lack of high-quality forged face datasets affects the training effect of the forgery detection model, making it difficult to identify forged faces. Compared with existing technologies, this application simulates agents with different profiles and memories to perform face forgery, thereby generating more realistic and diverse high-quality multimodal samples. This method effectively solves the problem of insufficient high-quality forged face datasets, enhances the generalization ability and accuracy of the forgery detection model, and thus improves the model's ability to recognize new types of forged attack faces.
[0079] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 3 Step S30 also includes steps S301 to S305:
[0080] Step S301: Input the candidate forged image into an external face forgery detector to obtain a first evaluation score;
[0081] It should be noted that the first evaluation score output by the external face forgery detector Xception is the probability dis of outputting the candidate forgery image being forged.
[0082] Step S302: Input the candidate forged image into the large language model to obtain a second evaluation score;
[0083] It should be noted that the second evaluation score output by the large language model is an image quality score for the candidate forgery model.
[0084] Step S303: Based on the first evaluation score and the second evaluation score, obtain the target evaluation score;
[0085] It should be noted that the target evaluation score of the candidate forged image is obtained by weighted fusion of two evaluation scores. In the first part, the first evaluation score is the objective deception difficulty score, which requires the candidate forged image to be input into a pre-set external face forgery detector Xception, and the first evaluation score is output (i.e., the probability dis of the candidate forged image being forged). In the second part, the second evaluation score is the subjective forgery quality score LLM, which requires the candidate forged image to be input into a large language model, and the second evaluation score is output.
[0086] In practice, the target evaluation score can be determined based on the first evaluation score and the second evaluation score. The specific calculation formula is as follows:
[0087]
[0088] In the formula, dis represents the first evaluation score, LLM represents the second evaluation score, and sn represents the target evaluation score.
[0089] Step S304: Determine whether the target evaluation score is greater than or equal to a preset threshold;
[0090] It should be noted that when the target evaluation score of the candidate forged image is greater than or equal to the preset threshold, the forgery attempt on the candidate forged image can be considered successful. In this case, the candidate forged image will be used as a multimodal sample to train the forgery detection model. When the target evaluation score of the candidate forged image is less than the preset threshold, the forgery attempt on the candidate forged image can be considered unsuccessful. In this case, the candidate forged image will not be used to train the forgery detection model.
[0091] In one feasible implementation, after the step of determining whether the target evaluation score is greater than or equal to a preset threshold, the method further includes: when the target evaluation score is greater than or equal to the preset threshold, determining the forgery attempt result of the candidate forgery image as a successful forgery attempt; when the target evaluation score is less than the preset threshold, determining the forgery attempt result of the candidate forgery image as a failed forgery attempt; using the first evaluation score, the second evaluation score, and the forgery attempt result of the candidate forgery image as emotional memory information in the target attempt information, and using the forgery operation type and forgery operation area of the candidate forgery image as factual memory information in the target attempt information; analyzing the target attempt information through a large language model to obtain common information of success and failure, and updating the memory information according to the common information of success and failure.
[0092] It should be noted that, as Figure 4As shown, when the target evaluation score is greater than or equal to the preset threshold, the forgery attempt of the candidate forged image is determined to be successful, which means that the forged image is visually realistic enough to deceive the forgery detector and the large language model; when the target evaluation score is less than the preset threshold, the forgery attempt of the candidate forged image is determined to be unsuccessful, which indicates that the forged image has obvious visual defects and is easily identified by the detector.
[0093] It is understandable that, regardless of whether the forgery attempt of the candidate image is successful or not, the result of this forgery attempt (including operations, images, text, evaluation scores, etc., i.e. target attempt information) will be memorized into the agent's memory model.
[0094] It should be noted that the target attempt information includes both emotional memory information and factual memory information. Emotional memory information includes the first evaluation score (objective deception difficulty score), the second evaluation score (subjective forgery quality score), and the forgery attempt result of the candidate forgery image. This information reflects the emotional and subjective evaluation of the forgery attempt, helping the agent understand the success or failure of the forgery attempt and the reasons for it. Factual memory information includes the forgery operation type and forgery operation area of the candidate forgery image. This information provides specific details about the forgery operation, such as which techniques were used and which parts of the image were modified.
[0095] Understandably, analyzing target attempt information using a large language model to extract common information about success and failure might include common operation types and regions in successful forgery attempts, as well as common defects in failed forgery attempts. By analyzing this common information, the agent can learn which operations are more likely to lead to successful forgery and which operations are more easily recognized by the detector, thereby optimizing its forgery strategy. Based on the extracted common information, the agent's memory information is updated. This update helps the agent make more informed decisions in future forgery attempts, improving the quality of forged images and generating more realistic and harder-to-detect forgeries.
[0096] In one feasible implementation, the step of analyzing the target attempt information using a large language model to obtain common success and failure information, and updating the memory information based on the common success and failure information, includes: obtaining a set of target attempt information within a short-term reflection period, wherein the set of target attempt information includes multiple target attempt information; obtaining a set of long-term attempt information, wherein the long-term attempt information set includes all target attempt information; analyzing the set of target attempt information and the set of long-term attempt information using the large language model to obtain common success and failure information, and updating the memory information based on the common success and failure information.
[0097] It should be noted that the target attempt information set within the short-term reflection period refers to the target attempt information collected in the most recent period (such as the most recent attempts) (i.e., short-term memory), which reflects the agent's performance and results in recent spoofing attempts. The long-term target attempt information set includes all target attempt information collected by the agent over a longer period (such as the entire training process) (i.e., long-term memory), which provides a more comprehensive perspective and helps the agent understand its performance and long-term trends at different time scales.
[0098] It should be noted that the agent periodically (e.g., after every N attempts) reflects on its memory. The LLM analyzes past successes and failures to adjust its subsequent forgery strategies. Specifically, this is performed internally by the agent through a collaborative process between its memory and action modules. The purpose is to enable the agent to learn from accumulated experience and dynamically adjust its future forgery strategies. This step is triggered under a pre-set periodic condition, for example, after the agent completes N forgery attempts. Once triggered, the agent's memory module initiates a reflection operation. In this operation, the memory module first retrieves the complete record of the most recent N forgery attempts from its stored factual and emotional memories. Subsequently, the core within the agent analyzes these records to identify recurring patterns and challenges. For example, the LLM core might analyze and summarize the commonalities in the failures of a certain forgery technique under specific conditions (such as processing large-angle profile images), thereby providing the agent with directions for improvement. Through this periodic reflection and analysis, the agent can learn from accumulated experience and dynamically adjust its future forgery strategies. This mechanism enables the agent to continuously adapt and optimize its behavior to improve success rate and efficiency. This self-learning and self-optimization capability is achieved through the collaborative work of its internal memory and action modules, which work together in the agent's decision-making process, allowing it to continuously improve and grow in complex and ever-changing environments.
[0099] In specific implementations, such as Figure 5 As shown, after identifying multiple candidate forged images through multi-round simulations using Adaptive Rejection Sampling (ARS), multiple agents with different roles interact with these candidate forged images to obtain the corresponding social text. Finally, a multimodal sample with complex image-text consistency challenges is constructed by combining the forged label (y=1), the social text, and the deceptive statement of the "Gemini Auditor".
[0100] Step S305: When the target evaluation score is greater than or equal to the preset threshold, generate multimodal samples based on the candidate forged image and the ordinary text.
[0101] Understandably, when the target evaluation score is greater than or equal to a preset threshold, the candidate forgery model needs to be processed to generate multimodal samples.
[0102] This embodiment obtains a first evaluation score by inputting the candidate forged image into an external face forgery detector; it then obtains a second evaluation score by inputting the candidate forged image into a large language model; based on the first and second evaluation scores, a target evaluation score is obtained; it is determined whether the target evaluation score is greater than or equal to a preset threshold; when the target evaluation score is greater than or equal to the preset threshold, multimodal samples are generated based on the candidate forged image and the ordinary text. By setting a threshold, the system can filter out high-quality forged images, which, along with their corresponding ordinary text descriptions, are used to generate multimodal samples. These samples not only enrich the diversity of training data but also improve the training quality and efficiency of the forgery detection model, thereby enhancing the model's ability to recognize forged images. Furthermore, this method promotes the agent's self-learning and optimization during the forgery process, enabling it to continuously adjust its forgery strategy based on the evaluation results to generate more realistic and harder-to-detect forged images, ultimately achieving a dual improvement in forgery image generation and detection technologies.
[0103] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the agent-based fake face detection method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0104] This application also provides a fake face detection device based on intelligent agents; please refer to [reference needed]. Figure 6 The agent-based fake face detection device includes:
[0105] The generation module 10 is used to generate fake operation instructions based on the profile information and memory information of the intelligent agent;
[0106] The processing module 20 is used to process the original image according to the forgery operation instruction to obtain a candidate forgery image and generate plain text of the candidate forgery image;
[0107] The generation module 10 is further configured to generate multimodal samples based on the candidate forged image and the ordinary text;
[0108] The training module 30 is used to train the forgery detection model based on the multimodal samples to obtain the target forgery detection model, and to detect forged faces based on the target forgery detection model.
[0109] The agent-based fake face detection device provided in this application employs the agent-based fake face detection method described in the above embodiments. It addresses the technical problem that the lack of high-quality fake face datasets affects the training effect of the fake face detection model, making it difficult to identify fake faces. Compared with the prior art, the beneficial effects of the agent-based fake face detection device provided in this application are the same as those of the agent-based fake face detection method provided in the above embodiments. Furthermore, other technical features of the agent-based fake face detection device are the same as those disclosed in the methods of the above embodiments, and will not be elaborated upon here.
[0110] This application provides an agent-based fake face detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the agent-based fake face detection method in the first embodiment described above.
[0111] The following is for reference. Figure 7 This document illustrates a structural diagram of an agent-based fake face detection device suitable for implementing embodiments of this application. The agent-based fake face detection device in this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 7 The agent-based fake face detection device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0112] like Figure 7As shown, the agent-based fake face detection device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the agent-based fake face detection device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the agent-based fake face detection device to communicate wirelessly or wiredly with other devices to exchange data. While the figure shows agent-based fake face detection devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0113] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0114] The agent-based fake face detection device provided in this application, employing the agent-based fake face detection method described in the above embodiments, can solve the technical problem that the lack of high-quality fake face datasets affects the training effect of the fake face detection model, making it difficult to identify fake faces. Compared with the prior art, the beneficial effects of the agent-based fake face detection device provided in this application are the same as those of the agent-based fake face detection method provided in the above embodiments, and other technical features in this agent-based fake face detection device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0115] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0116] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0117] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the agent-based fake face detection method in the above embodiments.
[0118] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0119] The aforementioned computer-readable storage medium may be included in an agent-based fake face detection device; or it may exist independently and not assembled into an agent-based fake face detection device.
[0120] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an agent-based fake face detection device, cause the agent-based fake face detection device to: generate a fake operation instruction based on the agent's portrait information and memory information; process the original image according to the fake operation instruction to obtain a candidate fake image and generate ordinary text of the candidate fake image; generate multimodal samples based on the candidate fake image and the ordinary text; train a fake detection model based on the multimodal samples to obtain a target fake detection model, and detect fake faces according to the target fake detection model.
[0121] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0123] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0124] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described agent-based fake face detection method. This solves the technical problem that the lack of high-quality fake face datasets affects the training effect of the fake face detection model, making it difficult to identify fake faces. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the agent-based fake face detection method provided in the above embodiments, and will not be repeated here.
[0125] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the agent-based fake face detection method described above.
[0126] The computer program product provided in this application can solve the technical problem that the lack of high-quality fake face datasets affects the training effect of fake face detection models, making it difficult to identify fake faces. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the agent-based fake face detection method provided in the above embodiments, and will not be repeated here.
[0127] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for detecting fake faces based on intelligent agents, characterized in that, The agent-based method for detecting fake faces includes: Based on the agent's profile and memory information, forged operation instructions are generated; The original image is processed according to the forgery operation instructions to obtain a candidate forgery image, and the plain text of the candidate forgery image is generated. Generate multimodal samples based on the candidate forged images and the ordinary text; Based on the multimodal samples, the forgery detection model is trained to obtain the target forgery detection model, and forged faces are detected according to the target forgery detection model. The step of generating forged operation instructions based on the profile information and memory information of the intelligent agent includes: Determine inherent forgery preferences based on the portrait information; Based on the memory information, the inherent forgery preferences are adjusted according to the forgery tree to generate forgery operation instructions. The forgery tree is used to simulate the decision-making process of an agent performing face forgery. The forgery tree is constructed with first-level branches at the root node, divided into local and whole-face branches. Second-level nodes correspond to the facial organs being operated on. Each decision node is associated with a success rate matrix in the memory bank, and the matrix element S of the success rate matrix... ij This represents the historical success rate of an agent performing forgery operations using the j-th parameter combination under the i-th type of lighting conditions; The step of generating multimodal samples based on the candidate forged image and the ordinary text includes: The candidate forged image is published in a simulated social environment, where multiple agents with different roles interact with the candidate forged image to generate social text, and a text description is determined based on the social text and the ordinary text. Determine the forged label based on the original image and the candidate forged image; Multimodal samples are generated based on the original image, the candidate forged image, the text description, and the forged label.
2. The agent-based fake face detection method as described in claim 1, characterized in that, The step of generating multimodal samples based on the candidate forged image and the ordinary text includes: The candidate forged image is input into an external face forgery detector to obtain a first evaluation score; The candidate forged image is input into the large language model to obtain a second evaluation score; Based on the first evaluation score and the second evaluation score, the target evaluation score is obtained; Determine whether the target evaluation score is greater than or equal to a preset threshold; When the target evaluation score is greater than or equal to the preset threshold, multimodal samples are generated based on the candidate forged image and the ordinary text.
3. The agent-based fake face detection method as described in claim 2, characterized in that, After the step of determining whether the target evaluation score is greater than or equal to a preset threshold, the method further includes: When the target evaluation score is greater than or equal to the preset threshold, the forgery attempt result of the candidate forgery image is determined as a successful forgery attempt; When the target evaluation score is less than the preset threshold, the forgery attempt result of the candidate forgery image is determined as a forgery attempt failure; The first evaluation score, the second evaluation score, and the forgery attempt result of the candidate forged image are used as emotional memory information in the target attempt information, and the forgery operation type and forgery operation area of the candidate forged image are used as factual memory information in the target attempt information. The target attempt information is analyzed using a large language model to obtain common information about success and failure, and the memory information is updated based on this common information.
4. The agent-based fake face detection method as described in claim 3, characterized in that, The step of analyzing the target attempt information using a large language model to obtain common success and failure information, and updating the memory information based on the common success and failure information, includes: Obtain a set of target attempt information within a short-term reflection period, wherein the set of target attempt information includes multiple target attempt information; Obtain a long-term attempt information set, wherein the long-term attempt information set includes all target attempt information; The target attempt information set and the long-term attempt information set are analyzed by the large language model to obtain common information on success and failure, and the memory information is updated according to the common information on success and failure.
5. A fake face detection device based on intelligent agents, characterized in that, The agent-based fake face detection device includes: The generation module is used to generate fake operation instructions based on the intelligent agent's profile information and memory information; The processing module is used to process the original image according to the forgery operation instruction to obtain a candidate forgery image and generate plain text of the candidate forgery image; The generation module is further configured to generate multimodal samples based on the candidate forged image and the ordinary text; The training module is used to train the forgery detection model based on the multimodal samples to obtain the target forgery detection model, and to detect forged faces based on the target forgery detection model. The generation module is further configured to: determine inherent forgery preferences based on the portrait information; adjust the inherent forgery preferences according to the forgery tree based on the memory information to generate forgery operation instructions, wherein the forgery tree is used to simulate the decision-making process of an intelligent agent when performing face forgery; construct the forgery tree, with the root node divided into first-level branches according to local-whole face, and second-level nodes corresponding to the operation of facial organs, each decision node being associated with a success rate matrix in the memory bank, and the matrix elements S of the success rate matrix being... ij This represents the historical success rate of an agent performing forgery operations using the j-th parameter combination under the i-th type of lighting conditions; The generation module is further configured to: publish the candidate forged image to a simulated social environment, where multiple agents with different roles interact with the candidate forged image to generate social text, and determine a text description based on the social text and the ordinary text; determine a forgery label based on the original image and the candidate forged image; and generate a multimodal sample based on the original image, the candidate forged image, the text description, and the forgery label.
6. A fake face detection device based on intelligent agents, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the agent-based fake face detection method as described in any one of claims 1 to 4.
7. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the agent-based fake face detection method as described in any one of claims 1 to 4.
8. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the agent-based fake face detection method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Cross-modal fusion-guided CLIP domain generalization face anti-counterfeiting method
CN119339447A
Cross-domain face anti-counterfeiting detection method and device based on multi-modal text enhancement
CN119441939A