Intelligent agent-based counterfeited face detection method and device, equipment, storage medium and product
Through the agent-based forged face detection method, high-quality multimodal samples are generated and the model is trained, which solves the problem of poor training effect of forged face detection model, and improves the generalization ability and recognition accuracy of the model.
Patent Information
- Application Number
- CN202510956277.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-07-11
AI Technical Summary
The existing forged face detection models lack high-quality forged face data sets, resulting in poor training results and it is difficult to identify forged faces.
By generating forgery operation instructions based on the image information and memory information of the agent, the original image is processed, candidate forgery images and ordinary text are generated, multi-modal samples are constructed, and the forgery detection model is trained based on these samples.
The generalization ability and accuracy of the forged detection model are improved, the recognition ability of new types of forged attacks is enhanced, and the problem of insufficient data sets of high-quality forged faces is solved.
Smart Images

Figure CN120452075A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of face detection technology, and in particular to an intelligent agent-based forged face detection method, apparatus, device, storage medium, and product. Background Art
[0002] In the current field of artificial intelligence, forged face detection technology is crucial for maintaining cybersecurity and social stability. However, due to the lack of large-scale, high-quality forged face datasets, the training results of existing forged face detection models are often unsatisfactory. This affects the accuracy of forged face detection models, making it difficult to identify forged faces.
[0003] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide an agent-based forged face detection method, device, equipment, storage medium and product, aiming to solve the technical problem that the lack of high-quality forged face datasets affects the training effect of the forged detection model, making it difficult to identify forged faces.
[0005] To achieve the above objectives, the present application proposes an agent-based forged face detection method, which includes: Generate forged operation instructions based on the agent's portrait information and memory information; Processing the original image according to the forgery operation instruction to obtain a candidate forged image, and generating a plain text of the candidate forged image; generating a multimodal sample according to the candidate forged image and the ordinary text; Based on the multimodal samples, a forgery detection model is trained to obtain a target forgery detection model, and forged faces are detected according to the target forgery detection model.
[0006] In one embodiment, the step of generating a forged operation instruction based on the portrait information and memory information of the agent includes: determining an inherent forgery preference based on the portrait information; Based on the memory information, the inherent forgery preference is adjusted according to a forgery tree to generate a forgery operation instruction, wherein the forgery tree is used to simulate the decision-making process of an intelligent agent when performing face forgery.
[0007] In one embodiment, the step of generating a multimodal sample based on the candidate forged image and the ordinary text includes: inputting the candidate forged image into an external face forgery detector to obtain a first evaluation score; Inputting the candidate forged image into a large language model to obtain a second evaluation score; Obtaining a target evaluation score based on the first evaluation score and the second evaluation score; Determine whether the target evaluation score is greater than or equal to a preset threshold; When the target evaluation score is greater than or equal to the preset threshold, a multimodal sample is generated according to the candidate forged image and the ordinary text.
[0008] In one embodiment, after the step of determining whether the target evaluation score is greater than or equal to a preset threshold, the method further includes: When the target evaluation score is greater than or equal to the preset threshold, determining the forgery attempt result of the candidate forged image as a successful forgery attempt; When the target evaluation score is less than the preset threshold, determining the forgery attempt result of the candidate forged image as a forgery attempt failure; using the first evaluation score, the second evaluation score, and the forgery attempt result of the candidate forged image as emotional memory information in the target attempt information, and using the forgery operation type and the forgery operation area of the candidate forged image as factual memory information in the target attempt information; The target attempt information is analyzed by a large language model to obtain common information of success and failure, and the memory information is updated according to the common information of success and failure.
[0009] In one embodiment, the step of analyzing the target attempt information using a large language model to obtain common information about success and failure, and updating the memory information based on the common information about success and failure includes: Acquire a target attempt information set within a short-term reflection period, wherein the target attempt information set includes a plurality of target attempt information; Acquire a long-term attempt information set, wherein the long-term attempt information set includes all target attempt information; The target attempt information set and the long-term attempt information set are analyzed by the large language model to obtain common information of success and failure, and the memory information is updated according to the common information of success and failure.
[0010] In one embodiment, the step of generating a multimodal sample based on the candidate forged image and the ordinary text includes: Posting the candidate forged image in a simulated social environment, having multiple intelligent agents with different roles interact with the candidate forged image to generate social text, and determining a text description based on the social text and the ordinary text; determining a forged label according to the original image and the candidate forged image; A multimodal sample is generated based on the original image, the candidate forged image, the text description, and the forged label.
[0011] In addition, to achieve the above-mentioned purpose, the present application also proposes an agent-based forged face detection device, the agent-based forged face detection device comprising: A generation module, used to generate forged operation instructions based on the agent's portrait information and memory information; a processing module, configured to process the original image according to the forgery operation instruction to obtain a candidate forged image, and generate a plain text of the candidate forged image; The generating module is further configured to generate a multimodal sample based on the candidate forged image and the ordinary text; The training module is used to train a forgery detection model based on the multimodal samples to obtain a target forgery detection model, and detect forged faces according to the target forgery detection model.
[0012] In addition, to achieve the above-mentioned purpose, the present application also proposes an agent-based fake face detection device, which includes: a memory, a processor, and a computer program stored on the memory and runnable on the processor, and the computer program is configured to implement the steps of the agent-based fake face detection method as described above.
[0013] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the agent-based fake face detection method as described above are implemented.
[0014] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the agent-based forged face detection method as described above.
[0015] One or more technical solutions proposed in this application have at least the following technical effects: The present application proposes an agent-based forged face detection method, apparatus, device, storage medium, and product. The method generates forged operation instructions based on the agent's portrait information and memory information; processes the original image according to the forged operation instructions to obtain a candidate forged image, and generates plain text of the candidate forged image; generates multimodal samples based on the candidate forged image and the plain text; trains a forged detection model based on the multimodal samples to obtain a target forged detection model, and detects forged faces based on the target forged detection model. This method solves the technical problem of difficulty in identifying forged faces due to the lack of a high-quality forged face dataset, which affects the training effect of the forged detection model. Compared to the existing technology, this application simulates agents with different portraits and different memories to perform face forgery, thereby generating more realistic and diverse high-quality multimodal samples. This method effectively solves the problem of insufficient high-quality forged face datasets, and can enhance the generalization and accuracy of the forged detection model, thereby improving the model's ability to recognize new types of forged attack faces. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 A flowchart of the first embodiment of the method for detecting fake faces based on an intelligent agent is provided in this application; Figure 2 A schematic diagram of a forged tree structure provided in Example 1 of the agent-based forged face detection method of this application; Figure 3 A flowchart illustrating the second embodiment of the method for detecting fake faces based on an intelligent agent is provided in this application; Figure 4 This is a schematic diagram of image quality assessment provided in Example 2 of the agent-based forged face detection method of this application; Figure 5 A schematic diagram of data construction provided for the second embodiment of the agent-based forged face detection method of this application; Figure 6 This is a schematic diagram of the module structure of the agent-based forged face detection device according to an embodiment of the present application; Figure 7Schematic diagram of the device structure of the hardware operating environment involved in the agent-based forged face detection method in the embodiment of the present application.
[0019] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0020] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0021] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0022] The main solution of the embodiment of the present application is: generating forgery operation instructions based on the portrait information and memory information of the intelligent body; processing the original image according to the forgery operation instructions to obtain a candidate forged image, and generating ordinary text of the candidate forged image; generating multimodal samples based on the candidate forged image and the ordinary text; training a forgery detection model based on the multimodal samples to obtain a target forgery detection model, and detecting forged faces based on the target forgery detection model.
[0023] As can be seen from the above embodiments, the present application generates forgery operation instructions based on the portrait information and memory information of the intelligent agent; processes the original image according to the forgery operation instructions to obtain a candidate forgery image, and generates ordinary text of the candidate forgery image; generates multimodal samples based on the candidate forgery image and the ordinary text; trains the forgery detection model based on the multimodal samples to obtain a target forgery detection model, and detects forged faces based on the target forgery detection model. This solves the technical problem that the lack of high-quality forgery face datasets affects the training effect of the forgery detection model, making it difficult to identify forged faces. Compared with the existing technology, the present application simulates intelligent agents with different portraits and different memories to perform face forgery, thereby generating more realistic and diverse high-quality multimodal samples. This method effectively solves the problem of insufficient high-quality forgery face datasets, can enhance the generalization ability and accuracy of the forgery detection model, and thus improves the model's ability to recognize new types of forged attack faces.
[0024] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution capabilities, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the aforementioned functions, such as an agent-based forged face detection device. This embodiment and the following embodiments will be described below using an agent-based forged face detection device as an example.
[0025] Based on this, the embodiment of the present application provides a method for detecting fake faces based on an intelligent agent, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the agent-based forged face detection method of this application.
[0026] In this embodiment, the agent-based forged face detection method includes steps S10 to S40: Step S10, generating a forging operation instruction based on the agent's portrait information and memory information; It should be noted that before constructing multimodal samples, the Agent4FaceForgery framework needs to be initialized. This includes defining a set of available forgery tools (e.g., Stable Diffusion, SBI, and other forgery methods), defining a set of participant roles in a social interaction environment (e.g., a student named Jack, 23 years old), and defining and initializing multiple LLM-driven agents (Agents). Each agent is assigned an initial profile, representing the profile of a specific social role. Specifically: the Observer agent (Watcher): Its profile is configured to tend to browse content and has weak judgment ability; the Explorer agent (Explorer): Its profile is configured to tend to discover inconsistencies by comparing relevant information; the Critic agent (Critic): Its profile is configured to tend to actively identify and point out signs of forgery; and the Gemini Auditor agent (Gemini Auditor): Its profile is specifically configured to issue deceptive statements that are contrary to the facts in response to forged content. Its role is to actively generate highly adversarial image-text inconsistency samples.
[0027] It should be noted that each agent's portrait information can be composed of a forgery preference and a priority attribute. The pseudo preference vector defines the agent's preference for different forgery types (e.g., local adjustment, full-face adjustment, or background replacement), while the priority attribute defines the specific facial regions that the agent focuses on when performing forgery (e.g., left eye, right eye, nose, or mouth). Memory information refers to the agent's historical experience, which includes factual memory information about past forgery attempts (e.g., performed editing operations) and emotional memory information. Finally, the agent's portrait information and memory information can be fused to obtain a final forgery operation instruction (i.e., specifying the specific visual editing operation to be performed on the original image). For example, a replacement operation is performed on the eye area of the original image.
[0028] In a feasible embodiment, the step of generating forgery operation instructions based on the portrait information and memory information of the intelligent agent includes: determining the inherent forgery preference according to the portrait information; based on the memory information, adjusting the inherent forgery preference according to the forgery tree to generate forgery operation instructions, wherein the forgery tree is used to simulate the decision-making process of the intelligent agent when performing face forgery.
[0029] In a specific implementation, the agent's inherent forgery preferences (such as the agent's preference for different forgery types and the specific facial areas that the agent focuses on when performing forgery) can be determined based on the agent's portrait information. This may involve a preference for specific facial features, such as face shape or glasses type. Then, based on the agent's memory information, which may include the results of past forgery attempts, these preferences are adjusted using the forgery tree. The structure of the forgery tree is as follows: Figure 2 As shown in [1], the root node is divided into first-level branches based on the local and whole face, and second-level nodes correspond to operable facial features (e.g., the glasses node includes nodes for frame type and reflective intensity). For example, starting from the root node, it is divided into two major branches: "local" and "whole face," which are further subdivided into specific features such as "face," "glasses," "nose," and "mouth." Each node represents a decision point. The agent chooses different paths in the tree, such as "ellipse" or "circle" under the "face" node, or "left" or "right" under the "glasses" node, to generate specific forgery operation instructions. This tree-based decision-making process enables the agent to systematically explore different forgery strategies and optimize its choices based on memory information, thereby generating more accurate and effective forgery operation instructions. This approach not only improves the quality and success rate of forged images, but also, by simulating the agent's decision-making process, makes the generated forgery operation instructions more consistent with the agent's experience and preferences.
[0030] In the specific implementation, the agent's tendency to forge specific facial areas (such as glasses and mouth shape) can be analyzed based on its portrait information to form an initial preference weight vector (i.e., the attention paid to each facial feature); then a forgery tree is constructed, and the root node is divided into first-level branches according to "partial-whole face", and the second-level nodes correspond to operable facial parts (such as the glasses node contains frame type and reflective intensity sub-nodes); each decision node is associated with the success rate matrix in the memory bank, and the matrix element S of the success rate matrix is ij It represents the historical success rate of the agent's forgery operation under the i-th lighting condition using the j-th parameter combination. When traversing the forgery tree, the path score needs to be calculated in real time:
[0031] Where Wp represents the agent's inherent preference weight vector (e.g., preference for glasses modification is 0.7, preference for lip shape modification is 0.3), X represents the feature vector of the original image (e.g., 512-dimensional features extracted by CNN), Indicates the degree of match between the feature vector of the original image and the inherent preference of the agent (if the glasses region feature in X is significant and the glasses weight in Wp is high, then this score is high).
[0032] It's important to note that when the path scores of multiple child nodes are similar (e.g., |Pscore1−Pscore2| < 0.05), the process must trace back to the nearest parent node to check whether its decision conditions are outdated (e.g., if the associated memory item hasn't been updated for more than six months). This requires recalculating the parent node's path score and selecting a new child node branch. Finally, the final path is converted into an executable set of operational instructions (i.e., forged operational instructions), including: region positioning coordinates (e.g., the BBox of the glasses area); feature modification parameters (e.g., lens reflectivity ±15%); and anti-detection injection points (e.g., adding noise of a specific frequency band to the edges of the glasses).
[0033] Step S20, processing the original image according to the forgery operation instruction to obtain a candidate forged image, and generating a plain text of the candidate forged image; It should be noted that the original image is randomly selected from a queue of real images to be processed; the queue of real images to be processed is composed of all images selected from an initial library of real face images (a portion of which is selected from public datasets such as FF++). Based on the forgery operation instructions, the original image is subjected to specific modifications or synthesis operations. These operations may include replacing facial features, modifying expressions, adjusting lighting, and replacing backgrounds to simulate real-world forgery scenarios, thereby generating candidate forged images.
[0034] In the specific implementation, the agent receives a forgery operation instruction and performs one or more forgery operations (using integrated forgery tools for face replacement, face generation, expression modification, etc.) to generate a candidate forged image. At the same time, during the visual editing process, the action module in the agent initiates a memory retrieval from the agent's memory module to obtain historical experience related to the forgery operation currently being performed. If the retrieved emotional memory indicates that the past use of a certain fusion parameter or technique (for example, a certain edge feathering radius) will result in obvious visual artifacts (for example, an assessment of an overly blurred edge area), the agent's action module will adjust the parameters or forgery synthesis method (such as changing the image lighting conditions or replacing the image background) in this operation to avoid repeating past mistakes. This adjustment is intended to dynamically optimize the execution details of the forgery operation based on historical experience to improve the realism of the final generated candidate forgery image. Specifically, the forged tool parameters (such as feather radius, light intensity, and GAN mixing coefficient) are constructed as a high-dimensional parameter space (that is, the discrete and isolated forged tool parameters are converted into continuous mathematical representations, so that the tool parameters and results of past historical operations are mapped into this high-dimensional space to form a data point distribution). The forged parameter tool recorded in the historical operation in the memory module is used to train a lightweight prediction model (such as a Bayesian optimization proxy model). When it is detected that the features of the current operation object (such as low-resolution faces and complex backgrounds) are similar to the forged failure cases in the memory, the model is called to generate parameter adjustment suggestions in real time.
[0035] In a specific implementation, plain text can be generated for the forged image based on the portrait information. There are two possibilities for the generated plain text: 1. If the agent's portrait information tends to be objective, a more accurate description of the facial edits performed on the original image will be generated. For example, if the visual edit performed is eye replacement, the generated plain text will be: The eyes of the person in the image have been replaced. 2. If the agent's portrait information tends to be deceptive, a misleading statement that contradicts the true state of the image will be generated. For example, even if the original image has been significantly forged, the generated plain text may be: This is an original image that has not been modified in any way.
[0036] Step S30, generating a multimodal sample based on the candidate forged image and the ordinary text; In a feasible implementation, the step of generating a multimodal sample based on the candidate forged image and the ordinary text includes: publishing the candidate forged image into a simulated social environment, having multiple intelligent agents with different roles interact with the candidate forged image to generate social text, and determining a text description based on the social text and the ordinary text; determining a forged label based on the original image and the candidate forged image; and generating a multimodal sample based on the original image, the candidate forged image, the text description, and the forged label.
[0037] In the implementation, candidate forged images that have been screened for training a forgery detection model are posted to a simulated social environment. Simulated user agents, each driven by a different LLM (e.g., Watcher, Critic, Explorer, GeminiAuditor), interact with the candidate forged images, generating comments, judgments, and sharing behaviors, thereby generating a series of social texts. Finally, the forgery label (y=1), social texts, and deceptive statements from the Gemini Auditor are combined to construct multimodal samples with complex image-text consistency challenges. For example, an image that has been pointed out as forged by a critic can be labeled as requiring the model to identify this inconsistency, even if the original text claims to be authentic.
[0038] It's important to note that agents with multiple different roles interact with candidate forged images. Specifically, an agent assigned the Critic portrait tends to generate cautious and questioning social text, potentially pointing out specific visual issues. For example, the agent might ask, "Wait a minute, the lighting directions of the person's left and right eyes in this image seem inconsistent, making them look like a post-production composite." Or, the edge where the neck and face meet is a bit blurry, which doesn't seem right. An agent assigned the Explorer portrait might generate social text that reflects judgments drawn through association or information comparison. For example, the agent might ask, "I think I've seen this background somewhere before, but the person in it isn't him. Is this image photoshopped?" An agent assigned the Gemini Auditor portrait is specifically designed to generate highly misleading and deceptive statements that contradict the facts.
[0039] Step S40: training a forgery detection model based on the multimodal samples to obtain a target forgery detection model, and detecting forged faces according to the target forgery detection model.
[0040] It's important to note that multimodal samples combine both image and text information, providing the model with more comprehensive forgery features, enabling it to learn and identify forged images from different dimensions. During training, the model analyzes these samples and continuously optimizes parameters to reduce prediction error, thereby improving its ability to identify forged images. After training, the resulting forgery detection model can accurately classify and identify new forged face images. The advantages of this approach include improved detection accuracy, enhanced adaptability to different forgery techniques, and improved generalization.
[0041] This embodiment generates forgery operation instructions based on the portrait information and memory information of the intelligent agent; processes the original image according to the forgery operation instructions to obtain a candidate forged image, and generates plain text of the candidate forged image; generates multimodal samples based on the candidate forged image and the plain text; trains a forgery detection model based on the multimodal samples to obtain a target forgery detection model, and detects forged faces based on the target forgery detection model. This solves the technical problem of difficulty in identifying forged faces due to the lack of high-quality forged face datasets. Compared with the existing technology, this application simulates intelligent agents with different portraits and different memories to perform face forgery, thereby generating more realistic and diverse high-quality multimodal samples. This method effectively solves the problem of insufficient high-quality forged face datasets, can enhance the generalization and accuracy of the forgery detection model, and thus improves the model's ability to recognize new types of forged attack faces.
[0042] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above introduction and will not be described in detail later. Figure 3 , step S30 also includes steps S301 to S305: Step S301, inputting the candidate forged image into an external face forgery detector to obtain a first evaluation score; It should be noted that the first evaluation score output by the external face forgery detector Xception is the probability dis that the candidate forged image is forged.
[0043] Step S302: inputting the candidate forged image into a large language model to obtain a second evaluation score; It should be noted that the second evaluation score output by the large language model is the image quality score for the candidate forged model.
[0044] Step S303: obtaining a target evaluation score based on the first evaluation score and the second evaluation score; It should be noted that the target evaluation score of the candidate forged image is obtained by weighted fusion of the two evaluation scores. In the first part, the first evaluation score is the objective deception difficulty score. The candidate forged image needs to be input into a preset external face forgery detector Xception to output the first evaluation score (that is, the probability dis that the candidate forged image is forged); in the second part, the second evaluation score is the subjective forgery quality score LLM. The candidate forged image needs to be input into the large language model to output the second evaluation score.
[0045] In a specific implementation, the target evaluation score can be determined based on the first evaluation score and the second evaluation score. The specific calculation formula is as follows:
[0046] Where dis represents the first evaluation score, LLM represents the second evaluation score, and sn represents the target evaluation score.
[0047] Step S304, determining whether the target evaluation score is greater than or equal to a preset threshold; It should be noted that when the target evaluation score of a candidate forged image is greater than or equal to a preset threshold, the forgery attempt on the candidate forged image can be considered successful. At this time, the candidate forged image will be used as a multimodal sample to train the forgery detection model. When the target evaluation score of a candidate forged image is less than the preset threshold, the forgery attempt on the candidate forged image can be considered a failure. At this time, the candidate forged image will not be used to train the forgery detection model.
[0048] In a feasible embodiment, after the step of determining whether the target evaluation score is greater than or equal to a preset threshold, the step further includes: when the target evaluation score is greater than or equal to the preset threshold, determining the forgery attempt result of the candidate forged image as a successful forgery attempt; when the target evaluation score is less than the preset threshold, determining the forgery attempt result of the candidate forged image as a failed forgery attempt; using the first evaluation score, the second evaluation score and the forgery attempt result of the candidate forged image as emotional memory information in the target attempt information, and using the forgery operation type and the forgery operation area of the candidate forged image as factual memory information in the target attempt information; analyzing the target attempt information through a large language model to obtain common information of success and failure, and updating the memory information based on the common information of success and failure.
[0049] It should be noted that if Figure 4As shown, when the target evaluation score is greater than or equal to the preset threshold, the forgery attempt result of the candidate forged image is determined to be successful, which means that the forged image is visually realistic enough to deceive the forgery detector and the large language model; when the target evaluation score is less than the preset threshold, the forgery attempt result of the candidate forged image is determined to be failed, which indicates that the forged image has obvious visual defects and is easily recognized by the detector.
[0050] It can be understood that whether the forgery attempt of the candidate forged image is successful or failed, the result of this forgery attempt (including operation, image, text, evaluation score, etc., that is, target attempt information) will be memorized into the memory model of the agent.
[0051] It should be noted that target attempt information includes both emotional memory information and factual memory information. The emotional memory information includes the candidate forged image's first evaluation score (objective deception difficulty score), second evaluation score (subjective forgery quality score), and forgery attempt results. This information reflects the emotional and subjective evaluation of the forgery attempt, helping the agent understand the success or failure of the forgery attempt and the reasons behind it. The factual memory information includes the forgery operation type and forgery operation area of the candidate forged image. This information provides specific details about the forgery operation, such as the techniques used and the parts of the image that were modified.
[0052] As can be understood, a large language model is used to analyze target attempt information and extract common characteristics between success and failure. This common information may include common operation types and operation areas in successful forgery attempts, as well as common defects in failed forgery attempts. By analyzing this common information, the agent can learn which operations are more likely to lead to successful forgeries and which operations are more easily identified by the detector, thereby optimizing its forgery strategy. Based on the extracted common information, the agent's memory information is updated. This update helps the agent make more informed decisions in future forgery attempts, improve the quality of forged images, and thus generate more realistic and difficult-to-detect forgeries.
[0053] In a feasible implementation, the steps of analyzing the target attempt information through a large language model to obtain common information on success and failure, and updating the memory information based on the common information on success and failure include: obtaining a target attempt information set within a short-term reflection period, wherein the target attempt information set includes multiple target attempt information; obtaining a long-term attempt information set, wherein the long-term attempt information set includes all target attempt information; analyzing the target attempt information set and the long-term attempt information set through the large language model to obtain common information on success and failure, and updating the memory information based on the common information on success and failure.
[0054] It should be noted that the target attempt information set within the short-term reflection period refers to the target attempt information collected within a recent period (e.g., the last few attempts) (i.e., short-term memory). This information reflects the agent's performance and results in recent forgery attempts. The long-term target attempt information set includes all target attempt information collected by the agent over a longer period of time (e.g., throughout the entire training process) (i.e., long-term memory). This information provides a more comprehensive perspective, helping the agent understand its performance at different time scales and long-term trends.
[0055] It should be noted that the agent periodically (for example, after every N attempts) reflects on its memory. The LLM analyzes past successes and failures and adjusts its subsequent forgery strategy. Specifically, this process is performed internally by the agent through a collaborative effort between the memory module and the action module. This process enables the agent to learn from accumulated experience and dynamically adjust its future forgery strategy. This step is triggered periodically under a preset condition, for example, after the agent completes N forgery attempts. Once triggered, the agent's memory module initiates a reflection operation. During this operation, the memory module first retrieves a complete record of the last N forgery attempts from its stored factual and emotional memories. The agent's internal core then analyzes this record to identify recurring patterns and challenges. For example, the LLM core might analyze and summarize the common failure patterns of a forgery technique under specific conditions (such as when processing high-angle profile images), thereby providing the agent with guidance for improvement. Through this periodic reflection and analysis, the agent is able to learn from accumulated experience and dynamically adjust its future forgery strategy. This mechanism enables the agent to continuously adapt and optimize its behavior to improve its success rate and efficiency. This self-learning and self-optimization capability is achieved through the collaborative work of its internal memory module and action module. Together, they influence the agent's decision-making process, enabling it to continuously improve and grow in complex and changing environments.
[0056] In the specific implementation, Figure 5 As shown in the figure, after multiple rounds of simulation using Adaptive Rejection Sampling (ARS) are used to identify multiple candidate forged images, multiple agents with different roles interact with these candidate forged images to obtain the corresponding social context. Finally, the forged labels (y=1), social context, and the deceptive statements of the "Gemini Auditor" are combined to construct multimodal samples that meet the complex image-text consistency challenge.
[0057] Step S305 : When the target evaluation score is greater than or equal to the preset threshold, a multimodal sample is generated according to the candidate forged image and the ordinary text.
[0058] It is understandable that when the target evaluation score is greater than or equal to a preset threshold, the candidate forged model needs to be processed to generate a multimodal sample.
[0059] This embodiment obtains a first evaluation score by inputting the candidate forged image into an external face forgery detector; obtains a second evaluation score by inputting the candidate forged image into a large language model; obtains a target evaluation score based on the first and second evaluation scores; determines whether the target evaluation score is greater than or equal to a preset threshold; and when the target evaluation score is greater than or equal to the preset threshold, generates a multimodal sample based on the candidate forged image and the plain text. By setting a threshold, the system can screen out high-quality forged images, which, along with their corresponding plain text descriptions, are used to generate multimodal samples. These samples not only enrich the diversity of training data but also improve the training quality and efficiency of the forgery detection model, thereby enhancing the model's ability to recognize forged images. In addition, this method promotes the self-learning and optimization of the intelligent agent during the forgery process, enabling it to continuously adjust its forgery strategy based on the evaluation results to generate more realistic and difficult-to-detect forged images, ultimately achieving a dual improvement in forged image generation and detection technology.
[0060] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the fake face detection method based on intelligent agents in the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0061] This application also provides a fake face detection device based on intelligent agent, please refer to Figure 6 , the agent-based forged face detection device comprises: A generation module 10, configured to generate a forged operation instruction based on the agent's portrait information and memory information; a processing module 20 for processing the original image according to the forgery operation instruction to obtain a candidate forged image and generating a plain text of the candidate forged image; The generating module 10 is further configured to generate a multimodal sample based on the candidate forged image and the ordinary text; The training module 30 is configured to train a forgery detection model based on the multimodal samples to obtain a target forgery detection model, and detect forged faces according to the target forgery detection model.
[0062] The agent-based forged face detection device provided in this application, which utilizes the agent-based forged face detection method described in the aforementioned embodiments, can address the technical issue of difficulty identifying forged faces due to a lack of high-quality forged face datasets, which hinders the training of forged face detection models. Compared to the prior art, the beneficial effects of the agent-based forged face detection device provided in this application are the same as those of the agent-based forged face detection method described in the aforementioned embodiments. Other technical features of the agent-based forged face detection device are the same as those disclosed in the aforementioned embodiments and are not further elaborated upon here.
[0063] The present application provides an agent-based forged face detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the agent-based forged face detection method of the above-mentioned embodiment one.
[0064] Reference below Figure 7 , which shows a schematic structural diagram of an agent-based forged face detection device suitable for implementing embodiments of the present application. The agent-based forged face detection device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The agent-based fake face detection device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0065] like Figure 7As shown, the agent-based forged face detection device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the agent-based forged face detection device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, a magnetic tape or hard disk; and a communication device 1009. The communication device 1009 can allow the agent-based forged face detection device to communicate wirelessly or wired with other devices to exchange data. Although the figure shows an agent-based forged face detection device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented or have alternatively.
[0066] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0067] The agent-based forged face detection device provided in this application utilizes the agent-based forged face detection method described in the aforementioned embodiment, resolving the technical issue of difficulty identifying forged faces due to a lack of high-quality forged face datasets, which hinders the training of forged detection models. Compared to the prior art, the beneficial effects of the agent-based forged face detection device provided in this application are the same as those of the agent-based forged face detection method described in the aforementioned embodiment. Other technical features of the agent-based forged face detection device are the same as those disclosed in the aforementioned embodiment and are not further elaborated upon here.
[0068] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0069] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0070] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, wherein the computer-readable program instructions are used to execute the agent-based forged face detection method in the above-mentioned embodiment.
[0071] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0072] The computer-readable storage medium may be included in the agent-based forged face detection device; or may exist independently without being assembled into the agent-based forged face detection device.
[0073] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the intelligent agent-based forged face detection device, the intelligent agent-based forged face detection device: generates forged operation instructions based on the intelligent agent's portrait information and memory information; processes the original image according to the forged operation instructions to obtain a candidate forged image, and generates ordinary text of the candidate forged image; generates multimodal samples based on the candidate forged image and the ordinary text; trains a forged detection model based on the multimodal samples to obtain a target forged detection model, and detects forged faces based on the target forged detection model.
[0074] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0075] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0076] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0077] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned agent-based forged face detection method. This computer-readable storage medium can address the technical issue of difficulty identifying forged faces due to a lack of high-quality forged face datasets, which hinders the training of forged face detection models. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are similar to those of the agent-based forged face detection method provided in the aforementioned embodiments and are not further elaborated here.
[0078] The present application also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the above-mentioned agent-based forged face detection method.
[0079] The computer program product provided in this application can address the technical issue of difficulty identifying forged faces due to the lack of high-quality forged face datasets, which hinders the training of forged face detection models. Compared to the prior art, the beneficial effects of the computer program product provided in this application are similar to those of the agent-based forged face detection method provided in the aforementioned embodiments, and are not further elaborated here.
[0080] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A method for detecting fake faces based on an intelligent agent, characterized in that: The agent-based forged face detection method comprises: Generate forged operation instructions based on the agent's portrait information and memory information; Processing the original image according to the forgery operation instruction to obtain a candidate forged image, and generating a plain text of the candidate forged image; generating a multimodal sample according to the candidate forged image and the ordinary text; Based on the multimodal samples, a forgery detection model is trained to obtain a target forgery detection model, and forged faces are detected according to the target forgery detection model.
2. The method for detecting fake faces based on an agent as claimed in claim 1, wherein: The step of generating a forged operation instruction based on the portrait information and memory information of the intelligent agent includes: determining an inherent forgery preference based on the portrait information; Based on the memory information, the inherent forgery preference is adjusted according to a forgery tree to generate a forgery operation instruction, wherein the forgery tree is used to simulate the decision-making process of an intelligent agent when performing face forgery.
3. The method for detecting fake faces based on an agent as claimed in claim 1, wherein: The step of generating a multimodal sample based on the candidate forged image and the ordinary text includes: inputting the candidate forged image into an external face forgery detector to obtain a first evaluation score; Inputting the candidate forged image into a large language model to obtain a second evaluation score; Obtaining a target evaluation score based on the first evaluation score and the second evaluation score; Determine whether the target evaluation score is greater than or equal to a preset threshold; When the target evaluation score is greater than or equal to the preset threshold, a multimodal sample is generated according to the candidate forged image and the ordinary text.
4. The method for detecting fake faces based on an agent as claimed in claim 3, wherein: After the step of determining whether the target evaluation score is greater than or equal to a preset threshold, the method further includes: When the target evaluation score is greater than or equal to the preset threshold, determining the forgery attempt result of the candidate forged image as a successful forgery attempt; When the target evaluation score is less than the preset threshold, determining the forgery attempt result of the candidate forged image as a forgery attempt failure; using the first evaluation score, the second evaluation score, and the forgery attempt result of the candidate forged image as emotional memory information in the target attempt information, and using the forgery operation type and the forgery operation area of the candidate forged image as factual memory information in the target attempt information; The target attempt information is analyzed by a large language model to obtain common information of success and failure, and the memory information is updated according to the common information of success and failure.
5. The method for detecting fake faces based on an agent as claimed in claim 4, wherein: The step of analyzing the target attempt information by using a large language model to obtain common information of success and failure, and updating the memory information according to the common information of success and failure includes: Acquire a target attempt information set within a short-term reflection period, wherein the target attempt information set includes a plurality of target attempt information; Acquire a long-term attempt information set, wherein the long-term attempt information set includes all target attempt information; The target attempt information set and the long-term attempt information set are analyzed by the large language model to obtain common information of success and failure, and the memory information is updated according to the common information of success and failure.
6. The method for detecting fake faces based on an agent as claimed in claim 1, wherein: The step of generating a multimodal sample according to the candidate forged image and the ordinary text includes: Posting the candidate forged image in a simulated social environment, having multiple intelligent agents with different roles interact with the candidate forged image to generate social text, and determining a text description based on the social text and the ordinary text; determining a forged label according to the original image and the candidate forged image; A multimodal sample is generated based on the original image, the candidate forged image, the text description, and the forged label.
7. An agent-based forged face detection device, characterized in that: The forged face detection device based on intelligent agent comprises: A generation module, used to generate forged operation instructions based on the agent's portrait information and memory information; a processing module, configured to process the original image according to the forgery operation instruction to obtain a candidate forged image, and generate a plain text of the candidate forged image; The generating module is further configured to generate a multimodal sample based on the candidate forged image and the ordinary text; The training module is used to train a forgery detection model based on the multimodal samples to obtain a target forgery detection model, and detect forged faces according to the target forgery detection model.
8. An agent-based forged face detection device, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the agent-based forged face detection method according to any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the forged face detection method based on an intelligent agent are implemented as described in any one of claims 1 to 6.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the agent-based forged face detection method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Face forgery detection method and device based on visual semantic information
CN113627233A
Lip shape synchronization face forgery generation method and system based on image completion
CN114663962A
Face forgery detection model construction method, face forgery detection method and face forgery detection device
CN115240243A
Face forgery attack detection method, face recognition method, device and equipment
CN117197857A
Face anti-counterfeiting recognition model training method, face anti-counterfeiting recognition method and face anti-counterfeiting recognition device
CN118799947A