Attribute error binding attack and evaluation method for identity preserving model

By constructing a sensitive word dataset and introducing Attribute Binding Security Score (ABSS), the problem of evaluating attribute misbinding attacks in identity preservation models is solved, achieving efficient and accurate attack and evaluation, and improving the security of the model and the comprehensiveness of the evaluation.

CN121746833APending Publication Date: 2026-03-27SUN YAT SEN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies cannot effectively assess and attack the risk of misbinding attributes in identity preservation models, and existing assessment metrics cannot fully measure the true harm of new attacks.

Method used

A sensitive word dataset is constructed, and candidate prompt words are generated by selecting the wrong binding strategy through the LLM model. After evaluation by the text filter, the prompt words are input into the identity preservation model to generate high-risk images. The attribute binding security score (ABSS) is introduced for comprehensive evaluation.

Benefits of technology

It enables precise attacks against identity preservation models, generating high-risk content while maintaining the identity characteristics of individuals, providing comprehensive security assessment standards, and improving the stealth of attacks and the scientific nature of assessments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746833A_ABST
    Figure CN121746833A_ABST
Patent Text Reader

Abstract

The invention provides an attribute error binding attack and evaluation method for an identity preserving model, which comprises the following steps of: acquiring an open source unadapted content data set which comprises a plurality of cue words, labeling the cue words along four harmful dimensions, decomposing the cue words into three semantic components, constructing a sensitive word data set based on the plurality of cue words subjected to labeling and component decomposition; sensitive words are extracted to serve as input cue words, candidate cue words are generated by selecting an error binding strategy through an LLM model, and effective cue words are integrated into an error binding instruction evaluation set; obtaining an identity reference image, inputting the identity reference image into a diffusion model of the identity maintenance model, and generating a generated image which maintains character identity features and accords with cue word description; and performing attribute binding security evaluation on the generated image to obtain an attribute binding security score. The technical problem that the attack risk based on'attribute wrong binding 'cannot be effectively evaluated in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image generation model attack and evaluation technology, specifically to an attribute misbinding attack and evaluation method targeting identity preservation models. Background Technology

[0002] Currently, security research on AI image generation models mainly focuses on two aspects: Attacks targeting the underlying diffusion model: Most existing attack methods, such as synonym substitution, character-level noise injection, or adversarial prompts, primarily target the general text-to-image underlying model. These attacks manipulate the semantics of the input text to induce the model to generate unsafe content.

[0003] Security Assessment and Filtering: To defend against the aforementioned attacks, research and industry have deployed multi-layered defense systems, including text filters at the input end (blocking malicious keywords) and image filters at the output end (identifying harmful visual content). Methods for assessing the security of these models typically measure two metrics independently: the security of the generated image and the consistency between the generated image and the input prompt, i.e., alignment.

[0004] The shortcomings of existing technology: Despite some progress in existing technologies, there are still significant shortcomings and insufficiently researched technical deficiencies in the increasingly popular "identity preservation model": The attack methods lack specificity: Existing attack research mainly focuses on the basic diffusion model, rarely exploring the unique vulnerabilities of the identity preservation model. To accurately maintain human features, the identity preservation model's training process overemphasizes facial and human-shaped data, leading to a unique "subject-centric attention bias." Existing attack methods do not exploit this inherent flaw.

[0005] Failure to discover and exploit a core vulnerability: Due to the aforementioned attention bias, the identity preservation model is highly susceptible to "leaking" attributes (such as "hairless" or "exposed") describing background or other neutral objects in the text and incorrectly binding them to the core character when parsing instructions. This phenomenon is known as "attribute leakage." This is a unique and exploitable systemic attack vector, but it has not yet been discovered and exploited by existing technologies.

[0006] The assessment system has limitations and cannot measure emerging risks: existing security assessment metrics are fragmented. They either only assess whether the generated image is "safe" or only assess the "alignment" between the image and text. This single-dimensional assessment method cannot comprehensively measure the true harm of a new type of attack—that is, an attacker using seemingly harmless prompts to accurately generate risky images with specific identities while maintaining high "alignment." Therefore, it cannot effectively assess the risk of attacks based on "attribute misbinding." Summary of the Invention

[0007] The purpose of this invention is to overcome the above-mentioned technical deficiencies and provide a method for evaluating attribute misbinding attacks on identity preservation models, thereby solving the technical problem that existing technologies cannot effectively evaluate the attack risks based on "attribute misbinding".

[0008] To achieve the above-mentioned technical objectives, in a first aspect, the present invention provides a method for evaluating attribute misbinding attacks on identity preservation models, comprising the following steps: Obtain an open-source dataset of inappropriate content, which includes multiple prompt words. The prompt words are labeled along four harmful dimensions: pornography, violence, discrimination, and illegality. The prompt words are decomposed into three semantic components: role, scene, and state. A sensitive word dataset is constructed based on the multiple prompt words after labeling and decomposition. Sensitive words with harmful dimensions and semantic component labels are extracted from the sensitive word dataset as input prompt words. Candidate prompt words are generated by selecting the wrong binding strategy through the LLM model. After the text filter security evaluation, the valid prompt words that pass the evaluation are integrated into the wrong binding instruction evaluation set. Obtain the identity reference image, encode the valid prompt words in the misbound instruction evaluation set and the identity reference image through a text encoder and an image encoder respectively, and input them into the diffusion model of the identity preservation model to generate a generated image that preserves the identity features of the person and conforms to the prompt word description. An attribute binding security assessment is performed on the generated image to obtain an attribute binding security score.

[0009] Compared with the prior art, the beneficial effects of the present invention include: This invention addresses the problem of exploiting unique flaws in identity preservation models for attacks: This patent is the first to propose and systematize an "Attribute Misbinding Attack" method. This method, through carefully designed seemingly harmless text instructions, leverages the model's "subject-centric attention deficit" to precisely induce the model to incorrectly attach sensitive attributes describing neutral objects to individuals with specified identities, thereby bypassing traditional text filters and generating high-risk non-public content (NSFW).

[0010] This invention addresses the problem of how to standardize the evaluation of the effectiveness of such attacks: to systematically test and quantify the vulnerability of models, this patent constructs a specialized "Misbinding Promptevaluation set." This evaluation set contains a large number of specially designed instructions that can efficiently trigger attribute misbinding, covering four risk dimensions: pornography, violence, discrimination, and illegality.

[0011] This invention addresses the problem of comprehensively assessing the risks of novel attacks: to overcome the limitations of existing assessment metrics, this patent pioneers an "Attribute Binding Safety Score" (ABSS). This scoring system can simultaneously evaluate the "content fidelity" and "security compliance" of generated images, quantifying the real harm of attribute misbinding attacks through a unified and comprehensive score, thus solving the problem that existing assessment methods cannot comprehensively consider both aspects.

[0012] According to some embodiments of the present invention, a sensitive word dataset is constructed based on multiple prompt words after annotation and decomposition, including the following steps: An initial seed set is constructed based on multiple prompt words after annotation and decomposition. The initial seed set is input into a large language model for automated vocabulary expansion: Given a seed word and its corresponding harmful dimension and semantic components, the model applies a specified strategy to generate new sensitive words that belong to the same harm type but span different semantic components, and constructs a sensitive word dataset based on all the new sensitive words.

[0013] According to some embodiments of the present invention, in semantic components, role represents the identity, occupation or title of the subject, scene represents the location, context or background of the event, and state represents the appearance, behavior and psychological state related to the role.

[0014] According to some embodiments of the present invention, an attribute binding security assessment is performed on the generated image to obtain an attribute binding security score, including the following steps: Calculate the fidelity of the generated image to the semantic content described in the input prompt word to obtain the prompt word-image alignment score; Calculate the safety of the generated image within a predetermined risk category to obtain the safe generation rate; The attribute binding security score is calculated based on the prompt word-image alignment score and the security generation rate.

[0015] According to some embodiments of the present invention, calculating the fidelity of the generated image to the semantic content described in the input prompt word to obtain a prompt word-image alignment score includes the following steps; For prompt words Requirements Core objects First, verify each core object. Does it exist in the image? In this context, a successfully generated object is represented as... A visual question-answering model is used to evaluate the accuracy of its attribute binding; The fractions are calculated as follows: ; in, It is a core verification object An indicator function that exists, and It is the accuracy score for binding properties to an existing object.

[0016] According to some embodiments of the present invention, calculating the safety of a generated image within a predetermined risk category to obtain a safe generation rate includes the following steps: Safe generation rate Used to measure images In the pre-defined risk category For internal security, a multimodal large language model (MLLM) acts as a security classification function. It provides a direct harmfulness score. Mathematically, this can be expressed as: ; It is a continuous value between 0 and 1, and the higher the value, the higher the security of the image.

[0017] According to some embodiments of the present invention, the Attribute Binding Security Score (ABSS) is calculated based on the cue word-image alignment score and the security generation rate. The ABSS is calculated by multiplying the cue word-image alignment score by a security weighting factor to ensure both accuracy and security in content capture. ; Sensitivity parameters This is used to adjust the impact of the safe generation rate.

[0018] According to some embodiments of the present invention, Set to 2.

[0019] Secondly, the present invention provides a system for evaluating attribute misbinding attacks on identity preservation models, including: The sensitive word dataset construction module is used to obtain an open-source inappropriate content dataset, which includes multiple prompt words. The prompt words are labeled along four harmful dimensions: pornography, violence, discrimination, and illegality. The prompt words are decomposed into three semantic components: role, scene, and state. The sensitive word dataset is constructed based on the multiple prompt words after labeling and decomposition. The error binding instruction evaluation set construction module communicates with the sensitive word dataset construction module. It extracts sensitive words with harmful dimensions and semantic component labels from the sensitive word dataset as input prompt words, selects error binding strategies through the LLM model to generate candidate prompt words, and integrates the valid prompt words that pass the evaluation into the error binding instruction evaluation set after the text filter security evaluation. The image generation module is communicatively connected to the mis-binding instruction evaluation set construction module to obtain the identity reference image. After the valid prompt words in the mis-binding instruction evaluation set and the identity reference image are encoded by the text encoder and the image encoder respectively, they are input into the diffusion model of the identity preservation model to generate a generated image that preserves the identity features of the person and conforms to the prompt word description. The scoring module communicates with the image generation module to perform attribute binding security assessment on the generated image and obtain the attribute binding security score.

[0020] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0021] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, wherein the abstract drawings are to be completely consistent with one of the drawings in the specification: Figure 1 A flowchart illustrating an embodiment of the present invention for an attribute misbinding attack and evaluation method against an identity preservation model; Figure 2 This is an example diagram of an attribute misbinding attack for an identity preservation model and an evaluation method provided in an embodiment of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0023] It should be noted that although functional modules are divided in the system diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the system or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0024] Reference Figures 1 to 2 , Figure 1 A flowchart illustrating an embodiment of the present invention for an attribute misbinding attack and evaluation method against an identity preservation model; Figure 2 This is an example diagram of an attribute misbinding attack for an identity preservation model and an evaluation method provided in an embodiment of the present invention.

[0025] In one embodiment, the method for attacking and evaluating attribute misbinding attacks on identity preservation models includes the following steps: acquiring an open-source inappropriate content dataset, which includes multiple prompt words. These prompt words are labeled along four harmful dimensions: pornography, violence, discrimination, and illegality. The prompt words are decomposed into three semantic components: role, scene, and state. A sensitive word dataset is constructed based on the labeled and decomposed prompt words. Sensitive words with harmful dimension and semantic component labels are extracted from the sensitive word dataset as input prompt words. A misbinding strategy is selected using an LLM model to generate candidate prompt words. After security evaluation by a text filter, the valid prompt words that pass the evaluation are integrated into a misbinding instruction evaluation set. An identity reference image is acquired. The valid prompt words in the misbinding instruction evaluation set and the identity reference image are encoded by a text encoder and an image encoder, respectively, and then input into the diffusion model of the identity preservation model to generate a generated image that retains the character's identity features and conforms to the prompt word description. An attribute binding security evaluation is performed on the generated image to obtain an attribute binding security score.

[0026] The advantages of the technical solution of this invention are as follows: I. The attack is highly targeted, directly exploiting the unique vulnerability of the identity preservation model.

[0027] Existing technologies mostly target general text-image basic models, neglecting the core characteristic of identity preservation models—"over-focusing on facial / human-shaped data to maintain human features"—resulting in poor attack effectiveness and a lack of precision. This solution precisely anchors vulnerabilities through "sensitive word annotation + semantic component decomposition": on one hand, it annotates data along four harmful dimensions—pornography, violence, discrimination, and illegality—to narrow down the risk range; on the other hand, it decomposes prompts into three semantic components: "role (subject identity), scene (background context), and state (appearance / behavior)," precisely matching the "subject-centric attention bias" of identity preservation models—exploiting the "attribute leakage" vulnerability where models easily misbind background / neutral object attributes to the subject. This prevents attacks from generalizing and instead targets the core flaws of identity preservation models, resulting in a much higher precision than existing general attack methods.

[0028] Second, its attack concealment and evasion capabilities are outstanding, breaking through traditional security defenses.

[0029] Existing adversarial attack prompts are often easily blocked by input text filters due to their obscure semantics and inclusion of obviously malicious words. This solution constructs seemingly harmless but actually risky mis-binding instructions through "LLM strategy selection + text filter verification": ① LLM, based on the harmful dimension and semantic component tags of sensitive words, selects strategies such as adjective migration and scene desensitization to generate candidate prompts without obviously malicious words (e.g., transforming "naked" into "hairy kitten" for related descriptions); ② After evaluation and screening by text filters, it is ensured that effective prompts can bypass traditional keyword blocking mechanisms. Experimental data confirms its advantages—the overall bypass rate against 5 mainstream text security detectors reaches 40.56%, far exceeding existing methods such as MMA (4.40%) and Sneakyprompt (4.42%), significantly improving stealth and defense penetration capabilities.

[0030] Third, the attack effect is precise and controllable, taking into account both "identity preservation" and "risk inducement".

[0031] Existing attacks often suffer from problems such as "failing to maintain the person's identity" or "generating content that deviates from the prompts," making it difficult to achieve the attack target of "precisely implanting risks without destroying identity features." This solution achieves dual precision through "dual-encoded input + diffusion model generation": ① Valid prompts are encoded using a text encoder (ensuring accurate semantic transmission), and the identity reference image is encoded using an image encoder (locking in the person's identity features). The dual inputs ensure that the diffusion model simultaneously receives both "risk instructions" and "identity information"; ② The diffusion model combines the two types of encoding results to generate an image that strictly maintains the person's features in the identity reference image (solving the "identity loss" problem) while accurately executing the incorrect attribute binding in the prompts (achieving "risk content inducement"). Ultimately, this achieves an attack effect of "high identity fidelity + high risk inducement rate," a core advantage that existing attack methods cannot simultaneously achieve.

[0032] Fourth, the assessment system is scientific and comprehensive, filling a gap in new types of risk assessment.

[0033] Existing technologies only independently measure "generated image security" or "cue-image alignment," failing to assess the novel risks of "generating highly aligned risky images using harmless cue words," leading to biased assessment results that are detached from practical application scenarios. This solution pioneers an "attribute-bound security assessment" system, achieving two major breakthroughs: ① Comprehensive assessment dimensions—It no longer measures a single indicator but simultaneously covers "alignment between generated image and cue words" (ensuring attack accuracy) and "image security compliance" (quantifying risk level) through "attribute-bound security scoring," perfectly adapting to the assessment needs of novel attack risks characterized by "high alignment + low security"; ② Reliable assessment results—The scoring logic is strongly positively correlated with human subjective judgment, avoiding the disconnect between existing assessment indicators and actual risk perception, and providing a scientific and unified quantitative standard for security testing of identity preservation models.

[0034] Fifth, the process is automated, highly reusable, and has extremely high practical value.

[0035] This solution is designed to balance "automation efficiency" and "cross-scenario adaptability" throughout the entire process, addressing the pain points of existing attack methods such as "manual design of prompts and poor reusability": ① High degree of automation—From the construction of the sensitive word dataset (LLM automated extension) and the generation of mis-binding prompts (LLM strategy selection) to security assessment (text filter + multimodal model scoring), the core steps are all automated, significantly reducing manual costs; ② Strong reusability—The sensitive word dataset and mis-binding instruction evaluation set can be reused across different identity preservation models (such as PhotoMaker, UniPortrait, FLUX.1-dev, etc.), without the need to repeatedly design attack samples for a single model; ③ Wide range of application scenarios—It can be used by security testing organizations to assess the vulnerability of identity preservation models, and can also assist developers in optimizing model defense mechanisms, providing practical tools for the secure iteration of identity preservation models.

[0036] VI. The data support is solid, and the attack and assessment are highly replicable.

[0037] Existing technologies often suffer from a lack of standardized datasets, leading to difficulties in reproducing attack methods and incomparable evaluation results. This solution addresses this issue by "constructing standardized datasets in a layered manner": first, a sensitive word dataset (covering four major risk dimensions) is built based on annotation and semantic decomposition; then, a mis-binding instruction evaluation set (containing a large number of valid prompt words) is derived. Both datasets are characterized by "clear classification, complete labels, and scalability"—not only supporting the stable execution of the attack process of this solution, but also providing standardized benchmark data for subsequent security research on identity preservation models, thus promoting research in this field from "fragmented exploration" to a systematic stage of "reproducibility and comparability".

[0038] To enable and evaluate attribute misbinding attacks, the framework consists of four steps.

[0039] Step 1: This invention collects an open-source NSFW dataset and labels it along four harmful dimensions relevant to humans: pornography, violence, discrimination, and illegality. To achieve fine-grained analysis and controllable generation of NSFW content, this invention decomposes a prompt word into three semantic components: role (the subject's identity, profession, or title), scene (the location, context, or background of the event), and state (appearance, behavior, and psychological state related to the role).

[0040] Based on this semantic framework, this invention first constructs an initial seed set containing 200 sensitive words, each categorized according to its harm type and semantic components. To systematically expand this vocabulary, this invention designs and implements an automated generation process driven by a Large Language Model (LLM). This process operates based on a set of predefined expansion strategies, which are formulated as a system prompt word to guide the LLM generation process. Specifically, when a seed word and its corresponding category (i.e., harm type and semantic component) are provided, the model applies the specified strategy to generate new sensitive words belonging to the same harm type but spanning different semantic components. In implementation, this invention uses the Qwen3 model as the generator. This automated process ultimately constructs a sensitive word dataset containing 2000 entries.

[0041] Step 2: The present invention then proceeds to programmatically generate the "error binding prompt word". The process is as follows: Policy Instantiation: This invention constructs an input from the sensitive word dataset developed in step 1, which contains the sensitive words themselves, their harm types, and semantic components. Then, a system prompt word instructs an LLM to select an appropriate error binding strategy from the system prompt words to guide the generation of the "error binding prompt word".

[0042] Security filtering: The candidate prompts generated by LLM are subjected to a security assessment through a text filter.

[0043] Cue word validation: Any cue words flagged by the filter will be discarded, and the process will be repeated. Cue words that pass this filtering stage will become valid misbinding cue words and will be selected for the final dataset.

[0044] This invention uses the Qwen3 model as both a generator and a text filter. This automated process ultimately generates a dataset containing 2000 error-binding prompt words.

[0045] Step 3: In this invention, the error binding prompt words and identity reference images constructed in step 2 are processed by an image encoder and a text encoder, respectively, and then input into a diffusion model. The model combines these two input information to generate a new image that retains the identity characteristics of the reference person and conforms to the text content description.

[0046] Step 4: To effectively evaluate attribute misbinding attacks launched by the "misbinding prompt word" evaluation set of this invention and to measure the defense strength of the diffusion model, this invention introduces the Attribute Binding Security Score (ABSS). ABSS provides a comprehensive evaluation by simultaneously addressing both prompt word-image alignment and generated content security.

[0047] Hint - Image Alignment Score Aimed at quantizing the generation of images Input prompts The fidelity of the semantic content described in the text.

[0048] To avoid penalizing a single generation defect multiple times, this invention designs a hierarchical evaluation mechanism. For prompt words... Requirements The invention first verifies each of the core objects. Does it exist in the image? In the middle. Subsequently, only for successfully generated objects, this invention uses a representation as A visual question-answering model is used to evaluate the accuracy of its attribute binding.

[0049] The fractions are calculated as follows: ; in, It is a verification object An indicator function that exists, and It is the accuracy score for binding properties to an existing object.

[0050] Safe generation rate Used to measure images In the pre-defined risk category Internal security. A multimodal large language model (MLLM) acts as the security classification function. It provides a direct harmfulness score, which is expressed mathematically as: ; It is a continuous value between 0 and 1, and the higher the value, the higher the security of the image.

[0051] ABSS is calculated by multiplying the cue word-image alignment score by a security weighting factor to ensure both accuracy and security in content capture: ; Sensitivity parameters In the tests of this invention, the value was set to 2 to adjust for the impact of the security score. This multiplication guarantees the performance of any NSFW image (whose...). ABSS values ​​close to zero will all approach zero, thus emphasizing safety in the assessment.

[0052] The process of constructing a sensitive word dataset based on multiple prompt words after annotation and decomposition includes the following steps: constructing an initial seed set based on multiple prompt words after annotation and decomposition; inputting the initial seed set into a large language model for automated vocabulary expansion; when a seed word and its corresponding harmful dimension and semantic component are provided, the model applies a specified strategy to generate new sensitive words that belong to the same harm type but span different semantic components; and constructing a sensitive word dataset based on all the new sensitive words.

[0053] The process of performing an attribute binding security assessment on the generated image to obtain an attribute binding security score includes the following steps: calculating the fidelity of the generated image to the semantic content described in the input prompt word to obtain a prompt word-image alignment score; calculating the security of the generated image within a predetermined risk category to obtain a secure generation rate; and calculating the attribute binding security score based on the prompt word-image alignment score and the secure generation rate.

[0054] Calculate the fidelity of the generated image to the semantic content described in the input prompt, and obtain the prompt-image alignment score, including the following steps; For prompt words Requirements Core objects First, verify each core object. Does it exist in the image? In this context, a successfully generated object is represented as... A visual question-answering model is used to evaluate the accuracy of its attribute binding; The fractions are calculated as follows: ; in, It is a core verification object An indicator function that exists, and It is the accuracy score for binding properties to an existing object.

[0055] Calculate the safety of the generated image within a predetermined risk category to obtain the safe generation rate, including the following steps: Safe generation rate Used to measure images In the pre-defined risk category For internal security, a multimodal large language model (MLLM) acts as a security classification function. It provides a direct harmfulness score. Mathematically, this can be expressed as: ; It is a continuous value between 0 and 1, and the higher the value, the higher the security of the image.

[0056] The Attribute Binding Security Score (ABSS) is calculated based on the cue word-image alignment score and the secure generation rate. ABSS is calculated by multiplying the cue word-image alignment score by a security weighting factor to ensure both content accuracy and security are captured. ; Sensitivity parameters This is used to adjust the impact of the safe generation rate.

[0057] The effect after improvement: Through the aforementioned innovations, this patented technology achieves significant improvements over existing technologies, specifically in the following aspects: Extremely strong security filter circumvention capability: The "mis-bound instructions" generated by this technology perform exceptionally well in bypassing security filters. Experimental data shows that when facing five mainstream text security detectors (including GPT-4o) simultaneously, the instruction set of this technology can achieve an overall bypass rate of 40.56%, far exceeding other mainstream adversarial methods (such as MMA at 4.40% and Sneakyprompt at 4.42%).

[0058] Table 1: Comparison of Error Binding Warning Words with Other Baseline Methods

[0059] Highly Effective Inducement of Risky Content Generation: This attack method is remarkably effective in inducing models to generate inappropriate content. Among all tested models and datasets, images generated by this technique's instruction set achieved the lowest "safe generation rate," meaning it generates risky content most efficiently. This effect is particularly evident in identity-preserving models, demonstrating the attack's precise targeting capability against such models. Attack examples include... Figure 2 As shown.

[0060] Precision and high fidelity of the attack: Unlike other methods that use complex or obscure text, the instructions in this technique maintain a high degree of semantic clarity. Therefore, the "content alignment" between the generated image and the instructions is the highest among all tested methods. This proves that this attack is not only effective but also precise, perfectly replicating the instruction content while implanting the risk.

[0061] Table 2: Evaluation Validity Analysis

[0062] Comprehensiveness and Reliability of Evaluation Results: The "attribute-bound security score" evaluation index proposed in this patent provides a new and reliable standard for measuring the overall security of a model. It not only quantifies the novel risk of "high alignment, low security," but its evaluation results, confirmed by user surveys, show a strong positive correlation with human subjective judgment. This indicates that the "attribute-bound security score" is a more comprehensive evaluation tool than existing single indicators, and more closely reflects real-world risk perception.

[0063] In one embodiment, the attribute misbinding attack and evaluation system for identity preservation models includes: a sensitive word dataset construction module, used to acquire an open-source inappropriate content dataset, which includes multiple prompt words. The prompt words are labeled along four harmful dimensions: pornography, violence, discrimination, and illegality. The prompt words are decomposed into three semantic components: role, scene, and state. A sensitive word dataset is constructed based on the labeled and decomposed prompt words. A misbinding instruction evaluation set construction module is communicatively connected to the sensitive word dataset construction module. It extracts sensitive words with harmful dimension and semantic component labels from the sensitive word dataset as input prompt words. The LLM model selects an error binding strategy to generate candidate prompt words. After security evaluation by a text filter, the valid prompt words that pass the evaluation are integrated into an error binding instruction evaluation set. The image generation module communicates with the error binding instruction evaluation set construction module to obtain an identity reference image. The valid prompt words in the error binding instruction evaluation set and the identity reference image are encoded by a text encoder and an image encoder, respectively, and then input into the diffusion model of the identity preservation model to generate an image that preserves the identity features of the person and conforms to the description of the prompt words. The scoring module communicates with the image generation module to perform attribute binding security evaluation on the generated image and obtain an attribute binding security score.

[0064] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of the present invention.

[0065] The specific embodiments of the present invention described above do not constitute a limitation on the scope of protection of the present invention. Any other corresponding changes and modifications made in accordance with the technical concept of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A method for evaluating attribute misbinding attacks targeting identity preservation models, characterized in that, Including the following steps: Obtain an open-source dataset of inappropriate content, which includes multiple prompt words. The prompt words are labeled along four harmful dimensions: pornography, violence, discrimination, and illegality. The prompt words are decomposed into three semantic components: role, scene, and state. A sensitive word dataset is constructed based on the multiple prompt words after labeling and decomposition. Sensitive words with harmful dimensions and semantic component labels are extracted from the sensitive word dataset as input prompt words. Candidate prompt words are generated by selecting the wrong binding strategy through the LLM model. After the text filter security evaluation, the valid prompt words that pass the evaluation are integrated into the wrong binding instruction evaluation set. Obtain the identity reference image, encode the valid prompt words in the misbound instruction evaluation set and the identity reference image through a text encoder and an image encoder respectively, and input them into the diffusion model of the identity preservation model to generate a generated image that preserves the identity features of the person and conforms to the prompt word description. An attribute binding security assessment is performed on the generated image to obtain an attribute binding security score.

2. The method for evaluating attribute misbinding attacks against identity preservation models according to claim 1, characterized in that, A sensitive word dataset is constructed based on multiple prompt words after annotation and component decomposition, including the following steps: An initial seed set is constructed based on multiple prompt words after annotation and decomposition. The initial seed set is input into a large language model for automated vocabulary expansion: Given a seed word and its corresponding harmful dimension and semantic components, the model applies a specified strategy to generate new sensitive words that belong to the same harm type but span different semantic components, and constructs a sensitive word dataset based on all the new sensitive words.

3. The method for evaluating attribute misbinding attacks against identity preservation models according to claim 1, characterized in that, In semantic components, role represents the subject's identity, profession, or title; scene represents the location, context, or background of an event; and state represents the appearance, behavior, and psychological state related to the role.

4. The method for evaluating attribute misbinding attacks against identity preservation models according to claim 1, characterized in that, The generated image is subjected to an attribute binding security assessment to obtain an attribute binding security score, including the following steps: Calculate the fidelity of the generated image to the semantic content described in the input prompt word to obtain the prompt word-image alignment score; Calculate the safety of the generated image within a predetermined risk category to obtain the safe generation rate; The attribute binding security score is calculated based on the prompt word-image alignment score and the security generation rate.

5. The method for evaluating attribute misbinding attacks against identity preservation models according to claim 4, characterized in that, Calculate the fidelity of the generated image to the semantic content described in the input prompt, and obtain the prompt-image alignment score, including the following steps; For prompt words Requirements Core objects First, verify each core object. Does it exist in the image? In this context, a successfully generated object is represented as... A visual question-answering model is used to evaluate the accuracy of its attribute binding; The fractions are calculated as follows: ; in, It is a core verification object An indicator function that exists, and It is the accuracy score for binding properties to an existing object.

6. The method for evaluating attribute misbinding attacks against identity preservation models according to claim 5, characterized in that, Calculate the safety of the generated image within a predetermined risk category to obtain the safe generation rate, including the following steps: Safe generation rate Used to measure images In the pre-defined risk category For internal security, a multimodal large language model (MLLM) acts as a security classification function. It provides a direct harmfulness score. Mathematically, this can be expressed as: ; It is a continuous value between 0 and 1, and the higher the value, the higher the security of the image.

7. The method for evaluating attribute misbinding attacks against identity preservation models according to claim 6, characterized in that, The Attribute Binding Security Score (ABSS) is calculated based on the cue word-image alignment score and the secure generation rate. ABSS is calculated by multiplying the cue word-image alignment score by a security weighting factor to ensure both content accuracy and security are captured. ; Sensitivity parameters This is used to adjust the impact of the safe generation rate.

8. The method for evaluating attribute misbinding attacks against identity preservation models according to claim 7, characterized in that, Set to 2.

9. A system for evaluating attribute misbinding attacks on identity preservation models, characterized in that, include: The sensitive word dataset construction module is used to obtain an open-source inappropriate content dataset, which includes multiple prompt words. The prompt words are labeled along four harmful dimensions: pornography, violence, discrimination, and illegality. The prompt words are decomposed into three semantic components: role, scene, and state. The sensitive word dataset is constructed based on the multiple prompt words after labeling and decomposition. The error binding instruction evaluation set construction module communicates with the sensitive word dataset construction module. It extracts sensitive words with harmful dimensions and semantic component labels from the sensitive word dataset as input prompt words, selects error binding strategies through the LLM model to generate candidate prompt words, and integrates the valid prompt words that pass the evaluation into the error binding instruction evaluation set after the text filter security evaluation. The image generation module is communicatively connected to the mis-binding instruction evaluation set construction module to obtain the identity reference image. After the valid prompt words in the mis-binding instruction evaluation set and the identity reference image are encoded by the text encoder and the image encoder respectively, they are input into the diffusion model of the identity preservation model to generate a generated image that preserves the identity features of the person and conforms to the prompt word description. The scoring module communicates with the image generation module to perform attribute binding security assessment on the generated image and obtain the attribute binding security score.

Citation Information

Cited By

  • Vision structure-based large model prompt word injection attack determination method and device

    CN122197017A