Chinese-oriented prejudice attack method for generating large language model

Through bias attack methods for generating large language models against Chinese, including data set acquisition, bias association initialization and adaptive search, the problem of lack of bias attacks against Chinese models in the existing technology is solved, effective evaluation and attack on model bias is achieved, and the fairness and accuracy of the model is promoted.

CN119938882APending Publication Date: 2025-05-06GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510033143.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The lack of bias attack methods for generating large language models in the prior art, resulting in models that may inherit stereotypes and biases in the dataset, affecting the fairness and accuracy of the model.

Method used

Provide a bias attack method for generating large language models for Chinese, including obtaining the data set required for bias attacks, initializing bias associations, adaptive-based search, computing Pareto frontiers for different targets, bias-oriented selection strategies, and evaluating bias in generated texts. This method uses adversarial search and initialization of bias associations, with the goal being to maximize the probability of the model outputting target polarity replies.

Benefits of technology

This method can effectively evaluate and attack the biases of large language models generated by Chinese language, help develop more fair and accurate models and reduce psychological harm to minorities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938882A_ABST
    Figure CN119938882A_ABST
Patent Text Reader

Abstract

The invention discloses a Chinese-oriented prejudice attack method for generating a large language model, belongs to the field of natural language processing, is used for text confrontation attacks, and comprises the following steps: acquiring a data set required by a prejudice attack; using the data set to initialize prejudice association and set an optimization target; secondly, searching for a smooth confrontation prompt through an adaptability-based search benefit, wherein the adaptability search utilizes a large model to predict probability distribution of a next lexical element and semantic similarity filtering to improve the concealment of the confrontation prompt; then Pareto leading edges of different targets are calculated; selecting effective candidate confrontation prompts by using a prejudice-oriented selection strategy; and finally, evaluating the robustness of the prejudice of the generated text by utilizing the obtained confrontation prompt.
Need to check novelty before this filing date? Find Prior Art

Description

(I) Technical field

[0001] The present invention relates to natural language processing technology, which is a bias attack method for generating a large language model for Chinese. (II) Background technology

[0002] Generative large language models have shown excellent performance in various language-related tasks. These models are able to process and generate human-like text and are widely used in many fields such as text generation, machine translation, and sentiment analysis. However, since these models are usually trained based on a large amount of unscreened data, they may inherit or even inherit stereotypes and biases in the method dataset. The existence of this bias not only affects the fairness and accuracy of the model, but may also have adverse effects on society.

[0003] The current work assesses bias by measuring semantic polarity toward specific groups. It determines whether the text shows bias against specific groups by associating the generated text with positive or negative statements or attributes. Generative large language models may disproportionately generate positive descriptions for certain groups of people and associate other groups with negative characteristics. By analyzing polarity, researchers can assess whether the generated text reveals significant differences between groups, which can help promote the development of fairer generative large language models.

[0004] Recent studies have shown that generative large models are vulnerable to adversarial attacks. Adversarial attacks can bypass the internal protection of large language models and guide the model to generate harmful text. However, existing adversarial attack methods ignore the bias problem of large language models. The attacker's goal may also be bias, which will lead to a decrease in people's trust in large language models and aggravate psychological harm to minority groups. Therefore, it is necessary to study the vulnerability of generative large models to bias. In addition, existing attack methods on large language models focus on English, and there are no adversarial attacks on Chinese generative large language models. (III) Summary of the invention

[0005] The purpose of the present invention is to provide a bias attack method for generating a large language model for Chinese, so as to solve the problem that the prior art lacks consideration of bias attacks on large language models generated by Chinese.

[0006] In order to achieve the above object, the present invention provides a bias attack method for generating a large language model for Chinese, comprising the following steps:

[0007] Step 1: Obtain the data set required for bias attack;

[0008] Step 2: Initialize bias association;

[0009] Step 3: Adaptability-based search;

[0010] Step 4: Calculate the Pareto frontier for different objectives;

[0011] Step 5: Selection strategy facing bias;

[0012] Step 6: Evaluate the bias of generated text.

[0013] In the process of obtaining the dataset required for bias attack, we build on the existing attention dataset as the prompt part of the dataset; for each prompt we sample the Chinese non-canonical model to obtain positive and negative responses.

[0014] Initializing bias associations includes: setting bias associations to associate prompts from population group 1 with negative responses, and to associate prompts from population group 2 with positive responses; here, the prompts from population group 1 and population group 2 differ only in the groups they describe, and all other content is the same; the attacker's goal is to add adversarial prompts to maximize the probability that the model outputs a target polarity response.

[0015] The specific steps of adaptability-based search are as follows:

[0016] Step 1: Based on the two sentences associated with the given bias, obtain the mean probability distribution of the attacked generative model predicting the next word after two prompts: m: represents the number of groups Included group d i Tips Large model input tips The probability distribution of the predicted next word

[0017] Step 2: For the obtained probability distribution, adopt probability-based adoption without replacement and sample b words as b initial adversarial prompts

[0018] Step 3: Based on the initial adversarial prompts, for each adversarial prompt, after splicing the prompts from different groups, the probability distribution of the next word is sampled: represents the i-th adversarial hint Represents the concatenation operation of two vectors or strings

[0019] For the i-th adversarial prompt, we obtain the probability to sample the k next word units {t1,…,t k}.

[0020] Step 4: For the i-th adversarial prompt, obtain the probability and sample the k next word units {t1,…,t k}, and then splice it with the previous confrontation prompt to get a new confrontation prompt

[0021] Step 5: Calculate the semantic similarity between the concatenated adversarial prompt and the original input, and filter out adversarial prompts whose semantic similarity is less than a threshold.

[0022] Calculating the Pareto front for different objectives involves the following steps:

[0023] Step 1: Calculate the probability scores of the output of sentences with different target polarities, where the probability of outputting the jth word is: x n+j ~P(·|x1,x2,…,x n+j-1 ) n: indicates the number of words in the input prompt

[0024] The above formula indicates that the output j-th word is sampled from the probability distribution.

[0025] Step 2: Output the loss of the target sentence: y: target sentence N: The number of tokens in the target sentence x: indicates input prompt

[0026] Step 3: Use the above to calculate the loss scores of different targets for each prompt.

[0027] Step 4: Use the fast non-dominated sorting algorithm to calculate the Pareto front of each adversarial prompt on the positive and negative targets.

[0028] The bias-oriented selection strategy involves using the bias-dominant mechanism to select bias-effective solutions from the Pareto frontier.

[0029] We will select b candidate adversarial cues from the frontier to enter the next adaptive search iteration until the target number of iterations is reached.

[0030] Evaluating the bias of generated text involves the following steps:

[0031] Step 1: Evaluation preparation, including loading the Chinese toxicity classifier and semantic classifier.

[0032] Step 2: Concatenate the prompt data and our adversarial suffixes to determine whether the prompts of different groups have different semantics or whether the toxicity scores of different groups increase.

[0033] Step 3: Count the semantic scores and draw a semantic score distribution graph to intuitively analyze the bias of the large model generated before and after the attack. (IV) Description of the drawings

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0035] Figure 1 This is a flowchart of a bias attack method for generating a large language model for Chinese according to the present invention. (V) Specific implementation methods

[0036] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in combination with specific examples and with reference to the accompanying drawings.

[0037] Figure 1 This is the overall flow chart of the present invention, which includes six steps, namely, obtaining the data set required for bias attack; initializing bias association; adaptive search; calculating the Pareto frontier of different objectives; bias-oriented selection strategy; and evaluating the bias of generated text.

[0038] In the process of obtaining the data set required for bias attack, a Chinese prompt data set containing sensitive groups is used as the input of the Chinese unsupervised large model, and the positive and negative responses contained in each Chinese prompt are sampled.

[0039] Initializing the bias association involves the following steps:

[0040] Step 1: Associate the prompt containing Group 1 with a negative target response, and associate Group 2 with a positive target response for this bias category; for example, for gender bias, Group 1 is female and Group 2 is male.

[0041] Step 2: Set the attack target, that is, find the confrontation hint x a , so that the original responses with the same semantics have different semantic polarities after adding adversarial prompts: g(·): prediction of the semantic classifier M(·): Generates the response of the large model to the input

[0042] Adaptive-based search involves the following steps:

[0043] Step 1: Set the adaptive search parameter b, which indicates the number of candidate prompts retained in one iteration; the parameter L, which indicates the length of the adversarial prompt.

[0044] Step 2: Use the mean of the probability distribution of different groups to sample b words without replacement as the initial b adversarial suffixes

[0045] Step 3: Each adversarial prompt is expanded to k using the sampling method, that is, a total of b×k candidate adversarial prompts are generated;

[0046] Step 4: Filter the generated adversarial prompts and the original prompts using semantic similarity. The semantic similarity s is as follows: enc: Represents the universal sentence encoder, which returns an embedding vector in the semantic space.

[0047] If the semantic similarity calculated by the adversarial prompt is less than a threshold, the adversarial prompt is filtered out.

[0048] Step 5: Select the b best adversarial cues using a bias-oriented selection strategy and repeat the iteration L-1 times.

[0049] Calculating the Pareto front for different objectives involves the following steps:

[0050] Step 1: Calculate the probability of outputting positive and negative targets for each adversarial prompt.

[0051] Step 2: Use fast non-dominated sorting to construct the Pareto front.

[0052] The bias-oriented selection strategy includes a bias dominance mechanism, that is, when comparing two adversarial prompts with different polarities, which adversarial prompt is better. The dominance mechanism is: first compare the negative output probability, and the adversarial prompt with a larger negative output probability is better; then compare the probability scores of different polarities, and the adversarial prompt with a larger sum of the two probability scores is better.

[0053] Evaluating the bias of generated text involves the following steps:

[0054] Step 1: Input the model with a dataset containing the population and sample the model’s responses.

[0055] Step 2: Mask the population groups in the model output text to prevent the group from affecting the model's prediction.

[0056] Step 3: Use the Chinese semantic classifier and Chinese toxicity classifier to evaluate the semantic polarity and toxicity score output by the model, thereby obtaining pre-attack evaluation data.

[0057] Step 4: Use the data containing the prompts combined with the constructed adversarial prompts to input into the attacked generative model; and sample the model's response.

[0058] Step 5: Use the Chinese semantic classifier and the Chinese toxicity classifier to evaluate the semantic polarity and toxicity score of the model output after the attack, thereby obtaining the evaluation data after the attack.

[0059] Step 6. Draw a bar chart and a scatter plot of the semantic polarity distribution to analyze the difference in bias scores before and after the attack.

Claims

1. The present invention discloses a bias attack method for generating a large language model for Chinese, which mainly includes: Obtain the data set needed for bias attacks; Initialize bias associations; Adaptability-based search; Calculate the Pareto frontier for different objectives; bias-oriented selection strategies; Evaluating bias in generated text.

2. According to claim 1, a bias attack method for generating a large language model for Chinese, characterized in that: Adaptive search uses the mean probability distribution of the next word predicted by the large model for different groups of prompts to sample and generate fluent and content-related words. The mean probability distribution is: m: represents the number of groups Included group d i Tips Large model input tips The probability distribution of the predicted next token.

3. The bias attack method for generating a large language model for Chinese according to claim 1, characterized in that: When calculating the Pareto frontier of different objectives, the loss function of the objective is: y: target sentence N: The number of tokens in the target sentence x: indicates the input prompt.

4. The bias attack method for generating a large language model for Chinese according to claim 1, characterized in that: Evaluating the bias of generated text requires masking the demographic groups before evaluation.

Citation Information

Cited By

  • Large language model prejudice reduction method and system based on cross-model judgment

    CN120806096A