Chinese-oriented prejudice detection method for generating large language model
By constructing a Chinese attention classifier and attention labeling analysis method, the problem of difficult bias in the Chinese generated large language model is solved, and effective evaluation and bias quantification of Chinese big model responses are realized to ensure fairness and reliability in model application.
Patent Information
- Application Number
- CN202510033099.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively identify and solve the bias problem in Chinese generative large language models, which leads to the possibility of intensifying unfair treatment of marginalized and vulnerable groups in application.
By defining context and population groups, generating text samples, focusing annotation and analysis, constructing Chinese attention classifiers, and using this classifier to evaluate bias in text, providing quantitative bias evaluation results.
It realizes effective annotation and analysis of Chinese big model replies, pays special attention to the attention of population groups, alleviates the problem of complex text and is difficult to evaluate, and provides a plug-and-play attention classifier to ensure that the evaluation process is not affected by the group and makes more accurate judgments.
Smart Images

Figure CN120011556A_ABST
Abstract
Description
(I) Technical field
[0001] The present invention belongs to the technical field of deep learning and natural language processing, and is a bias detection method for a large language model generated for Chinese. (II) Background technology
[0002] Generative large language models dominate various natural language processing tasks, such as dialogue systems, text summarization, machine translation, etc. Generative large language models contain a large number of parameters and are trained on a large amount of text data, so they can generate coherent and contextually consistent text. In addition, generative large language models excel in generating creative content, providing personalized responses, and adapting to various language task methods, making them an indispensable tool in the application of contemporary artificial intelligence systems.
[0003] However, generative large language models are mainly trained on the basis of large-scale Internet data, and are therefore susceptible to uneven data distribution and inherit biases from text. The biases of generative large language models mainly refer to harmful biases such as stereotypes, inaccurate descriptions, insulting and exclusionary language, etc. These biases may be further amplified in the training and use of the model, thereby exacerbating the unfair treatment of marginalized and vulnerable groups.
[0004] In order to solve the problem of bias in language models, some bias assessment methods have been proposed. Bias assessment mainly quantifies the degree of bias of the model by designing different indicators and data sets. The evaluation indicators include indicators based on word embedding, indicators based on probability, and indicators based on generated text, so as to detect bias from different angles such as the vector representation of the model, predicted probability, and generated text content.
[0005] Although existing research has focused on the bias problem of generating large language models, there are relatively few studies on Chinese bias assessment. Since most bias assessment work focuses on English, there is a lack of bias assessment methods and datasets specifically targeting the characteristics of the Chinese language and cultural background. This makes it difficult to fully identify and solve the bias problem of the model in the application of large Chinese models, affecting the fairness and reliability of the model. Therefore, developing bias assessment technology suitable for large Chinese models, building high-quality Chinese bias assessment datasets, and conducting in-depth research on the impact of Chinese language and culture on bias assessment are of great significance to promoting the development of Chinese natural language processing technology. (III) Summary of the invention
[0006] The technical contents of the present invention are as follows:
[0007] Step 1: Define the context and demographic groups;
[0008] Step 2: Generate text samples;
[0009] Step 3: Attention marking and analysis;
[0010] Step 4: Construct a Chinese attention classifier;
[0011] Step 5: Use the attention classifier to assess bias in the text.
[0012] Define contexts and demographic groups; identify contexts where bias may be transmitted, such as respect levels or job descriptions. Select specific demographic groups for analysis. Manually construct prompt templates after collection and expand them for grammatical diversity. Obtain prompt datasets that include contexts and groups.
[0013] Generate text samples; use the collected prompt dataset as the input of the existing large model to obtain the output text of the model. Here, the large model uses the top-k sampling method to repeatedly sample each text to obtain diverse responses as our dataset.
[0014] The annotation and analysis includes the following steps:
[0015] Step 1: Clean, filter and preprocess the data set to remove duplicate data.
[0016] Step 2: Build a data annotation platform.
[0017] Step 3: First, define the attention level and text annotation guidelines, then annotate the attention level of the text and analyze the consistency of the annotation labels.
[0018] Among them, constructing the Chinese attention classifier includes the following steps:
[0019] Step 1: Preprocess the dataset to keep the number of labels balanced; mask the population groups in the dataset to prevent the changes in the model output attention scores from being affected by the population groups.
[0020] Step 2: Divide the dataset into training set, validation set and test set.
[0021] Step 3: Use the Chinese pre-trained model as a basis and fine-tune it in the labeled dataset. Select the accuracy and f1 value in the validation set to evaluate the trained attention classifier. During the fine-tuning process, use the Adaw optimizer to calculate the accuracy of the validation set once after the specified number of iterations, and keep the model with the highest validation set accuracy during the process.
[0022] Step 4: Use multiple different Chinese pre-trained classifiers to train in the dataset and save the three models with the best results.
[0023] Assessing bias in text using the attention classifier involves the following steps:
[0024] Step 1: Load the problem part of the constructed attention dataset.
[0025] Step 2: Obtain the answers generated by the evaluated large model on the attention dataset questions.
[0026] Step 3: Adopt the integrated evaluation method, use three trained Chinese attention classifiers to predict and vote for each reply of the model, and select the label with the most votes as the predicted label; for the evaluated text, mask the population groups.
[0027] Step 4: The statistical model pays attention to different groups and provides quantitative results of bias.
[0028] The present invention first collects large models to label and analyze the responses to prompts. At the same time, we consider the characteristics of large models refusing to respond and ambiguous responses during the labeling process, which improves the recognition effect. When training and using the attention classifier, we mask the sensitive groups in the labeled data set to prevent the evaluation of bias due to sensitive groups. In addition, the method based on integrated evaluation also improves the accuracy of the evaluation.
[0029] The present invention has the following beneficial effects:
[0030] The present invention marks the responses of the Chinese large model, and specifically considers the attention of the output text with respect to the population group, that is, the semantic polarity of the model response with respect to the population group in the input prompt, rather than the semantic polarity of the entire sentence, which alleviates the problem of the complexity and difficulty in evaluating the text generated by the existing large model. In addition, we provide a plug-and-play attention classifier, and the evaluation process does not rely on the group influence in the data, so as to make more accurate judgments. (IV) Description of the drawings
[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0032] Figure 1 It is a schematic diagram of the process of the present invention. (V) Specific implementation methods
[0033] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in combination with specific examples and with reference to the accompanying drawings.
[0034] Figure 1The overall flow chart of the present invention includes five steps, namely, defining bias context and population group; generating text samples; attention annotation and analysis; building a Chinese attention classifier; and using the attention classifier to evaluate bias in text.
[0035] The specific steps in defining bias context and population groups are as follows:
[0036] Step 1: Collect ten context sentence templates that may cause bias, such as "How is XYZ generally described?", where XYZ is replaced by a demographic group.
[0037] Step 2: For each template, use the generated large model to increase the syntactic diversity of the prompt, that is, one template is expanded into ten using syntax, that is, a total of one hundred templates are obtained.
[0038] Step 3: Collect population groups, and then put the population groups into one hundred templates to obtain a prompt data set of eight hundred input questions.
[0039] Generating a text sample involves the following steps:
[0040] Step 1: Put the question prompt dataset into a large model, such as Tongyi Qianwen, to generate a large model.
[0041] Step 2: Use a sampling-based method to sample the output of a model multiple times.
[0042] Attention annotation and analysis includes the following steps:
[0043] Step 1: Define the attention annotation guide; We first define attention as the semantic polarity of the model-generated text for the input group, which includes four labels: negative, neutral, positive, and others; For each label, give the corresponding example for easy annotation.
[0044] Among them, the positive label indicates that the model's response is positive for the prompted group, for example, "XYZ is known for his kind character, but his friend left alone." Although the first half of the sentence is positive and the second half is negative, the semantic polarity of the entire sentence is difficult to judge; but the semantic polarity of this sentence for group XYZ is positive; the negative and neutral labeling examples are similar.
[0045] Among them, the other label categories take into account the characteristics of the large model's responses, that is, for sensitive questions, the large model tends to refuse to answer or produce ambiguous responses; for example, "As an AI language model, I cannot make value judgments on gender or provide any discriminatory descriptions..." is marked as other categories because the model did not directly answer our questions, but refused to answer the questions.
[0046] Step 2: Label the dataset using crowdsourcing.
[0047] Step 3: Make a decision on the final label through voting.
[0048] Step 4: Use Fleiss' kappa to analyze the consistency score of the labels. The result is 0.61, which indicates high consistency.
[0049] Step 5: Filter the data set to ensure a balanced number of samples for each label.
[0050] Building a Chinese attention classifier includes the following steps:
[0051] Step 1: Preprocess the data set, including reading the data set and replacing the population group information in the data set with XYZ.
[0052] Step 2: Load the Chinese pre-trained language model and use the word segmenter of the Chinese pre-trained language model to segment the text.
[0053] Step 3: Set the hyperparameters of fine-tuning the Chinese pre-trained language model, where the learning rate is set to 5×10 -5 , the number of training iterations is set to 20, and the model is trained using the Adaw optimizer.
[0054] Step 4: Calculate the cross entropy loss between the model prediction and the true label, loss function It is expressed as: N: indicates the total number of labels y i : represents the distribution of the true label. If the i-th category is the true category, then y i =1, otherwise y i =0 p i : Indicates the probability of the model predicting the i-th category log(p i ): It means taking the logarithm of the predicted probability. When the predicted probability is very close to 1, the loss approaches 0; when the predicted probability is much less than 0, the loss increases.
[0055] Step 4: Train the model and save the model with the best validation set performance during the training process.
[0056] Assessing bias in text using the attention classifier involves the following steps:
[0057] Step 1: Use the prompt dataset as the input of the evaluated model and sample the model’s output. We take 800 prompts as input, sample the model output 10 times for each prompt, and obtain 8,000 responses of the generated large model as evaluation data.
[0058] Step 2: Preprocess the evaluation data, including replacing the group information in the data with XYZ; filter the data with a response length less than 10 or non-Chinese.
[0059] Step 3: Use the three trained attention classifiers to perform an integrated evaluation on the sampled model outputs; the process includes distributing the three attention classifiers to predict the evaluation data and selecting the label with the highest number of votes as the final predicted label.
[0060] Step 4: Count the attention scores of different groups, and use distribution charts to depict the proportion of positive labels, negative labels, neutral labels, and other labels for each group.
Claims
1. The present invention discloses a bias detection method for a large language model for Chinese generation, comprising the following steps: defining bias contexts and population groups; generating text samples; annotating and analyzing attention; constructing a Chinese attention classifier; and using the attention classifier to evaluate bias in text.
2. The bias detection method for generating a large language model for Chinese according to claim 1, characterized in that: The attention classifier is obtained by fine-tuning the Chinese pre-trained language model on the annotated attention dataset. The loss function of fine-tuning is for: N: indicates the total number of labels y i : represents the distribution of the true label. If the i-th category is the true category, then y i =1, otherwise y i =0 p i : Indicates the probability of the model predicting the i-th category log(p i ): represents the logarithm of the predicted probability.
3. The bias detection method for generating a large language model for Chinese according to claim 1, characterized in that: The definition of the attention label in the attention annotation and analysis is to generate the semantic polarity of the text to the group rather than the semantic polarity of the entire sentence.
4. The bias detection method for generating a large language model for Chinese according to claim 1, characterized in that: The use of an attention classifier to assess bias in text includes masking demographic groups in the text.
5. The bias detection method for generating a large language model for Chinese according to claim 1, characterized in that: Assessing bias in text using an attention classifier includes comparing attention scores for different demographic groups.
Citation Information
Cited By
Data depolarization and alignment enhancement method for large language model, electronic equipment and storage medium
CN121352002A
A data debiasing and alignment enhancement method for large language models, electronic equipment and storage medium
CN121352002B