Method and system for bias mitigation of large language models across model adjudication
Patent Information
- Application Number
- CN202510943528.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-07-09
AI Technical Summary
[0007]本发明提供一种跨模型裁决的大语言模型偏见消减方法及系统,以解决当前偏见检测和消减方法通常基于大量数据集和度量标准对LLM进行微调,由于固有的不灵活性(固定的测试内容,例如文本)和潜在的规避性(基准本身可能已被用于训练模型),导致无法实现有效的偏见消除的问题
1)提出一种跨模型同行评审协议,其中多个LLM使用动态校准的度量迭代评估彼此的响应;
Smart Images

Figure CN120806096B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of large language model bias reduction technology, specifically to a method and system for cross-model adjudication of large language model bias reduction. Background Technology
[0002] Modern artificial intelligence systems, particularly LLMs, are increasingly integrated into critical social infrastructures, necessitating a rigorous review of their operational characteristics and potential impacts in this embodiment. These models now frequently inform or automate high-risk decisions affecting millions of people, spanning areas from employment screening to financial credit allocation. However, mounting evidence suggests a significant challenge: current LLMs often inherit, reflect, and potentially amplify human social biases present in their vast training datasets.
[0003] This phenomenon is not merely a technological flaw, but has far-reaching social implications. For example, commercial and open-source models exhibit statistically significant biases, such as a persistent preference for female characters in short stories, or a significantly higher frequency of associating African American vernacular English with negative or criminal situations than standard American English dialects. Such biases can cause real harm across various sectors. In employment, biased LLM models used for resume screening or job recommendation systems may lead to persistent discriminatory hiring practices based on gender, race, or age against specific demographic groups. In finance, models used for loan approval or credit scoring may lead to unfair outcomes and limit access to capital for minority groups. In healthcare, biased LLM models that inform medical diagnoses or treatment recommendations may exacerbate existing health disparities related to gender or race. Furthermore, in the criminal justice system, risk assessment tools using LLM models for parole or sentencing decisions may reinforce systemic biases, resulting in unfair outcomes for certain groups.
[0004] Despite the widespread recognition of this problem, current bias detection and reduction methods face substantial limitations, often providing only a partial view of the complex nature of LLM bias. Supervised fine-tuning, a common approach, heavily relies on manual annotation, which is not only costly but may also unintentionally embed implicit or explicit biases of the annotator into the model. Adversarial filtering techniques aim to remove sensitive attribute information but may impair model utility by stripping away valuable contextual or cultural nuances. Furthermore, post-processing correction methods applied after model training typically introduce significant computational overhead, increase inference latency, and primarily address symptoms rather than the root causes of bias embedded in the model's core representation.
[0005] The most fundamental problem is a critical weakness prevalent in current evaluation paradigms: the cyclical validation problem. In many cases, models are evaluated using metrics, benchmarks, or even simulated interactions that are similar to those used in their training or development. This self-referential evaluation cycle creates significant blind spots, particularly regarding biases stemming from shared architecture choices, pre-trained data, or fine-tuning objectives. For example, external audits have shown that well-known models like LLaMA2 perform significantly worse than results reported by their internal self-evaluation procedures.
[0006] This invention addresses a key question: Can LLMs collectively identify and correct biases that a single dataset or model cannot perceive on its own? It hypothesizes that, just as democratic deliberation can reveal truths that no single individual can grasp, structured collaboration between heterogeneous models can expose implicit biases through adversarial scrutiny. The crucial insight lies in architectural diversity—by coordinating different models to critique each other's outputs, this invention creates a correction mechanism where biases cancel each other out, allowing genuine consensus to emerge. Summary of the Invention
[0007] This invention provides a method and system for reducing bias in large language models across models, addressing the problem that current bias detection and reduction methods typically fine-tune LLMs based on large datasets and metrics, which fail to achieve effective bias elimination due to inherent inflexibility (fixed test content, such as text) and potential avoidance (the benchmark itself may have already been used to train the model).
[0008] According to a first aspect, one embodiment provides a method for reducing large language model bias across model adjudication, the method comprising: Given a set of multiple different LLM models, input the preset prompts into the different LLM models respectively, and obtain the response generated by each LLM model; Based on a composite bias metric, the bias level of the responses generated by all other LLM models in the peer group is evaluated using each LLM model, and a bias assessment score is obtained. Based on the obtained bias assessment score, the arrival score of each response is calculated using an improved arrival counting mechanism, and the response with the highest arrival score is selected as the consensus target response with the least collective bias. Based on the obtained low-biased consensus target response, the parameters of each participating LLM model are fine-tuned.
[0009] Furthermore, given a set of multiple different LLM models, preset prompts are input into different LLM models respectively, and the responses generated by each LLM model are obtained, specifically including: A set of N different LLM models is represented as M = {M1, M2, ..., M}.N}; For a given input cue q, each model M m ∈M generates a response r m =(y1,y2,...,y L ), where L is the sequence length.
[0010] Furthermore, given a set of multiple different LLM models, preset prompts are input into different LLM models respectively, and the responses generated by each LLM model are obtained, specifically including: The generated responses are randomly sampled to encourage diversity while maintaining coherence, specifically using kernel sampling combined with temperature scaling: ; Where z m,t,yt Model M m At time step t, the term y is... t The generated logits, where τ is the temperature parameter.
[0011] Furthermore, based on a composite bias metric, the bias level of the responses generated by all other LLM models in the peer group was evaluated using each LLM model, specifically including: The composite bias metric provides the LLM model with a set of guiding principles or conceptual dimensions presented through prompts; Model M j The original response r is evaluated based on different dimensions of the established composite bias metric, including relevance, vocabulary, and context. k The bias level, k∈{1,...,N}, j ≠ k, and N is the number of LLM models.
[0012] Furthermore, based on a composite bias metric, the bias level of the responses generated by all other LLM models in the peer group was evaluated using each LLM model, specifically including: The relevance bias dimension is used to indicate the model M. j Evaluation response r k Whether stereotyped associations are established implicitly or explicitly between demographic groups and specific attributes, roles, or concepts that reflect social biases, the corresponding dimension aims to capture the characteristics of associated biases measured in the embedding space, by requiring LLM to identify subtle semantic links; Lexical bias dimension used in cue indication model M j Identify response r k Does the language contain overtly biased, stereotypical, derogatory, or harmful language? This includes examining words or phrases associated with negative stereotypes or unfair generalizations, and using LLM's linguistic knowledge to detect relevant problem terms in context. Contextual bias dimension is used to prompt the guidance model M j Evaluation response r k The overall narrative, emotion, and implied meaning include assessing whether the text subtly reinforces stereotypes, presents different groups in an unbalanced or unfair manner, or normalizes biased viewpoints, even without explicitly using biased terms; the corresponding dimensions aim to capture features that measure deviations from neutral or equitable distributions in a broader context.
[0013] Furthermore, when evaluating the bias level of responses generated by all other LLM models in the peer group based on a composite bias metric, the method specifically includes: Guiding revisions: The guidelines explicitly require model M to be correct. j Generate response r k The revised version is represented as The aim is to minimize biases identified in the evaluation process while preserving the original intent and core information. Counterfact generation: Hints require model M j Generate a modified version by systematically changing sensitive attributes. Alternative version .
[0014] Furthermore, a bias assessment score is obtained, specifically including: Evaluation scores jk From M j Comprehensive analysis: ; in, Representation Model M j Assessment Response The process, Representation Model Evaluating the response based on different dimensions of the composite bias metric The numerical score output at that time. It is a model Optional confidence weights, This represents a small amount of Gaussian noise added to the fraction, where σ is a hyperparameter. The output of this stage is an N×N fractional matrix S, which reflects the original response {r1,...,r...} N Collective judgments on the level of bias, resulting in revisions and counterfactual As a supplementary data collection.
[0015] Furthermore, based on the obtained bias assessment score, an improved arrival score is used to calculate the arrival score of each response. The response with the highest arrival score is selected as the consensus target response with the least collective bias, specifically including: For each response r k Its wavelet fraction b k It is calculated based on pairwise comparisons in the ranking of each evaluator model. Specifically, for each evaluator model M... j According to the bias assessment score s ji For all responses {r i} i≠j Rank, respond r k From the evaluator model M j The points earned are based on M j Bias assessment scores ji How many other responses did it beat? i Determined: ; Where I(·) is the indicator function; b k The calculation formula summarizes the response r k In all evaluator models M except for self-evaluation j The "wins" in pairwise comparisons; With the highest arrival fraction b k response Selected as the consensus target response: ; This serves as a low-biased target response extracted from the collective judgment of the model group for a given cue q.
[0016] Furthermore, based on the obtained low-biased consensus target response, the parameters of each participating LLM model are fine-tuned, specifically including: The LoRA mechanism is used for efficient fine-tuning of model parameters: for each model M m LoRA introduces a low-rank decomposition matrix B for a specific weight matrix W0 in the original model. m and A m Model updates are restricted to ; Model M m The fine-tuning objective is to minimize a distillation loss, i.e., the model's predicted distribution of response terms is close to the target consensus response r. ∗ Cross-entropy loss between: ; in It's a hint and its corresponding consensus target response right, It is the standard cross-entropy loss generated from the sequence, while the original model weights W0 remain frozen.
[0017] According to a second aspect, one embodiment provides a large language model bias reduction system for cross-model adjudication, the system comprising: The response generation module is used to take a set of multiple different LLM models, input the preset prompts into the different LLM models respectively, and obtain the response generated by each LLM model. The cross-model assessment module is used to evaluate the level of bias in the responses generated by all other LLM models in the peer group based on a composite bias metric, and to obtain a bias assessment score. The consensus extraction module is used to calculate the arrival score of each response based on the obtained bias evaluation score and using an improved arrival counting mechanism. The response with the highest arrival score is selected as the consensus target response with the least bias in collective consensus. The model fine-tuning module is used to fine-tune the parameters of each participating LLM model based on the obtained low-bias consensus target response.
[0018] This invention provides a method and system for reducing bias in large language models across models, which has the following beneficial effects: 1) Propose a cross-model peer review protocol in which multiple LLMs use dynamically calibrated metrics to iteratively evaluate each other's responses; 2) A consensus extraction algorithm is proposed, which identifies the most unbiased output through improved wave count; 3) Efficient parameter adaptation: Fine-tuning each model with these extracted unbiased consensuses, adding only 0.3% of additional parameters.
[0019] The cross-model adjudication framework of this invention requires no manual annotation, preserves cultural specificity, and operates as a universal bias “perspective” applicable to any LLM architecture. Attached Figure Description
[0020] Figure 1 A flowchart illustrating a method for reducing bias in large language models across models, as provided in one embodiment of the present invention; Figure 2 This is a schematic diagram of the logical structure of a large language model bias reduction system for cross-model adjudication, provided as an embodiment of the present invention. Detailed Implementation
[0021] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings. Similar elements in different embodiments are referred to by associated similar element reference numerals. In the following embodiments, many details are described to facilitate a better understanding of the invention. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other elements, materials, or methods. In some cases, certain operations related to the present invention are not shown or described in the specification. This is to avoid obscuring the core parts of the invention with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; they can fully understand the related operations based on the description in the specification and general technical knowledge in the art.
[0022] Furthermore, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can be rearranged or adjusted in a manner obvious to those skilled in the art. Therefore, the various orders in the specification and drawings are only for the clear description of a particular embodiment and do not imply a necessary order, unless otherwise stated that a particular order must be followed.
[0023] The first embodiment of this invention provides a method for reducing large language model bias in cross-model adjudication. It utilizes Lissajous trajectories to visualize current signals, captures series arc fault characteristics under different load conditions, determines the power mode type based on observed current harmonic components, relies on abnormal current pulses for initiation, and finally completes series arc fault detection based on image recognition technology. The following is a related description... Figure 1 A detailed explanation will be provided.
[0024] like Figure 1 As shown, in step S100, given a set of multiple different LLM models, preset prompts are input into different LLM models respectively, and the response generated by each LLM model is obtained.
[0025] Response generation: per model M m ∈M independently generates a response r for the input prompt q. m .
[0026] The above steps specifically include: S110, for a given input cue q, each model M m Generate a response r m =(y1,y2,...,y L ), where L is the sequence length.
[0027] S120, generation employs random sampling to encourage diversity while maintaining consistency. Specifically, this embodiment uses kernel sampling combined with temperature scaling: ; Where z m,t,yt Model M m At time step t, the term y is... t The generated logits, where τ is the temperature parameter, are used. Kernel sampling is applied to a cumulative probability threshold p=0.9, retaining the smallest set of tokens with a cumulative probability exceeding p. This ensures a diverse yet reasonable continuation of the relevance to the cue q.
[0028] like Figure 1 As shown, in step S200, based on the composite bias metric, the bias level of the responses generated by all other LLM models in the peer group is evaluated using each LLM model, and a bias assessment score is obtained.
[0029] Cross-model evaluation: for each model M j Evaluate all other models in the peer group {r k} k≠j The generated responses. This evaluation utilizes the Composite Bias Metric (CBM) proposed in this embodiment to evaluate each response r. k The level of bias. Crucially, M j Without evaluating one's own response r j To prevent self-assessment bias.
[0030] The above steps specifically include: This phase implemented the core peer review mechanism, which includes not only evaluation but also guidance revisions and counterfactual generation. For each response r k (k∈{1,...,N}), each other model M j (j=k) Execute a multi-step task prompted by specific instructions: S210, Bias Assessment: Model M j First, the original response r is evaluated based on the dimensions of the Composite Bias Measurement (CBM) (relevance, vocabulary, context). k The level of bias.
[0031] Composite Bias Measurement (CBM): The core of cross-model evaluation lies in guiding each evaluator model M j Evaluate peer response r k The level of bias. To provide a structured and multidimensional basis for this assessment, this embodiment defines CBM as a set of guidelines or dimensions of LLM presented to the evaluator through its prompts, rather than externally calculated scores. Specifically, M jBased on three key dimensions of potential bias (relevance, vocabulary, and context), r k The overall assessment is as follows: 1) Association bias dimension (semantic stereotype): used to prompt and indicate model M j Evaluation response r k Whether stereotypical associations are established implicitly or explicitly between demographic groups (e.g., gender, race, age) and specific attributes, roles, or concepts that reflect social biases. This dimension aims to capture the characteristics of association biases that are typically measured in the embedding space, which is achieved by requiring the LLM to identify subtle semantic links.
[0032] 2) Lexical bias dimension (explicit language): used for cueing and indicating model M j Identify response r k Does the language contain overtly biased, stereotypical, derogatory, or harmful language? This includes examining words or phrases associated with negative stereotypes or unfair generalizations and using the LLM's linguistic knowledge to detect such problematic terms in context.
[0033] 3) Contextual bias dimension (implicit bias and sentiment): used to prompt and guide model M j Evaluation response r k The overall narrative, emotion, and implied meaning. This includes assessing whether the text subtly reinforces stereotypes, presents different groups in an unbalanced or unfair manner, or normalizes biased viewpoints, even without explicitly using biased terms. This dimension aims to capture features that measure deviations from a neutral or equitable distribution in a broader context.
[0034] Provided to M j Assessment prompts can take many forms, such as: • After presenting the CBM dimension as the standard, a single, overall bias score is required (e.g., a range of 1 to 10, where 1 indicates the least bias).
[0035] • Requires separate scores for each dimension, then by M j Aggregation is performed either in subsequent steps or externally (e.g., using λ). E ,λ L ,λ C (Weighted average) to calculate the final score s jk .
[0036] In this implementation, the overall scoring method is mainly used to instruct the evaluator M. j Provide a representative r k A single score for the overall level of bias, derived from its assessment of these CBM dimensions.
[0037] Therefore, the bias assessment scores calculated later Representative model M j When prompted to evaluate r according to the CBM standard k The output is a numerical score. It is an LLM-generated assessment designed to capture the multifaceted nature of bias, rather than a direct calculation of mathematical formulas. This approach fully leverages the contextual understanding and reasoning capabilities of LLMs for peer review, directly aligning with the core concepts of cross-model adjudication.
[0038] S220, Amendment / Revision: Explicitly requires M j Generate r k The revised version is represented as The aim is to minimize bias identified during the evaluation process while preserving the original intent and core information. For example, an instruction might be: "Please rewrite the following text to eliminate any detected bias while remaining faithful to the original topic: 'r k '。 S230, Counterfactual Generation: Hints may also require M j Ideally, revised text is generated by systematically altering sensitive attributes (e.g., replacing "man" with "woman," "doctor" with "nurse" for gender analysis, or changing names / pronouns associated with different races or ages). An alternative version. This helps to detect fairness and consistency across population groups. These will be represented as... For example, the instruction: "Now, generate the revised text corresponding to the version with gender='female', age='elderly', race='Asian', ... Therefore, the evaluation score s jk From M j Comprehensive analysis: ; in Representation Model M j Evaluation response r k The process. In fact, this involves prompting M. j Scoring based on CBM dimensional instructions. k . The evaluator M represents j Calculated or estimated response r k The CBM score. This implicitly involves M. j Evaluate r in the embedding, lexical, and contextual dimensions of CBM k w j It is the evaluator M j Optional confidence weights. Initially uniform (w j =1 / N), which can be determined according to M jThe consensus is updated dynamically based on historical consistency, although static weights are used for simplicity in the main experiment. This represents a small amount of Gaussian noise added to the scores. This serves as a stochastic regularization technique, potentially improving robustness and preventing degenerate evaluation strategies that might be tacitly agreed upon by the models. σ is a small hyperparameter. The output at this stage is still an N×N score matrix S (ignoring s). jj It reflects the original response {r1,...,r} N Collective judgment on the level of bias. Generated revisions. and counterfactual As a supplementary data collection.
[0039] like Figure 1 As shown, in step S300, based on the obtained bias assessment score, the arrival score of each response is calculated using an improved arrival counting mechanism, and the response with the highest arrival score is selected as the consensus target response with the least collective bias.
[0040] Consensus Refinement: The evaluation scores obtained in step S200 are aggregated using an improved waveguide counting mechanism to identify the consensus target response r. ∗ The response r ∗ This represents the output with the "least bias" in response to the collective consensus of cue q, and serves as the target for subsequent fine-tuning.
[0041] The above steps specifically include: S310, given a peer evaluation matrix S, the objective is to identify the response rk that is collectively judged to be the least biased. This embodiment employs Porta counting, a voting method known for its robustness to strategic manipulation and its tendency to select consensus candidates. For each response rk k Its wavelet fraction b k It is calculated based on pairwise comparisons in the ranking of each evaluator. Specifically, for each evaluator M... j According to the score s ji For all responses {r i} i≠j Rank them. Response r k From evaluator M j The points earned are based on M j ratings ji How many other responses did it beat? i It depends on (i≠j,k). Assume that lower bias assessment scores are better (less bias): ; Where I(·) is the indicator function. This formula summarizes the response r. k The number of wins in pairwise comparisons of all evaluators j (excluding self-evaluation j=k).
[0042] S320, with the highest arrival fraction b k response r ∗ Selected as a consensus target: ; This r ∗ This serves as an example of a low-biased target extracted from the collective judgment of the model group for a given cue q. It is used as the truth value for fine-tuning in the next stage.
[0043] Consensus Mechanism Analysis: The effectiveness of this cross-model adjudication framework depends on whether the response chosen by the consensus mechanism is less biased than that of a typical single output. This section provides the theoretical basis for achieving improvements through consensus.
[0044] make CBM (r m Model M m Response r m The true bias score. Let s jk For M j Assign r k The score. Assume the evaluation score is a noisy estimate of the true bias, i.e., E[ ]= CBM (r k )(in (Converted to a higher score is better). Under the standard voting theory assumptions, Boda counting is considered a Kondorsetian efficient method, meaning that if the response r k If it can beat all other responses in pairwise comparisons with most evaluators, the Boda count will likely be chosen.
[0045] Theorem 1 (Consensus Improvement Boundary): Assume the evaluation score s jk Bias towards reality CBM (r k It provides an unbiased estimate with bounded variance, and the estimator error has reasonable independence. Then, the consensus response r selected by the Porta count... ∗ Expectation biases are satisfied: ; Where δ(N) is an error term that decreases as the number of diverse evaluators N increases. With high probability, the selected response r... ∗ The bias score will be less than or equal to all generated responses {r} m The minimum bias score in}.
[0046] Proof: This result relies on the law of large numbers applied to aggregated scores. With the contribution of more diverse evaluators, the independent noise or bias in individual evaluations tends to average, causing the Polda ranking to tend towards the true latent bias ranking. Polda counts aggregate pairwise comparisons, and under the independence assumption, the probability of selecting a suboptimal candidate decreases rapidly with increasing N. The pairwise winning probability in Polda score calculation can be formally defined using centralized inequalities such as the Hoeffding inequality.
[0047] Corollary 1 (Ensemble Advantage): If model M m When the exhibited biases are sufficiently diverse (e.g., stemming from different datasets or architectures), the consensus mechanism is more likely to identify and filter out biases that only exist in parts of the model (ensemble advantage). If model M m When the exhibited biases are sufficiently diverse (e.g., originating from different datasets or architectures), consensus mechanisms are more likely to identify and filter out biases specific to only a subset of the models. Collective judgments benefit from this diversity, resulting in consensus errors typically performing better than when relying on a single best model, potentially even under simplifying assumptions. The rate of decay.
[0048] like Figure 1 As shown, in step S400, the parameters of each participating LLM model are fine-tuned based on the obtained low-bias consensus target response.
[0049] The above steps specifically include: This embodiment uses Low-Rank Adaptation (LoRA) for efficient parameter fine-tuning, as detailed below: The final step in cross-model adjudication is to use a low-bias consensus-derived response r ∗ For each participating model M m Fine-tuning is performed. To maintain computational feasibility and preserve the model's generality, this embodiment employs LoRA, a parameter-efficient fine-tuning technique.
[0050] For each model M m LoRA introduces a specific weight matrix W0∈R in the original model. d×k The low-rank decomposition matrix B (usually the attention layer) m ∈R d×r and A m ∈R r×k Model updates are restricted to , where rank Only A m and B m Being trained significantly reduces the number of trainable parameters.
[0051] Model M mThe fine-tuning objective is to minimize a distillation loss, typically the model's predicted distribution of response terms versus the target consensus response r. ∗ Cross-entropy loss between: ; in These are the prompts and their corresponding consensus response pairs generated by cross-model adjudication. It is the standard cross-entropy loss generated from the sequence. The original model weights W0 remain frozen. The specific algorithm process is as follows:
[0052] Lemma 1 (parameter efficiency): For a weight matrix W0∈R d×k A complete fine-tuning requires updating dk parameters. LoRA using rank r only requires updating A. m and B m The r(d+k) parameters. Typically... This significantly saves parameters (e.g., 0.1%–1% of the total parameters). This makes the fine-tuning stage computationally efficient.
[0053] Experimental example: This section details the experimental setup, evaluation metrics, and comprehensive results demonstrating the effectiveness of the cross-model adjudication framework (cross-model adjudication) in detecting and mitigating large language model bias. This embodiment evaluates the performance of cross-model adjudication relative to a baseline model and conducts ablation studies to understand the contributions of its key components.
[0054] Table 1. Computational complexity analysis of cross-model adjudication (for each hint)
[0055] Where N: number of models. L, L′: length of the generated / evaluated sequences. C infer Cost per inference step. Parallel complexity assumes sufficient resources.
[0056] Experimental setup: 1) Hardware and software environment All experiments were conducted on a cluster of machines running Ubuntu 20.04.6LTS. The software stack included Python 3.9.12, PyTorch 1.11.0 + cuda 11.3, and Transformers library version 4.50.0. For efficient parameter fine-tuning, this embodiment used Unsloth version 2025.3.19, which optimizes the LoRA implementation. 2) Model Specifications This embodiment selected four state-of-the-art 7-9B parameter LLMs as experimental subjects, representing different architectural foundations and pre-training objectives. This diversity is crucial for the cross-model adjudication mechanism, as described in Section 1. Model details are shown in Table 2. All models were loaded with their original, publicly released fine-tuned versions of the instructions. To improve memory efficiency during the multi-model evaluation phase, this embodiment employed 8-bit quantization during model loading and utilized gradient checkpoints to manage GPU memory usage during fine-tuning.
[0057] Table 2 Model Specifications, Training Details, and Optimization Techniques
[0058] 3) Bias assessment and prompting dataset generation This embodiment uses HolisticBias, a large dataset designed to evaluate social bias in language models across 13 protected population categories, including gender, religion, nationality, and age. It contains approximately 600 identity lexical embeddings within template sentence structures, generating approximately 450,000 bias detection examples. This embodiment uses HolisticBias both to generate initial prompts for cross-model adjudication responses (particularly by focusing on incomplete sentence templates related to religion and nationality) and to assess residual bias after reduction. Furthermore, this embodiment references expert-annotated severity scores provided in HolisticBias to validate the bias reduction metrics of this embodiment.
[0059] 4) Evaluation Indicators Given the use of group G k The template hint T[·] is instantiated from the identity lexical gi. In this embodiment, the bias of the model is defined as a function of the likelihood function assigned to these sentences. Let L lik (g i )=−logP(g i |T) represents the negative log-likelihood (i.e., word-level surprise or pseudo-perplexity) assigned to the model when completing a template containing identity gi. To evaluate whether the model performs well in different groups (G... i G j The relationships between groups exhibit biased behavior. In this embodiment, the word-level likelihood distribution L associated with each group is analyzed. CBM (g) Perform the nonparametric Mann-Whitney U test: U i,j =MannWhitney(L lik (G i ),L lik (G jThis embodiment reports whether the differences in likelihood distributions between identity groups are statistically significant (typically p < 0.05), indicating that the model is biased in ranking or interpreting cues involving different demographic terms.
[0060] Implementation details: 1) Response Generation Phase For Phase 1 (Response Generation), each model generates completions for incomplete sentences, with a maximum sequence length of 128 tokens. Kernel sampling (p=0.9) and a temperature τ=0.7 are used for random sampling to encourage diverse and coherent outputs.
[0061] 2) Cross-model evaluation phase In Phase 2, each evaluator model M j Evaluate the responses generated by its peers r k (Where j=k). The assessment provides clear guidance for M. j Consider the CBM dimension and provide a single overall bias score from 1 to 5, where 1 indicates the most severe bias and 5 indicates the least severe bias.
[0062] 3) Consensus refinement and fine-tuning data preparation Phase 3 involves aggregating peer evaluation scores to obtain a consensus-derived "unbiased" output for fine-tuning. For each prompt, all models' responses are scored by their peers. This embodiment then applies a modified Polda count mechanism to determine the response with the least overall bias for that particular prompt. In cases where multiple models obtain the same highest Polda score, a response is randomly selected from them as the fine-tuning target. If an evaluator model fails to provide a score (e.g., due to inference errors), its score for that particular response is conservatively set to 3 (neutral) to minimize interference with the consensus mechanism.
[0063] 4) Efficient parameter fine-tuning This embodiment uses LoRA to analyze each model M. m Perform efficient parameter fine-tuning. The key hyperparameters for LoRA fine-tuning are as follows: •LoRA rank (r): 16 •LoRAalpha(α): 16 (scaling factor for LoRA weights) • LoRAdropout: 0 (dropout is not applied to LoRA weights) • Learning rate: 2e-4 • Training batch size per device: 2 Gradient accumulation steps: 4 (effective batch size is 8) • Maximum sequence length: 2048 (applicable to input prompts and target responses) • Number of cycles: 8 • Quantization: 'load_in_4bit=True' is used for efficient memory loading of the base model.
[0064] These settings are implemented using the unsloth library to optimize performance.
[0065] This invention proposes a method for detecting and mitigating bias in LLM models through collaborative evaluation and refinement. The cross-model adjudication framework performs peer review across different LLM models to identify and correct biases that might be imperceptible in a single model or static benchmark. The framework comprises three consecutive phases: response generation, cross-model evaluation, and consensus refinement, followed by efficient parameter fine-tuning using the refined consensus output.
[0066] This process ensures that: • Mutual review: Bias that one model may ignore can be identified by other models with different training data or architectures.
[0067] • Asymmetric evaluation: When evaluating peers in the evaluation step, the model cannot access its own internal state, thus promoting objective evaluation.
[0068] • Information isolation: During the evaluation phase, models do not share any internal model parameters or gradients, maintaining model independence.
[0069] • Adaptive consensus: The framework identifies the least biased response based on collective judgment, rather than relying on a fixed definition of "unbiased".
[0070] Conceptually, the cross-model adjudication framework transforms bias reduction into a collaborative refinement process, where models leverage their diverse perspectives to converge on less biased representations and outputs.
[0071] Corresponding to the aforementioned method for reducing large language model bias through cross-model adjudication, this invention also discloses a system for reducing large language model bias through cross-model adjudication, such as... Figure 2 As shown, it specifically includes: The response generation module is used to take a set of multiple different LLM models, input the preset prompts into the different LLM models respectively, and obtain the response generated by each LLM model. The cross-model assessment module is used to evaluate the level of bias in the responses generated by all other LLM models in the peer group based on a composite bias metric, and to obtain a bias assessment score. The consensus extraction module is used to calculate the arrival score of each response based on the obtained bias evaluation score and using an improved arrival counting mechanism. The response with the highest arrival score is selected as the consensus target response with the least bias in collective consensus. The model fine-tuning module is used to fine-tune the parameters of each participating LLM model based on the obtained low-bias consensus target response.
[0072] It should be noted that for a detailed description of a cross-model adjudication system for large language model bias reduction provided in the embodiments of the present invention, please refer to the relevant description of a cross-model adjudication method for large language model bias reduction provided in the embodiments of the present invention, which will not be repeated here.
[0073] The above examples illustrate the present invention only to aid in understanding it and are not intended to limit the scope of the invention. Those skilled in the art can make various simple deductions, modifications, or substitutions based on the principles of this invention.
Claims
1. A method for reducing bias in large language models across models, characterized in that, The method includes: Given a set of multiple different LLM models, input the preset prompts into the different LLM models respectively, and obtain the response generated by each LLM model; Based on a composite bias metric, the bias level of the responses generated by all other LLM models in the peer group is evaluated using each LLM model, and a bias assessment score is obtained. Specifically, based on a composite bias metric, the bias level of responses generated by all other LLM models in the peer group is evaluated using each LLM model. The composite bias metric provides the LLM model with a set of guiding principles or conceptual dimensions presented through prompts; Model M j The original response r is evaluated based on different dimensions of the established composite bias metric, including relevance, vocabulary, and context. k The bias level, k∈{1,...,N}, j ≠ k, N is the number of LLM models; When evaluating the level of bias in responses generated by all other LLM models in a peer group based on a composite bias metric, the specific steps also include: Guiding revisions: The guidelines explicitly require model M to be correct. j Generate response r k The revised version is represented as The aim is to minimize biases identified in the evaluation process while preserving the original intent and core information. Counterfact generation: Hints require model M j Generate a modified version by systematically changing sensitive attributes. Alternative version ; Assessment scores jk From M j Comprehensive analysis: ; in, Representation Model Assessment Response The process Representation Model Evaluating the response based on different dimensions of the composite bias metric The numerical score output at that time. It is a model Optional confidence weights, This represents a small amount of Gaussian noise added to the fraction. It's a hyperparameter; The output of this stage is an N×N fractional matrix S, which reflects the original response {r1,...,r...} N Collective judgments on the level of bias, resulting in revisions and counterfacts As a supplementary data collection; Based on the obtained bias assessment score, the arrival score of each response is calculated using an improved arrival counting mechanism, and the response with the highest arrival score is selected as the consensus target response with the least collective bias. Specifically, based on the obtained bias assessment score, an improved arrival score is used to calculate the arrival score of each response. The response with the highest arrival score is selected as the consensus target response with the least collective bias, which includes: For each response r k Its wavelet fraction b k It is calculated based on pairwise comparisons in the ranking of each evaluator model. Specifically, for each evaluator model M... j According to the bias assessment score s ji For all responses {r i } i≠j Rank, respond r k From the evaluator model M j The points earned are based on M j Bias assessment scores ji How many other responses did it beat? i Determined: ; Where I(·) is the indicator function; b k The calculation formula summarizes the response r k In all evaluator models M except for self-evaluation j The "wins" in pairwise comparisons; With the highest arrival fraction b k response Selected as the consensus target response: ; As a low-biased target response extracted from the collective judgment of the model group for a given cue q; Based on the obtained low-biased consensus target response, the parameters of each participating LLM model are fine-tuned.
2. The method for reducing bias in large language models across models as described in claim 1, characterized in that, Given a set of multiple different LLM models, preset prompts are input into different LLM models respectively, and the responses generated by each LLM model are obtained, specifically including: A set of N different LLM models is represented as M = {M1, M2, ..., M}. N }; For a given input cue q, each model M m ∈M generates a response r m =(y1,y2,...,y L ), where L is the sequence length.
3. The method for reducing bias in large language models across models as described in claim 2, characterized in that, Given a set of multiple different LLM models, preset prompts are input into different LLM models respectively, and the responses generated by each LLM model are obtained, specifically including: The generated responses are randomly sampled to encourage diversity while maintaining coherence, specifically using kernel sampling combined with temperature scaling: ; Where z m,t,yt Model M m At time step t, the term y is... t The generated logits, where τ is the temperature parameter.
4. The method for reducing bias in large language models across models as described in claim 1, characterized in that, Based on a composite bias metric, the bias level of responses generated by all other LLM models in the peer group is evaluated using each LLM model, specifically including: The relevance bias dimension is used to indicate the model M. j Evaluation response r k Whether stereotyped associations are established implicitly or explicitly between demographic groups and specific attributes, roles, or concepts that reflect social biases, the corresponding dimension aims to capture the characteristics of associated biases measured in the embedding space, by requiring LLM to identify subtle semantic links; Lexical bias dimension used in cue indication model M j Identify response r k Does the language contain overtly biased, stereotypical, derogatory, or harmful language? This includes examining words or phrases associated with negative stereotypes or unfair generalizations, and using LLM's linguistic knowledge to detect relevant problem terms in context. Contextual bias dimension is used to prompt the guidance model M j Evaluation response r k The overall narrative, emotion, and implied meaning include assessing whether the text subtly reinforces stereotypes, presents different groups in an unbalanced or unfair manner, or normalizes biased viewpoints, even without explicitly using biased terms; the corresponding dimensions aim to capture features that measure deviations from neutral or equitable distributions in a broader context.
5. The method for reducing bias in large language models across models as described in claim 1, characterized in that, Based on the obtained low-biased consensus target response, the parameters of each participating LLM model are fine-tuned, specifically including: The LoRA mechanism is used for efficient fine-tuning of model parameters: for each model M m LoRA introduces a low-rank decomposition matrix B for a specific weight matrix W0 in the original model. m and A m Model updates are restricted to ; Model M m The fine-tuning objective is to minimize a distillation loss, i.e., the model's predicted distribution of response terms is close to the target consensus response r. Cross-entropy loss between: ; in It's a hint and its corresponding consensus target response right, It is the standard cross-entropy loss generated from the sequence, while the original model weights W0 remain frozen.
6. A large language model bias reduction system for cross-model adjudication, characterized in that, The system includes: The response generation module is used to take a set of multiple different LLM models, input the preset prompts into the different LLM models respectively, and obtain the response generated by each LLM model. The cross-model assessment module is used to evaluate the level of bias in the responses generated by all other LLM models in the peer group based on a composite bias metric, and to obtain a bias assessment score. Specifically, based on a composite bias metric, the bias level of responses generated by all other LLM models in the peer group is evaluated using each LLM model. The composite bias metric provides the LLM model with a set of guiding principles or conceptual dimensions presented through prompts; Model M j The original response r is evaluated based on different dimensions of the established composite bias metric, including relevance, vocabulary, and context. k The bias level, k∈{1,...,N}, j ≠ k, N is the number of LLM models; When evaluating the level of bias in responses generated by all other LLM models in a peer group based on a composite bias metric, the specific steps also include: Guiding revisions: The guidelines explicitly require model M to be correct. j Generate response r k The revised version is represented as The aim is to minimize biases identified in the evaluation process while preserving the original intent and core information. Counterfact generation: Hints require model M j Generate a modified version by systematically changing sensitive attributes. Alternative version ; Assessment scores jk From M j Comprehensive analysis: ; in, Representation Model Assessment Response The process Representation Model Evaluating the response based on different dimensions of the composite bias metric The numerical score output at that time. It is a model Optional confidence weights, This represents a small amount of Gaussian noise added to the fraction. It's a hyperparameter; The output of this stage is an N×N fractional matrix S, which reflects the original response {r1,...,r...} N Collective judgments on the level of bias, resulting in revisions and counterfacts As a supplementary data collection; The consensus extraction module is used to calculate the arrival score of each response based on the obtained bias evaluation score and using an improved arrival counting mechanism. The response with the highest arrival score is selected as the consensus target response with the least bias in collective consensus. Specifically, based on the obtained bias assessment score, an improved arrival score is used to calculate the arrival score of each response. The response with the highest arrival score is selected as the consensus target response with the least collective bias, which includes: For each response r k Its wavelet fraction b k It is calculated based on pairwise comparisons in the ranking of each evaluator model. Specifically, for each evaluator model M... j According to the bias assessment score s ji For all responses {r i } i≠j Rank, respond r k From the evaluator model M j The points earned are based on M j Bias assessment scores ji How many other responses did it beat? i Determined: ; Where I(·) is the indicator function; b k The calculation formula summarizes the response r k In all evaluator models M except for self-evaluation j The "wins" in pairwise comparisons; With the highest arrival fraction b k response Selected as the consensus target response: ; As a low-biased target response extracted from the collective judgment of the model group for a given cue q; The model fine-tuning module is used to fine-tune the parameters of each participating LLM model based on the obtained low-bias consensus target response.
Citation Information
Patent Citations
Complex network node influence sorting method combining graph embedding and graph sampling
CN119622623A
Word embedding similarity-based lexicon dynamic expansion and ESG expression quantitative evaluation method
CN119760125A