A large language model context hallucination mitigation method based on bidirectional frequency alignment

CN122175002BActive Publication Date: 2026-10-09HUALAN DESIGN GRP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610295570.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-03-11
Publication Date
2026-10-09
Estimated Expiration
2046-03-11

AI Technical Summary

Technical Problem

[0005]本发明的目的是提供一种基于双向频率对齐的大语言模型上下文幻觉缓解方法,旨在解决或改善上述技术问题中的至少之一

Benefits of technology

本发明公开了一种基于双向频率对齐的大语言模型上下文幻觉缓解方法,所述方法通过构建上下文感知的分类器与双向频率校准的创新架构,解决了现有大语言模型解码阶段偏好全局高频低相关token的技术痛点。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122175002B_ABST
    Figure CN122175002B_ABST
Patent Text Reader

Abstract

The application discloses a large language model context illusion relief method based on bidirectional frequency alignment, and relates to the technical field of natural language processing. The method comprises the following steps: constructing anti-illusion positive and negative sample pairs in a training stage; constructing a word feature vector by extracting the multi-head attention ratio and normalized frequency of a word; generating a frequency-weighted training sample set by performing hierarchical resampling according to the feature distribution; training an optimal discriminant weight vector and an optimal discriminant classifier with local L2 regularization constraints by minimizing a frequency-weighted loss function; obtaining an original probability distribution generated by a large language model for an input question and answer pair in an inference stage, calculating a correlation score based on the optimal discriminant classifier, adjusting by bidirectional frequency alignment, obtaining a final probability distribution, and outputting a final answer. Through the construction of the innovative architecture of the context-aware classifier and bidirectional frequency calibration, the technical pain point that the existing large language model decoding stage prefers global high-frequency low-relevance tokens is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing technology, and in particular to a method for mitigating contextual illusions in large language models based on bidirectional frequency alignment. Background Technology

[0002] Large Language Models (LLMs) have demonstrated outstanding performance in natural language processing tasks such as text generation, question answering, and summarization. Their generation quality is highly dependent on the lexical probability distribution calibration strategy during the decoding stage.

[0003] Existing decoding methods generally suffer from a key flaw: large language models tend to select high-frequency lexical units from the global corpus that are not highly relevant to the current context, such as common verbs like "is" and "has," while ignoring low-frequency but highly relevant key lexical units in the context, such as entity information like time and amount. This flaw can lead to contextual illusions in the generated content, reducing the factual accuracy of the output text.

[0004] Although some studies have noted the impact of word frequency on decoding results, existing technologies have not systematically integrated frequency signals into the decoding process, making it impossible to fundamentally correct the word ordering distortion caused by statistical priors within the model, and making it difficult to balance the conflict between model statistical regularity and contextual semantic evidence. Summary of the Invention

[0005] The purpose of this invention is to provide a method for mitigating contextual illusions in large language models based on bidirectional frequency alignment, which aims to solve or improve at least one of the above-mentioned technical problems.

[0006] To achieve the above objectives, the present invention provides the following solution: A method for mitigating contextual illusions in large language models based on bidirectional frequency alignment includes: During the training phase, positive and negative anti-hallucination sample pairs are constructed; multi-head attention ratio and normalized frequency of word units are extracted to construct word unit feature vectors; hierarchical resampling is performed according to the feature distribution to generate frequency-weighted training sample sets; and the optimal discriminant weight vector and optimal discriminant classifier with local L2 regularization constraint are trained by minimizing the frequency-weighted loss function. During the reasoning phase, the original probability distribution of the input question-answer pair generated by the large language model is obtained, and the relevance score is calculated based on the optimal discriminant classifier. Bidirectional frequency alignment adjustment is performed to obtain the final probability distribution, and the final answer is output.

[0007] Furthermore, constructing anti-hallucination positive and negative sample pairs includes: Obtain multiple question-answer pairs as the training dataset; For each question-answer pair, add the words in the answer to the positive sample pool; Extract all lexical units from the input context of the question-answer pair to generate a negative sample candidate pool; use a named entity recognition tool to filter the negative sample candidate pool and retain entity lexical units; remove a preset number of lexical units with the highest probability ranking in the original output of the model, isolate high-frequency interference items that are likely to cause hallucinations, and finally determine the negative samples.

[0008] Furthermore, the multi-head attention ratio and normalized frequency of word units are extracted to construct word unit feature vectors, including: Extract the weight distribution of each attention head and calculate the normalized attention ratio to capture the global semantic relevance of lexical units. The expression is as follows: In the formula, Let be the attention ratio of the i-th attention head; is the average attention weight of word w in the i-th attention head; k is the total number of attention heads in the model; Count the frequency of word w in the context and calculate the normalized frequency, expressed as: In the formula, Normalized frequency; This represents the number of times the word w appears in the input context; The set of lexical terms for the input context; It is the most frequent word element; Finally, word feature vectors are generated. .

[0009] Furthermore, hierarchical resampling is performed based on the feature distribution to generate a frequency-weighted training sample set, including: Using a uniform distribution strategy, the sampling frequency of positive samples is obtained, expressed as: In the formula, The sampling frequency for positive samples; The set of positive samples; Construct an anti-hallucination sampling score, expressed as: In the formula, Sampling score for anti-hallucination; Average attention weight; Normalized frequency; It is a smoothing factor; This is the frequency penalty coefficient; Normalizing the anti-hallucination sampling score yields the negative sample sampling frequency, expressed as: In the formula, The sampling frequency for negative samples; For the negative sample set; The traversal variable in the negative sample set; Resampling is performed based on the sampling frequencies of positive and negative samples to obtain a frequency-weighted training sample set, which can also be expressed as: In the formula, The training sample set is frequency-weighted. It is a weighted random sampling function; The target number of samples.

[0010] Furthermore, the frequency-weighted loss function is expressed as follows: In the formula, This is the total loss function; The weight vector of the model to be learned; The training sample set is frequency-weighted. For binary cross-entropy loss; It is the sigmoid activation function; The true label of the i-th instance; Let be the lexical feature vector of the i-th instance; This is the regularization strength parameter; It is the square of the L2 norm.

[0011] Furthermore, obtain the original probability distribution of the input question-answer pair generated by the large language model, including: Using a pre-trained large language model, the context and question of the question-answer pair are input, and forward propagation is performed to obtain the original probability distribution of the next word in the model output, expressed as: In the formula, This represents the original probability distribution; It is a normalized exponential function; This represents the original logarithmic probability.

[0012] Furthermore, bidirectional frequency alignment adjustment is performed to obtain the final probability distribution, and the final answer is output, including: Calculate the global debiasing factor and context enhancement factor for lexical units; Based on the global debiasing factor, context enhancement factor, and relevance score, the original probability distribution is linearly fused in the Logits space and then normalized by Softmax to obtain the calibrated final probability distribution. A greedy decoding strategy is adopted to select the word with the highest probability from the calibrated final probability distribution as the generation result of the current step; Append the result generated in the current step to the context sequence, update the context state, and repeat the above steps until an end symbol is generated or the maximum length is reached, then output the final answer.

[0013] Furthermore, the expression for the global debiasing factor is: In the formula, This is the global debiasing factor; Let w be the global frequency of the lexical unit w; For smoothing terms; It is the natural logarithm function.

[0014] Furthermore, the expression for the context enhancement factor is: In the formula, For context enhancement factors; To enhance the intensity weighting factor; The normalized context frequency of the lexical unit w; Context frequency of lexical w The expression is: In the formula, The frequency of occurrence of the word element w; For the current context sequence; For indicator functions; For contextual lexical sets; For traversing variables; is the normalized context frequency of the lexical w.

[0015] Furthermore, the final probability distribution is expressed as follows: In the formula, This represents the final probability distribution; It is a normalized exponential function; This represents the original probability distribution of the model. This is the global debiasing factor; For context enhancement factors; Assess relevance scores; To balance the hyperparameters.

[0016] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects: This invention discloses a method for mitigating contextual illusion in large language models based on bidirectional frequency alignment. The method solves the technical pain point of existing large language models in the decoding stage by constructing a context-aware classifier and an innovative architecture of bidirectional frequency alignment, which favors globally high-frequency, low-relevance tokens.

[0017] This invention does not require modification of the original model architecture or task-specific fine-tuning. It can improve the factual accuracy and contextual relevance of the generated text simply by dynamically calibrating the probability distribution during the decoding stage. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of the overall process of the method of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] The purpose of this invention is to provide a method for mitigating contextual illusions in large language models based on bidirectional frequency alignment, which aims to solve or improve at least one of the above-mentioned technical problems.

[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0023] like Figure 1 As shown, this invention provides a method for mitigating contextual illusions in large language models based on bidirectional frequency alignment, comprising: During the training phase, anti-hallucination positive and negative sample pairs are constructed; multi-head attention ratios and normalized frequencies of lexical units are extracted to construct lexical feature vectors; hierarchical resampling is performed based on the feature distribution to generate a frequency-weighted training sample set; by minimizing the frequency-weighted loss function, the optimal discriminant weight vector and optimal discriminant classifier with local L2 regularization constraints are trained, including: S1, Obtain multiple question-answer pairs as the training dataset, and construct training sample pairs using a positive and negative sample separation strategy, including: S11, Select the NaturalQuestions dataset as the training data; the NaturalQuestions dataset includes multiple question-answer pairs, each of which includes the input context, the question, and the corresponding correct answer; S12, for each question-answer pair, construct training sample pairs using a positive-negative sample separation strategy, including: Positive sample construction: For each question-answer pair, add the words in the answer to the positive sample pool; Negative sample construction: All lexical units are extracted from the input context to generate a negative sample candidate pool; the negative sample candidate pool is filtered using a named entity recognition tool to retain entity lexical units; a preset number of lexical units with the highest probability ranking in the original model output are removed, and high-frequency interference items that are prone to causing hallucinations are isolated to finally determine the negative samples. In this embodiment, the preset number is 10.

[0024] S2, for each word, an attention-frequency dual-modal fusion method is used to construct a word feature vector, including: Extract the weight distribution of each attention head and calculate the normalized attention ratio to capture the global semantic relevance of lexical units. The expression is as follows: In the formula, The attention ratio of the i-th attention head eliminates the dimensional differences in the absolute values ​​of attention between different samples; is the average attention weight of word w in the i-th attention head, reflecting the degree of attention the attention head pays to word w; k is the total number of attention heads in the model; The frequency of word w in the context is counted, and the normalized frequency is calculated to capture explicit repetition patterns of word w. The expression is as follows: In the formula, Normalized frequency, range of values A value of 1 indicates that the word w is the most frequently occurring word in the current context; This represents the number of times the word w appears in the input context; The set of lexical terms for the input context; It is the most frequent word element; Finally, word feature vectors are generated. .

[0025] S3, calculate the differentiated sampling priority for each lexical unit, perform hierarchical resampling, and generate a frequency-weighted training sample set, including: Positive sample oversampling: A uniform distribution strategy is adopted to ensure that sparse positive samples are fully covered, resulting in the positive sample sampling frequency, expressed as: In the formula, The sampling frequency for positive samples; The set of positive samples; Dynamic weighting of negative samples: Constructing an anti-hallucination sampling score, the expression is: In the formula, Sampling score for anti-hallucination; Average attention weight; Normalized frequency; This is a smoothing factor, a very small constant, to prevent the denominator from being zero; This is the frequency penalty coefficient. The larger the value, the stronger the suppression of high-frequency noise; Normalizing the anti-hallucination sampling score yields the negative sample sampling frequency, expressed as: In the formula, The sampling frequency for negative samples; For the negative sample set; is the traversal variable in the negative sample set, representing any negative sample word in the set; Resampling is performed based on the sampling frequencies of positive and negative samples to obtain a frequency-weighted training sample set, expressed as: In the formula, The training sample set is frequency-weighted. It is a weighted random sampling function; The target number of samples.

[0026] S4, based on the frequency-weighted training sample set and word feature vectors, construct an L2-regularized logistic regression classifier, including: The optimization objective is to minimize the frequency-weighted loss function, expressed as: In the formula, This is the total loss function; The weight vector of the model to be learned; The training sample set is frequency-weighted. For binary cross-entropy loss; It is the sigmoid activation function; The true label of the i-th instance; Let be the lexical feature vector of the i-th instance; This is the regularization strength parameter; It is the square of the L2 norm.

[0027] During the inference phase, the original probability distribution of the input question-answer pair generated by the large language model is obtained. Based on the optimal discriminant classifier obtained during the training phase, a relevance score is calculated, and bidirectional frequency alignment adjustment is performed to obtain the final probability distribution. Finally, the final answer is output, including: S1 uses a pre-trained large language model, inputting the context and question of the question-answer pair, performing forward propagation to obtain the original probability distribution of the next word element output by the model, expressed as: In the formula, This represents the original probability distribution; It is a normalized exponential function; is the original log probability, representing the unnormalized log probability vector of the output layer of the large language model; S2, obtain the feature vector of each word, input it into the optimal discriminant classifier, and obtain the relevance score of the word; S3 calculates the global debiasing factor and context enhancement factor for lexical units to correct for statistical bias, including: Applying a probability penalty to globally high-frequency words to reduce their selection probability, the expression for the global debiasing factor is: In the formula, This is the global debiasing factor; Let w be the global frequency of word w, and let w represent the statistical frequency of word w in a large-scale pre-training corpus. To smooth out the terms and avoid a denominator of zero; It is the natural logarithm function; The context enhancement factor increases the generation probability of high-frequency keywords in the current context. Its expression is: In the formula, For context enhancement factors; To enhance the intensity weighting factor, used to adjust the influence intensity of context frequency; The normalized context frequency of the lexical unit w; Context frequency of lexical w The expression is: In the formula, The frequency of occurrence of the word element w; For the current context sequence; For indicator functions, when the condition The value is 1 if the condition is met, and 0 otherwise. The context lexical set represents the set of all non-repeating lexical elements in the current context; For traversing variables; The normalized context frequency of the lexical unit w; S4. Based on the global debiasing factor, context enhancement factor, and relevance score, the original probability distribution is linearly fused in the Logits space. After Softmax normalization, the calibrated final probability distribution is obtained, expressed as: In the formula, This represents the final probability distribution; It is a normalized exponential function; The original probability distribution of the model represents the initial Softmax probability output by the LLM without intervention; This is the global debiasing factor; For context enhancement factors; Assess relevance scores; To balance the hyperparameters.

[0028] S5 employs a greedy decoding strategy, selecting the term with the highest probability from the calibrated final probability distribution as the generation result for the current step. The expression is: In the formula, This represents the result generated in the current step. S6, generate the result of the current step. Add it to the context sequence, update the context state, and repeat S1-S5 until an end symbol is generated or the maximum growth is reached, then output the final answer.

[0029] To verify the effectiveness of this invention, on the NaturalQuestionsShort dataset, the method of this invention improved the F1 score by 3.02% compared to the DAGCD decoding method, proving that this method can effectively improve the accuracy of question answering tasks.

[0030] In summary, the large language model decoding method based on bidirectional frequency alignment disclosed in this invention solves the technical pain point of existing large language model decoding stages that favor globally high-frequency, low-relevance tokens by constructing a context-aware classifier and an innovative architecture of bidirectional frequency calibration.

[0031] This invention does not require modification of the original model architecture or task-specific fine-tuning. It can improve the factual accuracy and contextual relevance of the generated text simply by dynamically calibrating the probability distribution during the decoding stage.

[0032] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0033] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for alleviating contextual illusions in large language models based on bidirectional frequency alignment, characterized in that, include: During the training phase, positive and negative anti-hallucination sample pairs were constructed, and the NaturalQuestions dataset was selected as the training data. The NaturalQuestions dataset includes multiple question-answer pairs, each containing the input context, the question, and the corresponding correct answer; multi-head attention ratios and normalized frequencies of lexical units are extracted to construct lexical feature vectors; according to The feature distribution is subjected to hierarchical resampling to generate a frequency-weighted training sample set; the optimal discriminant weight vector and the optimal discriminant classifier with local L2 regularization constraint are trained by minimizing the frequency-weighted loss function. During the inference phase, the original probability distribution of the input question-answer pair generated by the large language model is obtained, and a relevance score is calculated based on the optimal discriminant classifier. Bidirectional frequency alignment adjustment is then performed to obtain the final probability distribution, and the final answer is output, including: Calculate the global debiasing factor and context enhancement factor for lexical units; Based on the global debiasing factor, context enhancement factor, and relevance score, the original probability distribution is linearly fused in the Logits space and then normalized by Softmax to obtain the calibrated final probability distribution. The expression for the global debiasing factor is: In the formula, This is the global debiasing factor; The global frequency of the lexical w; For smoothing terms; It is the natural logarithm function; The expression for the context enhancement factor is: In the formula, For context enhancement factors; To enhance the intensity weighting factor; The normalized context frequency of the lexical unit w; Context frequency of lexical w The expression is: In the formula, The frequency of occurrence of the word element w; For the current context sequence; For indicator functions; For contextual lexical sets; For traversing variables; The normalized context frequency of the lexical unit w; A greedy decoding strategy is adopted to select the word with the highest probability from the calibrated final probability distribution as the generation result of the current step; Append the result generated in the current step to the context sequence, update the context state, and repeat the above steps until an end symbol is generated or the maximum length is reached, then output the final answer.

2. The method for alleviating contextual illusions in a large language model based on bidirectional frequency alignment as described in claim 1, characterized in that, The construction of anti-hallucination positive and negative sample pairs includes: Obtain multiple question-answer pairs as the training dataset; For each question-answer pair, add the words in the answer to the positive sample pool; Extract all lexical units from the input context of the question-answer pair to generate a negative sample candidate pool; use a named entity recognition tool to filter the negative sample candidate pool and retain entity lexical units; remove a preset number of lexical units with the highest probability ranking in the original output of the model, isolate high-frequency interference items that are likely to cause hallucinations, and finally determine the negative samples.

3. The method for alleviating contextual illusions in a large language model based on bidirectional frequency alignment as described in claim 1, characterized in that, The construction of word feature vectors by extracting word units using multi-head attention ratios and normalized frequencies includes: Extract the weight distribution of each attention head and calculate the normalized attention ratio to capture the global semantic relevance of lexical units. The expression is as follows: In the formula, Let be the attention ratio of the i-th attention head; is the average attention weight of word w in the i-th attention head; k is the total number of attention heads in the model; Count the frequency of word w in the context and calculate the normalized frequency, expressed as: In the formula, The frequency of occurrence of the word element w; For the current context sequence; For indicator functions; For contextual lexical sets; For traversing variables; The normalized context frequency of the lexical unit w; Finally, word feature vectors are generated. .

4. The method for alleviating contextual illusions in a large language model based on bidirectional frequency alignment as described in claim 1, characterized in that, The step of generating a frequency-weighted training sample set by performing hierarchical resampling based on the feature distribution includes: Using a uniform distribution strategy, the sampling frequency of positive samples is obtained, expressed as: In the formula, The sampling frequency for positive samples; The set of positive samples; For word elements; Construct an anti-hallucination sampling score, expressed as: In the formula, Sampling score for anti-hallucination; Average attention weight; Normalized frequency; For smoothing terms; This is the frequency penalty coefficient; Normalizing the anti-hallucination sampling score yields the negative sample sampling frequency, expressed as: In the formula, The sampling frequency for negative samples; The set of negative samples; The traversal variable in the negative sample set; Resampling is performed based on the sampling frequencies of positive and negative samples to obtain a frequency-weighted training sample set, which can also be expressed as: In the formula, The training sample set is frequency-weighted. It is a weighted random sampling function; The target number of samples.

5. The method for alleviating contextual illusions in a large language model based on bidirectional frequency alignment according to claim 1, characterized in that, The frequency-weighted loss function is expressed as follows: In the formula, This is the total loss function; The weight vector of the model to be learned; The training sample set is frequency-weighted. For binary cross-entropy loss; It is the sigmoid activation function; The true label for the j-th instance; Let be the lexical feature vector of the j-th instance; This is the regularization strength parameter; It is the square of the L2 norm.

6. The method for alleviating contextual illusions in a large language model based on bidirectional frequency alignment as described in claim 1, characterized in that, The acquisition of the original probability distribution of the input question-answer pair generated by the large language model includes: Using a pre-trained large language model, the context and question of the question-answer pair are input, and forward propagation is performed to obtain the original probability distribution of the next word in the model output, expressed as: In the formula, This represents the original probability distribution; It is a normalized exponential function; The original logarithmic probability; It is a word element.

7. The method for alleviating contextual illusions in a large language model based on bidirectional frequency alignment as described in claim 1, characterized in that, The expression for the final probability distribution is: In the formula, This represents the final probability distribution; It is a normalized exponential function; This represents the original probability distribution of the model. This is the global debiasing factor; For context enhancement factors; Assess relevance scores; To balance the hyperparameters.

Citation Information

Patent Citations

  • Conversational data generation method and natural language reasoning data classification method and system

    CN118312589A

  • Big language model illusion relieving method based on contrast decoding

    CN118964552A