Large model chat illusion mitigation method based on SND
By dynamically identifying and discarding activation of neurons with large fluctuations in the hidden layer of large language models, the problem of the model producing hallucinations in chat conversations is solved, significantly improving the accuracy and reliability of the model and enhancing user trust.
Patent Information
- Application Number
- CN202510161375.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-06
AI Technical Summary
The hallucinations produced by large language models in chat conversations result in the generated replies that may contain misinformation or false statements, affecting user trust and system reliability, and may have serious consequences in applications in high-risk areas.
Using a SND (Selective Neuron Dropout)-based method, the activation of neurons is dynamically monitored in the N-1 layer of the hidden layer, and those neurons with large fluctuations in different contexts are identified and selectively discarded, thereby reducing the occurrence of hallucinations.
It significantly reduces the probability of the model generating errors or false information, enhances the interpretability within the model, improves user interaction trust, and ensures the accuracy and reliability of content generated in applications in high-risk areas.
Smart Images

Figure CN120104734A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of artificial intelligence, relates to natural language processing, and particularly to a large model chat hallucination reduction method based on SND. Background Art
[0002] The hallucination phenomenon of large language models (LLMs) refers to the fact that the content they generate may be inconsistent with real-world facts, user input, or training data seen in the past. In machine chat conversations, the existence of hallucinations means that the responses generated by the model may contain wrong information or false statements, which will affect the user's trust and dependence on the system. Hallucinations not only reduce the reliability of the model in practical applications, but also may have serious consequences, especially in high-risk fields such as medicine and law. Therefore, how to effectively identify and suppress hallucinations has become an important issue to improve the performance and security of large models.
[0003] At present, in order to reduce the phenomenon of hallucination, the mainstream methods are mostly focused on fine-tuning instructions, adjusting model parameters or introducing external knowledge bases. Although these methods reduce the occurrence of hallucinations to a certain extent, they do not fundamentally solve the potential problems within the model. Summary of the invention
[0004] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a large-model chat hallucination reduction method based on SND, aiming to fundamentally reduce the hallucination phenomenon from within the model, make the machine chat dialogue process more accurate and reliable, and enhance the interactive trust with users.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is:
[0006] A method for alleviating chat hallucination of a large model based on SND comprises the following steps:
[0007] Step 1, obtaining chat sentence samples as input of a large model, wherein the large model is based on a neural network architecture and includes an input layer, a hidden layer and an output layer, wherein the hidden layer has a total of N layers;
[0008] Step 2: In the N-1th layer of the hidden layer, the activation matrix R formed in the previous layer during the forward propagation process is converted into a sentence embedding vector E k ;
[0009] Step 3: Extract the sentence embedding vector E k The activation change between adjacent embedded elements is obtained by difference calculation, which is expressed as checkpoint Δe;
[0010] Step 4: Calculate the variability of the activation values of each embedding element between each checkpoint, sort the variability from large to small, and discard one or several embedding elements with the highest ranking;
[0011] Step 5: Chat using the model after discarding the embedding meta-data.
[0012] The activation matrix R is composed of activation features of multiple neurons, with a dimension of m×n, and is expressed as:
[0013]
[0014] m and n are the columns and rows of the activation matrix R, respectively. n,m represents the activation feature of the nth row and mth column;
[0015] The sentence embedding vector is a column vector obtained by column compression of the activation matrix R, consisting of n embedded elements, expressed as:
[0016]
[0017] Each embedding element is compressed from the activation features of m neurons, and the i-th embedding element is represented by e i =f(r i,1 ,r i,2 ,…,r i,m ).
[0018] The activation matrix R is compressed to obtain the sentence embedding vector The compression formula is expressed as:
[0019]
[0020] In the formula, represents the activation value of the i-th neuron in the N-1th layer of the hidden layer, Represents the activation value of the mth neuron in the N-1th layer of the hidden layer.
[0021] The activation change between adjacent embedding elements is expressed as:
[0022] Δe=|e i -e i-1 |
[0023] In the formula, || represents absolute value calculation.
[0024] In step 4, the i-th embedding element e i Variability V i It is defined as the degree of fluctuation between each checkpoint, and is calculated by e i Checkpoint Δe i With e iThe variance of the means of the two checkpoints involved in the calculation is obtained and is calculated as follows:
[0025]
[0026] is the mean of the checkpoints.
[0027] In step 4, the top 10% of the embedding elements are discarded.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] 1. Reduce hallucinations: This invention effectively reduces the probability of the model generating errors or false information by introducing a special neuron recognition and processing mechanism, significantly reducing the occurrence of hallucinations. By analyzing the variability of hidden layer activation values, potential errors in the model can be discovered and corrected in a timely manner, thereby avoiding the generation of unreal content.
[0030] 2. Enhance the internal interpretability of the model: This invention makes the internal state of the model more transparent by analyzing the neuron activation values, and can provide a deep understanding of the model decision-making process. This interpretability not only helps optimize the model, but also improves the user's trust in the model decision-making process, which is especially important in applications in high-risk fields (such as medical and legal fields).
[0031] 3. Real-time adjustment and correction: The present invention operates at the hidden layer N-1, which can detect and correct problems in the early stages of their occurrence, thereby avoiding delayed corrections. This mechanism can ensure that during user interaction, the model can correct errors in a timely manner and ensure more accurate responses.
[0032] 4. Improve user interaction trust: The present invention improves the trust between users and the system by reducing hallucinations and enhancing the reliability of the model. Especially when facing complex dialogue tasks, users can rely on the real and accurate content generated by the model, thereby enhancing the practicality and acceptability of the system.
[0033] 5. Compatibility with external knowledge bases: The present invention can be seamlessly integrated with external knowledge bases and data sources to further enhance the knowledge updating capability and adaptability of the model. By using external information to assist in correcting the content generated by the model, the performance of the model in different fields can be effectively improved, avoiding illusions caused by lack of the latest information.
[0034] 6. Applicable to multiple application scenarios: The present invention can be applied not only to conventional dialogue generation tasks, but also to a variety of fields such as law, medical treatment, and education, providing more accurate and reliable chat services. In these high-risk scenarios, reducing hallucinations is of great significance.
[0035] 7. Simplified model adjustment process: Compared with traditional fine-tuning or parameter adjustment methods, the present invention simplifies the adjustment process of model performance by analyzing and optimizing the neurons inside the model. This method is more efficient and can significantly reduce the hallucination phenomenon while maintaining the efficiency of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a schematic diagram of the principle of the present invention. DETAILED DESCRIPTION
[0037] The embodiments of the present invention are described in detail below with reference to the accompanying drawings and examples.
[0038] Chat replies from large language models often contain incorrect information or false statements. In addition, when faced with interactions in multilingual and multicultural backgrounds, due to a lack of understanding of context and cultural nuances, they are prone to misunderstanding user intent, further amplifying the problem of fictitious facts, which directly affects the applicability and credibility of large language models. Not only that, the problem of fictitious facts is more significant in high-load environments. For example, when users participate in large-scale promotional activities or hot events, a large number of consultation requests will significantly increase the workload of the intelligent system, causing large language models to more easily generate answers that deviate from the facts or are irrelevant, thereby affecting the user experience.
[0039] In order to meet these challenges, the present invention provides a large-model chat hallucination mitigation method based on Selective Neuron Dropout (SND) technology, which fundamentally reduces the occurrence of fictitious facts by dynamically screening and suppressing potential hallucination generation paths in the model. This method can not only improve the accuracy and credibility of answers, but also adapt to a variety of complex practical application scenarios, thereby helping large language models maintain stable performance in multi-round conversations and high-load environments.
[0040] The principle of the present invention is to dynamically monitor the activation of neurons during model training, identify and selectively discard neurons with large activation fluctuations in different contexts. These special neurons are often the source of hallucinations due to their instability, and easily cause the model to generate answers that do not meet user expectations. Through this method, it is possible to effectively reduce the generation of unrealistic content in the model generation process, reduce the hallucination phenomenon of large language models when generating answers, and retain the overall performance and generation diversity of the model. Compared with the traditional method of randomly discarding neurons, the SND optimization strategy can more accurately locate the neurons that cause hallucinations, thereby significantly improving the accuracy and consistency of answers without losing the overall model performance. This provides a more reliable criterion for determining the termination of training, ensuring that the model not only achieves loss convergence, but also exhibits a stable fact confidence.
[0041] The present invention specifically comprises the following steps:
[0042] Step 1: Data collection and preprocessing.
[0043] The present invention is based on the actual application scenarios of large models, and uses a large amount of chat records, FAQ (Frequently Asked Questions) data and user feedback as training data sources. By preprocessing and normalizing these data, invalid information and redundant data are removed, thereby ensuring the quality and diversity of input data. This data preparation process can help the system better understand the diverse needs of users in different scenarios and improve the model's responsiveness to user inquiries.
[0044] The large model of the present invention is based on a neural network architecture, including an input layer, a hidden layer and an output layer. The input layer takes the acquired chat sentence samples as input, the hidden layer has a total of N layers, and the processing is transmitted layer by layer, and the output layer obtains its processing results.
[0045] Step 2: Get special neurons.
[0046] In this step, first, in the N-1th layer of the hidden layer, the activation matrix R formed in the previous layer during the forward propagation process is converted into a sentence embedding vector E k ; Then extract the sentence embedding vector E k The embedded element e in the example is used to calculate the activation change between adjacent embedded elements, which is expressed as checkpoint Δe. Finally, the variability of the activation value of each embedded element between each checkpoint is calculated and compared with the preset threshold. The embedded element with variability higher than the threshold is a special neuron. Alternatively, the variability is arranged from large to small, and the top one or several embedded elements are special neurons.
[0047] The N-1th layer of the hidden layer is the layer closest to the output layer. If a problem is found in the output content of the Nth layer of the hidden layer, it is obviously too late to modify it. Therefore, the present invention operates on the second to last layer of the hidden layer, that is, the N-1th layer, so that it can be modified in time when a problem is found.
[0048] The activation matrix of a large model usually refers to the matrix formed by the activation state of each layer of neurons in the forward propagation process in the neural network. Each layer of neurons will generate an activation value under the action of the input signal, and the activation matrix is a matrix structure that stores these activation values. The main function of the activation matrix is to record the response of the model to different input data. By analyzing the characteristics of the activation matrix, the performance of the model can be understood and optimized, especially when diagnosing the "hallucination" phenomenon in the model, the activation matrix can serve as a key signal source. The activation matrix R composed of the activation characteristics of m×n neurons is expressed as:
[0049]
[0050] m×n is also the dimension of the activation matrix R, m is the column, that is, the width, n is the row, r n,m represents the activation feature of the nth row and mth column, that is, the activation value of the nth neuron in the mth dimension, which is obtained by the activation calculation of the neuron. Therefore, r n,m It represents the output value of a neuron (in the nth row) on a feature dimension (in the mth column). Multiple neurons work together to generate a complete activation matrix.
[0051] In order to analyze the internal state of the model, the present invention converts the activation matrix R of the input (sentence, question, text, etc.) into a sentence embedding vector The conversion formula is as follows:
[0052]
[0053] Where the sentence embedding vector E k It is a column vector obtained after the activation matrix R is column compressed. The penultimate layer of the hidden layer is the layer closest to the output probability. Due to its rich information about the output certainty, it is the main focus of the hallucination analysis of the present invention.
[0054] By converting, the present invention converts the activation matrix R into a vector E that is easier to handle k , can help understand and optimize model performance, especially when diagnosing "hallucination" phenomena in the model. Sentence embedding vector E k It consists of n embedded units, expressed as:
[0055]
[0056] Each embedding element is compressed from the activation features of m neurons, and the i-th embedding element is represented by e i =f(r i,1 ,r i,2 ,…,r i,m ).
[0057] Furthermore, the sentence embedding vector E in the present invention k The activation change between adjacent embedding elements can reflect their oscillatory behavior. Based on this behavior, those embedding elements with significant changes can be extracted. These embedding elements are considered to be special neurons related to the hallucination phenomenon. The checkpoint Δe reflects the state of the model at a certain moment in the training process. By comparing the activation changes of neurons at different checkpoints, it is possible to analyze which neuronal changes are related to the "hallucination" phenomenon. The present invention expresses the activation change between adjacent embedding elements as: Δe = |e i -e i-1 |, where || represents absolute value calculation.
[0058] It should be noted that the checkpoint is calculated by the difference between two adjacent embedding elements, that is, except for the first and last embedding elements, other embedding elements will be interpolated twice (for example, there are m embedding elements in total, and there are m-1 checkpoints after interpolation).
[0059] Special neurons refer to neurons whose activation values fluctuate greatly between multiple training checkpoints during training, indicating high variability during this period.
[0060] The present invention embeds the variability V i It is defined as the degree of fluctuation between each checkpoint, by calculating the i-th embedding element e i Checkpoint Δe i With e i The variance of the means of the two checkpoints involved in the calculation is obtained and is calculated as follows: in is the mean of the checkpoints.
[0061] The present invention measures the variability by variance. Furthermore, if the activation value of an embedding element fluctuates greatly (high variance) and changes significantly (high variance) in the past C checkpoints, the compiled value of the embedding element will be higher. By combining the checkpoints with the variance (fluctuation stability), the variability can be further accurately calculated.
[0062] For example, the present invention defines the embedding units with top 10% variability as special neurons.
[0063] Step 3, discard special neurons and use the model after discarding the embedding element to chat, thereby alleviating the probability of hallucinations in the model from the root.
[0064] Special neurons are important components of the model and can be adjusted to adjust the training process to reduce the hallucination changes of the model and improve the overall confidence when the model is inferring. In essence, the special point is in the sentence embedding vector E k The embedding indices in the training set undergo drastic changes between checkpoints / epochs of training, which we believe is related to the oscillatory behavior in the hallucination performance. Therefore, we choose to discard the embedding points with the top 10% of oscillation amplitudes, as follows: Where M is the embedding vector size, k is the desired percentile threshold, and in this embodiment, k = 10. Since the embedding element is compressed from the neuron, discarding the embedding element is essentially discarding the special neuron.
[0065] The present invention uses SND to control the discard rate and only remove neurons that affect the hallucination phenomenon, rather than blindly reducing the number of neurons overall. This method can improve the accuracy and response speed of user consultation without affecting the overall understanding ability of the model, and can also significantly improve the user experience.
[0066] Figure 1 The specific principle of the present invention is shown. For dialogue generation or content prediction, in the input layer, an input sequence is split into different token1, token2, ..., tokenk. In the hidden layer, each neuron in each layer processes each token, and the generated dialogue result or predicted content is output in the Nth layer. In the second to last layer of the hidden layer, that is, the N-1th layer, the present invention extracts the activation matrix and compresses it to obtain the sentence embedding vector E k , Figure 1 The first embedded element is shown in 1 To the nth embedded element e n , then calculate the checkpoints. In this embodiment, there are 137 values in total. Sort the checkpoints from large to small. The first embedded element e 1 The variability of is ranked sixth. Finally, the top 10% of the checkpoints are selected and represented as special neurons. After deleting these 14 special neurons, the remaining embedding elements are obtained. In this embodiment, the first two embedding elements are deleted, and the hidden layer architecture obtained is as follows Figure 1 As shown on the right, the model is finally used for chatting to alleviate the hallucination phenomenon.
[0067] According to the present invention, for example, one of the most popular questions in the big model in 2024 is: "Which is bigger, 9.9 or 9.11?" The most popular and most models on the market will reply that 9.11 is bigger, and give the following explanation: "In comparing numbers, if the integer parts are the same, we will compare the decimal parts. Although 9.11 has two decimal parts and 9.9 has only one decimal part, 0.11 is greater than 0.9. Therefore, 9.11 is bigger."
[0068] This example shows that the model's previous reasoning is correct, but in the subsequent comparison, the model mistakenly converts the decimal comparison into an integer comparison, resulting in an incorrect result. After the correction of the present invention, the model's answer should be: "9.9 is greater". Specifically, the integer parts of 9.11 and 9.9 are both 9, so they are equal and no further comparison is required. Next, compare their decimal parts: the tenth digit of the decimal part of 9.11 is 1, while the tenth digit of the decimal part of 9.9 is 9. Since 9 is greater than 1, 9.9 is greater than 9.11. Therefore, the present invention can effectively solve the problem of hallucinations in the model during the reasoning process.
Claims
1. A method for alleviating chat hallucination based on large models of SND, characterized in that: The steps include: Step 1, obtaining chat sentence samples as input of a large model, wherein the large model is based on a neural network architecture and includes an input layer, a hidden layer and an output layer, wherein the hidden layer has a total of N layers; Step 2: In the N-1th layer of the hidden layer, the activation matrix R formed in the previous layer during the forward propagation process is converted into a sentence embedding vector E k ; Step 3: Extract the sentence embedding vector E k The activation change between adjacent embedded elements is obtained by difference calculation, which is expressed as checkpoint Δe; Step 4: Calculate the variability of the activation values of each embedding element between each checkpoint, sort the variability from large to small, and discard one or several embedding elements with the highest ranking; Step 5: Chat using the model after discarding the embedding meta-data.
2. According to claim 1, the method for alleviating the large model chat hallucination based on SND is characterized in that: In step 2, the activation matrix R is composed of activation features of multiple neurons, with a dimension of is m×n, expressed as: m and n are the columns and rows of the activation matrix R, respectively. n,m represents the activation feature of the nth row and mth column; The sentence embedding vector is a column vector obtained by column compression of the activation matrix R, consisting of n embedded elements, expressed as: Each embedding element is compressed from the activation features of m neurons, and the i-th embedding element is represented by e i =f(r i,1 ,r i,2 ,…,r i,m ).
3. According to claim 2, the method for alleviating the large model chat hallucination based on SND is characterized in that: The activation matrix R is compressed to obtain the sentence embedding vector The compression formula is expressed as: In the formula, represents the activation value of the i-th neuron in the N-1th layer of the hidden layer, Represents the activation value of the mth neuron in the N-1th layer of the hidden layer.
4. The method for alleviating the large-model chat illusion based on SND according to claim 2 or 3, characterized in that: In step 3, the activation change between adjacent embedding elements is expressed as: Δe=|e i -e i-1 | In the formula, || represents absolute value calculation.
5. According to claim 4, the method for alleviating the large model chat illusion based on SND is characterized in that: In step 4, the i-th embedding element e i Variability V i It is defined as the degree of fluctuation between each checkpoint, and is calculated by e i Checkpoint Δe i With e i The variance of the means of the two checkpoints involved in the calculation is obtained and is calculated as follows: is the mean of the checkpoints.
6. The method for alleviating the large-model chat hallucination based on SND according to claim 1, characterized in that: In step 4, the top 10% of the embedding elements are discarded.
Citation Information
Cited By
Systems and methods for generating and deploying specialized expert small models
US20260178576A1