Self-iterative Training Method for Large Language Model Hallucination Detector Based on Self-consistent Voting
Through the self-iteration training method of hallucination detectors based on self-consistent voting, the problem of difficulty in improving the scale and detector accuracy of hallucination data sets in the prior art is solved, and the cycle process of data set amplification and detector accuracy is realized, and the accuracy and performance of hallucination detectors are improved.
Patent Information
- Application Number
- CN202410842355.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-06-27
AI Technical Summary
In the prior art, the training method of the large language model hallucination detector is difficult to simultaneously improve the scale of the hallucination data set and the detector accuracy, resulting in weak hallucination detection capabilities.
The self-iteration training method of the large language model hallucination detector based on self-consistent voting is adopted. By obtaining the original hallucination data set, converting it into a preset format, conducting multi-stage training, and adding new data to the data set by self-consistent voting, the cyclical process of data set amplification and detector accuracy improvement is realized.
The data set scale and detector accuracy are achieved simultaneously, solving the problem that the simultaneous improvement of the simultaneous improvement of the simultaneous precision of the simultaneous improvement of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of the simultaneous development of
Smart Images

Figure CN118839134B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of model hallucination detection, and in particular to a self-iterative training method for a large language model hallucination detector based on self-consistent voting. Background Art
[0002] Hallucination detection of large language models refers to using certain technical means to detect hallucination phenomena in the responses of large language models. These hallucinations may be misleading information, inaccurate descriptions, misunderstandings or ambiguities, etc., which may have a negative impact on the understanding and dissemination of information. Currently, large language models generally have the problem of "hallucinations", that is, when answering user questions, especially when answering questions that require a large amount of knowledge, the model will generate information that sounds credible but is not true or meaningless, which greatly hinders the application of large language models in the real world.
[0003] The disadvantages of the existing technologies are all reflected in the problem of weak hallucination detection ability. The internal reason is the lack of a relatively large-scale training dataset. However, since the cost of pure manual annotation is unaffordable, building a large hallucination dataset requires a powerful hallucination detector. Thus, a dead loop is formed, that is, if the scale of the dataset needs to be expanded, a stronger hallucination detector is required, and to train a stronger hallucination detector, a larger dataset is required.
[0004] In summary, there is currently a lack of a training method for large language model hallucination detectors to solve or partially solve the problem that it is difficult to simultaneously improve the scale of the hallucination dataset and the accuracy of the detector. Summary of the Invention
[0005] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a self-iterative training method for a large language model hallucination detector based on self-consistent voting to solve or partially solve the problem that it is difficult to simultaneously improve the scale of the hallucination dataset and the accuracy of the detector.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] One aspect of the present invention provides a self-iterative training method for a large language model hallucination detector based on self-consistent voting, including the following steps:
[0008] Obtain the original hallucination dataset and convert it into a preset data format, and perform the first-stage training on the large language model hallucination detector based on the converted hallucination dataset;
[0009] Use the problem information in the input data as the input for multiple different large language models, collect the corresponding answer information, construct new input data, implement the annotation of the new input data based on self-consistent voting and add it to the hallucination dataset, and conduct the second stage of training on the large language model hallucination detector based on the current hallucination dataset;
[0010] Use multiple question information generated based on multiple topics as the input for multiple different large language models, collect the corresponding answer information, construct new input data, implement the annotation of the new input data based on self-consistent voting and add it to the hallucination dataset, and conduct the third stage of training on the large language model hallucination detector based on the current hallucination dataset.
[0011] As a preferred technical solution, the annotation process of the input data includes the following steps:
[0012] For each unannotated input data, generate multiple candidate output data in the form of triples based on the pre-constructed prompt words;
[0013] For the candidate output data, screen and obtain the final output data in the form of triples through self-consistent voting to achieve the annotation of the hallucination dataset,
[0014] Among them, the output data in the form of triples includes factual information, reference point information, and hallucination type information.
[0015] As a preferred technical solution, the process of screening and obtaining the final output data in the form of triples through self-consistent voting includes the following steps:
[0016] Through majority voting, preliminarily screen the candidate output data corresponding to the hallucination type with the largest proportion;
[0017] For the candidate output data after preliminary screening, select the candidate output data with the semantic similarity closest to the average semantic similarity of the reference points of all current candidate output data as the final output data.
[0018] As a preferred technical solution, the hallucination types include no hallucination, unverifiable, no fact, and contradiction.
[0019] As a preferred technical solution, the input data includes question information, answer information, and reference information.
[0020] As a preferred technical solution, the process of generating candidate output data in the form of triples includes the following steps:
[0021] Based on the question information and answer information in the input data, judge whether there are objective information points. If not, label the hallucination type of the input data as no fact and end the annotation. If so, generate factual information;
[0022] Extract reference key point information associated with the question information and answer information from the reference information of the input data;
[0023] Based on the reference key point information, determine the hallucination type, generate hallucination type information, and complete the generation of candidate output data in the form of triples.
[0024] As a preferred technical solution, the process of using the question information in the input data as the input of multiple different large language models and collecting the corresponding answer information includes the following steps:
[0025] Input the question information in the input data into multiple different large language models, and collect the answer information in the scenarios of RAG and non-RAG.
[0026] As a preferred technical solution, the original hallucination data set is the ANAH data set.
[0027] Another aspect of the present invention provides an electronic device, including: one or more processors and a memory, wherein the memory stores one or more programs, and the one or more programs include instructions for executing the aforementioned self-iterative training method of the large language model hallucination detector based on self-consistent voting.
[0028] Another aspect of the present invention provides a computer-readable storage medium, including one or more programs for execution by one or more processors of an electronic device, and the one or more programs include instructions for executing the aforementioned self-iterative training method of the large language model hallucination detector based on self-consistent voting.
[0029] Compared with the prior art, the present invention has at least one of the following beneficial effects:
[0030] (1) Achieve simultaneous improvement of the data set scale and detector accuracy: Through the self-iterative training method of the cycle of data - training model - using the new model to label new data - using the new data to update the model, the present invention realizes the amplification of the hallucination data set on the one hand, and on the other hand, trains based on the amplified hallucination data set, effectively ensuring the accuracy and performance of the hallucination detector, and solving the problem that it is difficult to simultaneously improve the data set scale and detector accuracy.
[0031] (2) High scalability: By adopting the self-iterative training method, the present invention combines the two processes of data annotation and training, can expand the data set while training, and effectively improves the scalability and training efficiency.
[0032] (3) High annotation accuracy: The present invention adopts an inference annotation method based on self-consistent voting, which conducts a preliminary screening based on voting and a re-screening based on self-consistency check for the candidate output data, and obtains the final output data by comparing semantic similarities, which can effectively improve the annotation accuracy. Description of the Drawings
[0033] Figure 1 It is a schematic diagram of the self-iterative training method of the large language model hallucination detector based on self-consistent voting in the embodiment;
[0034] Figure 2 It is a schematic diagram of the prompt words used in the fact existence check stage in the embodiment;
[0035] Figure 3 It is a schematic diagram of the prompt words used in the reference information extraction stage in the embodiment;
[0036] Figure 4 It is a schematic diagram of the prompt words used in the hallucination type judgment stage in the embodiment;
[0037] Figure 5 It is a schematic diagram of the self-iterative training of the hallucination annotator in the embodiment. Detailed Embodiment
[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0039] The terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, "a plurality" means two or more, unless otherwise specifically defined.
[0040] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitations, the element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the said element.
[0041] Example 1
[0042] Aiming at the problem that it is difficult to improve the scale of the hallucination dataset and the accuracy of the detector simultaneously, this embodiment provides a self-iterative training method for the large language model hallucination detector based on self-consistent voting. By adopting a self-iterative training mode, a cycle of "labeling data - training the model - labeling new data with the new model - training and updating the model with the new data" is constructed to improve the accuracy of the detector while expanding the dataset.
[0043] See Figure 1 , this method mainly includes three parts: the construction of the hallucination annotation process (corresponding to the right side of Figure 1 ), the hallucination annotation inference (corresponding to the left side of Figure 1 ), and the self-iterative training of the hallucination annotator (corresponding to the middle of Figure 1 ). The details of each part are as follows.
[0044] (1) Construction of the hallucination annotation process: The hallucination annotation process refers to providing a reference document and the information to be annotated. The model needs to combine the factual content in the reference document to give the hallucination type of the information to be annotated. In this embodiment, the types of hallucinations are divided into four categories: no hallucination, cannot be verified, no fact, and contradiction. By deconstructing the entire annotation process, this embodiment provides a new annotation method that is more in line with human understanding. This part divides the hallucination annotation process into three steps, namely: fact existence check, reference information extraction, and hallucination type judgment. The data format of the original hallucination dataset is converted through these three steps. For specific reference, see the right part of Figure 1
[0045] In the fact existence check stage, by providing the information to be annotated, based on the question and the reply to be annotated, the annotator is required to judge whether there are any facts that can be judged, that is, whether there are specific objective information points that can be verified through data, research results, or other reliable sources. If there are no facts that can be judged, the hallucination type of this information is directly labeled as "no fact", and then the entire annotation process terminates; if there are facts, the subsequent annotation continues. The specific prompt words are as shown in Figure 2 .
[0046] In the reference information extraction stage, based on the provided reference document, question, and the reply to be annotated, the annotator is required to extract the reference points related to the question and the reply to be annotated from the reference document. If no relevant information can be found in the document, the hallucination type is directly labeled as "cannot be verified", and then the entire annotation process terminates; otherwise, the found reference points are output, and the subsequent annotation continues. The specific prompt words are as shown in 3.
[0047] In the hallucination type judgment stage, the reference points obtained in the previous step are used, and the annotator is required to judge the hallucination type of the information to be annotated based on this. If the sentence contains factual information and is consistent with the reference information, its type is "no hallucination". If the sentence contradicts the reference document, its type is "contradictory hallucination". If the sentence lacks supporting evidence and cannot be verified, its type is "unverifiable hallucination". Specific prompts are as Figure 4 shown.
[0048] (2) Hallucination annotation reasoning: This embodiment provides a reasoning method based on self-consistent voting, which can effectively improve the annotation accuracy. Specifically, for a triple input (question, answer, reference document), the model will generate a triple output (factual information, reference points, hallucination type). By using the top-k sampling method, the model generates multiple candidate outputs for an input, and then conducts self-consistent voting to select the final hallucination annotation result.
[0049] Specifically, in the process of self-consistent voting, first use the majority voting method to screen the hallucination type and select the hallucination type with the largest proportion. If the selected hallucination type is without facts, there is no need to continue voting, that is, it is determined that the annotation information of the current input is without facts; otherwise, continue with the self-consistency check and screen the samples with this hallucination type in all candidate outputs again. Then, by comparing the semantic similarity, select the most representative reference points (that is, the sample whose semantic similarity is closest to the average semantic similarity of all candidate reference points). Since the factual information is strongly correlated with the reference points and the hallucination type, no additional processing is performed. Finally, the final output is obtained.
[0050] (3) Self-iteration of the hallucination annotator: This embodiment provides a self-iteration training scheme for the hallucination annotator. Through continuous cycles of data annotation and training, while expanding the scale of the hallucination dataset, the accuracy of the hallucination annotator is improved.
[0051] The whole process is as Figure 4 shown, including three training stages.
[0052] In the first stage, a relatively small hallucination dataset annotated by humans is used. Specifically, in this embodiment, ANAH is adopted. It should be noted that, on the premise of no conflict, other existing datasets can also be used for substitution, and the training method in part (1) is used to train a relatively weak annotator.
[0053] In the second stage, expansion is carried out in the dimension of the model responses. Using the same questions as in the first stage, responses of 13 open-source models of different scales and series are collected. For each model, responses are collected in both RAG (Retrieval-Augmented Generation for AI content generation) and non-RAG scenarios. Then, the responses are annotated using the annotator in the first stage, and the annotated data is added to the dataset for training a stronger annotator.
[0054] In the third stage, expansion is carried out in the dimension of topics. By collecting information on more topics, some questions are generated for each topic, and the responses of each model are collected using the same configuration as in the second stage, and the responses are annotated using the annotator obtained in the second stage. Finally, the new data is added to the dataset to train the final version of the annotator. Specifically, in this stage, Google's Ngram Viewer is used to automatically select topics according to the occurrence frequency. The selected topic categories are mainly divided into four major categories: celebrities, events, locations, and things, and cover multiple fields, such as politics, art, nature, economy, health, and sports, etc.
[0055] After the above hallucination annotation data construction process, a large hallucination dataset is obtained, which contains more than 3k topics, 196k model responses, and 822k sentence-level hallucination annotations. At the same time, a large language model hallucination annotator with extremely high accuracy is also obtained.
[0056] It should be noted that the training method of using "labeled data - training the model - using the new model to label new data - using the new data to train and update the model" for hallucination detector training is theoretically a variant design of the present invention.
[0057] Compared with some solutions, the methods used to generate hallucination data are to use special leading prompts to induce the large language model to generate hallucinations in the responses, resulting in the inability to reflect the hallucination level of the model under normal use. The hallucination data in this embodiment completely comes from the natural responses of the large language model in normal Q&A, making our dataset distribution more suitable for the training and inference of the large language model.
[0058] Compared with some solutions, the hallucination annotation process directly allows the model to output reference information, hallucination types, and correction measures within one round of conversation. Such a mixed task mode is not conducive to model training due to the lack of internal logical derivation, and it cannot clearly show the relationship between the reference point and hallucination judgment, resulting in unsatisfactory annotation accuracy. The annotation method in this embodiment is closer to the human judgment mode.
[0059] Compared with some solutions that use relatively small-scale training datasets and lack effective expansion methods, it is very difficult to further improve the detection ability of the model. The self-iterative training method of this embodiment can expand the scale of training data while continuously improving the accuracy of the detector, and has good scalability.
[0060] To verify the effectiveness of this method, hallucination detector training was carried out based on the Internlm2-7B model, and the performance of the hallucination detector at different stages was tested, which proved the effectiveness of this training scheme.
[0061] Referring to Table 1, the test set in ANAH was used to test the level of the hallucination detector at each stage. The prediction accuracy ACC and F1 were used as indicators, and the current state-of-the-art model GPT4 was compared.
[0062] Among them, Stage1, 2, and 3 respectively represent different iterative stages, that is, the first stage, the second stage, and the third stage in part (3).
[0063] Table 1 Training Results of Hallucination Detector
[0064] Model F1 ACC GPT4 87.11 86.97 Stage1 84.45 84.85 Stage2 87.75 88.18 Stage3 89.30 89.55
[0065] It can be found that after using the self-iterative training method of this embodiment, the detection accuracy of the model at each stage has been steadily increasing, and it surpassed the current strongest model GPT4 at the second stage, and reached the current highest prediction accuracy of 89.55% and F1 score of 89.30% at the third stage.
[0066] In summary, this method provides an effective method for expanding the hallucination dataset to address the current widespread lack of fine-grained training corpora for large language model hallucination detectors, which can efficiently expand the dataset at a relatively low cost. To address the problem of the weak ability of current hallucination detection models, a self-iterative training framework for large language model hallucination detectors is provided, which can effectively improve the hallucination detection level of large models.
[0067] Embodiment 2
[0068] This embodiment provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the self-iterative training method of the large language model hallucination detector based on self-consistent voting as described in Embodiment 1.
[0069] In a typical configuration, an electronic device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0070] The memory may include non - permanent memory in the form of computer - readable media, such as random access memory (RAM) and / or non - volatile memory, such as read - only memory (ROM) or flash RAM. The memory is an example of computer - readable media.
[0071] Embodiment 3
[0072] This embodiment provides a computer - readable storage medium, including one or more programs for execution by one or more processors of an electronic device, and the one or more programs include instructions for performing the self - iterative training method of the large - language model hallucination detector based on self - consistent voting as described in Embodiment 1.
[0073] Computer - readable media includes both permanent and non - permanent, removable and non - removable media and can implement information storage by any method or technology. The information can be computer - readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase - change memory (PRAM), static random - access memory (SRAM), dynamic random - access memory (DRAM), other types of random - access memory (RAM), read - only memory (ROM), electrically erasable programmable read - only memory (EEPROM), flash memory or other memory technologies, compact disc read - only memory (CD - ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non - transitory media that can be used to store information accessible by a computing device. As defined herein, computer - readable media does not include transitory computer - readable media, such as modulated data signals and carrier waves.
[0074] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A self-iterative training method for a large language model hallucination detector based on self-consistent voting, characterized in that: The steps include: Obtain the original hallucination dataset and convert it into a preset data format, and perform the first phase of training on the large language model hallucination detector based on the converted hallucination dataset; The question information in the input data is used as the input of multiple different large language models, the corresponding answer information is collected, and new input data is constructed. The new input data is annotated based on self-consistent voting and added to the hallucination dataset of the first stage. The large language model hallucination detector is trained in the second stage based on the current hallucination dataset. Multiple question information generated based on multiple topics is used as the input of multiple different large language models, the corresponding answer information is collected, new input data is constructed, and the new input data is annotated based on self-consistent voting and added to the second-stage hallucination dataset. The third-stage training of the large language model hallucination detector is carried out based on the current hallucination dataset.
2. According to the self-consistent voting-based large language model hallucination detector self-iterative training method of claim 1, characterized in that: The input data annotation process includes the following steps: For each unlabeled input data, multiple candidate output data in the form of triples are generated based on pre-constructed prompt words; For the candidate output data, the final output data in the form of triples is obtained through self-consistent voting screening to achieve the annotation of the hallucination data set. The output data in triple form includes factual information, reference point information and hallucination type information.
3. According to claim 2, a self-iterative training method for a large language model hallucination detector based on self-consistent voting is characterized in that: The final output data in the form of triples obtained by self-consistent voting screening comprises the following steps: Through majority voting, the candidate output data corresponding to the most common hallucination types are preliminarily screened; For the candidate output data after preliminary screening, the candidate output data whose semantic similarity is closest to the average semantic similarity of the reference points of all current candidate output data is selected as the final output data.
4. According to claim 3, a self-iterative training method for a large language model hallucination detector based on self-consistent voting is characterized in that: The types of hallucinations described include no hallucination, unverifiable, no fact, and contradiction.
5. The self-iterative training method of a large language model hallucination detector based on self-consistent voting according to claim 2 is characterized in that: The input data includes question information, answer information and reference information.
6. The self-iterative training method of a large language model hallucination detector based on self-consistent voting according to claim 5, characterized in that: The process of generating candidate output data in the form of triples includes the following steps: Based on the question information and answer information in the input data, determine whether there is an objective information point. If not, mark the input data illusion type as no fact and end the marking. If yes, generate fact information; Extracting reference key information associated with question information and answer information from reference information of input data; Based on the reference key point information, the hallucination type is determined, hallucination type information is generated, and the generation of candidate output data in the form of triples is completed.
7. The self-iterative training method of a large language model hallucination detector based on self-consistent voting according to claim 1, characterized in that: The process of using the question information in the input data as the input of multiple different large language models and collecting corresponding answer information includes the following steps: The question information in the input data is input into multiple different large language models to collect answer information in RAG and non-RAG scenarios.
8. The self-iterative training method of a large language model hallucination detector based on self-consistent voting according to claim 1, characterized in that: The original hallucination dataset is the ANAH dataset.
9. An electronic device, characterized in that: include: One or more processors and a memory, wherein the memory stores one or more programs, wherein the one or more programs include instructions for executing the self-iterative training method of a large language model hallucination detector based on self-consistent voting as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that: It includes one or more programs for execution by one or more processors of an electronic device, and the one or more programs include instructions for executing the self-iterative training method of a large language model hallucination detector based on self-consistent voting as described in any one of claims 1-8.
Citation Information
Patent Citations
Label value determination method and device, equipment and storage medium
CN117574286A
Parameter Efficient Prompt Tuning for Efficient Models at Scale
US20230325725A1