Method and device for screening instruction-following data for large language model hallucination phenomenon
By performing semantic consistency detection and clustering on the responses generated by the large language model, high-quality instruction-following data is screened out, which solves the model's hallucination problem during the instruction fine-tuning stage, and enables the model to effectively follow instructions and improve training efficiency.
Patent Information
- Application Number
- CN202510309563.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-03-17
AI Technical Summary
Large language models in existing technologies are prone to hallucinations during the instruction fine-tuning stage, and existing methods cannot effectively solve the dilemma of ensuring that the model follows instructions while reducing hallucinations. In particular, reinforcement learning methods are inefficient and ignore the model's ability to follow instructions.
By extracting the response set generated by a large language model, detecting semantic consistency and clustering responses, and screening out semantically equivalent response clusters, high-quality instruction-following data is determined based on semantic consistency and the proportion of responses within the cluster, which is used to train the model to reduce hallucinations.
It significantly reduces the hallucination phenomenon of large language models while maintaining the model's ability to follow instructions, improving training efficiency and model credibility.
Smart Images

Figure CN120256623B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and device for screening instruction-following data targeting large language model hallucination phenomena. Background Art
[0002] Alignment is a critical process for ensuring that large language models (LLMs) follow user instructions. A practical alignment approach is to fine-tune instructions on instruction data. However, state-of-the-art LLMs sometimes generate sentences that appear plausible but are actually incorrect, a phenomenon known as hallucination. This hallucination can undermine the credibility of LLMs in real-world applications.
[0003] Training LLMs with data containing unknown knowledge during the instruction fine-tuning phase may lead to overconfidence in the model and promote hallucinations. That is, if the knowledge in the instruction data is not learned during pre-training, the fine-tuned LLMs tend to make more errors when generating responses.
[0004] Therefore, there is a dilemma in instruction fine-tuning: on the one hand, LLMs need to learn how to follow user instructions at this stage, which is crucial for user interaction in real-world applications; on the other hand, using high-quality data (either manually labeled or generated by other state-of-the-art LLMs) for instruction fine-tuning may introduce knowledge that is unfamiliar to LLMs and may thus promote hallucinations.
[0005] How to align LLMs during the instruction fine-tuning stage so that they can follow instructions and reduce hallucinations? In the related art, some studies apply reinforcement learning (RL) to teach LLMs to reduce hallucinations after the instruction fine-tuning stage. For example, some studies take advantage of the self-assessment ability of LLMs and use GPT-3.5-turbo to create preference data, and then align LLMs through direct preference optimization (DPO). However, these methods require additional corpora and API costs of closed-source LLMs, making them inefficient. Such reinforcement learning methods ignore the potential impact on the model's ability to follow instructions. Unlike RL methods, an intuitive strategy is to screen out unfamiliar instruction data for instruction fine-tuning. These selected high-quality data may present more unknown knowledge to the LLM, thereby further promoting hallucinations, because these data may contain expert-level knowledge responses and often go deep into high-level details. As a result, the methods in the related art cannot effectively solve the hallucinations of large language models. Summary of the Invention
[0006] The main purpose of the present invention is to provide a method and device for filtering instruction-following data for the large language model hallucination phenomenon, so as to solve the shortcomings existing in the related art.
[0007] In order to achieve the above-mentioned purpose, according to a first aspect of the present invention, there is provided a method for screening instruction-following data for the large language model hallucination phenomenon, comprising: for any instruction, extracting the reply corresponding to the any instruction to obtain a reply set, wherein the reply corresponding to the any instruction is generated by a large language model based on the any instruction; detecting the semantic consistency of the replies in the reply set to obtain a semantic consistency detection score in the embedding space; clustering the replies in the reply set to obtain semantic clusters, and performing semantic equivalence detection on the target reply corresponding to the any instruction and the replies in the reply set, to determine the target semantic cluster to which the target reply belongs; determining a second score based on the number of replies generated in the target semantic cluster and the number of all replies generated; and screening out instruction-following data corresponding to any instruction based on the semantic consistency detection score and the second score.
[0008] Optionally, for any instruction, extract the reply corresponding to the any instruction; detect the semantic consistency of the reply, and obtain the semantic consistency detection score in the embedding space, including: for any instruction r, extract the reply [r′_1, r′_2, ..., r′_K] corresponding to the any instruction, and obtain a reply set, wherein K represents the quantity, and [r′_1, r′_2, ..., r′_K] is generated based on a large language model; determine the final sentence embedding for the reply, wherein the final sentence embedding E = [e_1, e_2, ..., e_K] is determined for the internal state of the K generated replies and the last token of the last layer of the large language model; and after processing the final sentence embedding into a multivariate Gaussian distribution, determine the differential entropy, wherein the differential entropy is used as a semantic consistency detection score to measure the semantic consistency in the continuous embedding space.
[0009] Optionally, after processing the final sentence embedding into a multivariate Gaussian distribution, determining the differential entropy includes: processing the sentence embedding E into a multivariate Gaussian distribution E~N(μ, ∑); based on Determine the differential entropy, where det(∑) represents the determinant of the covariance matrix Σ and d is the dimension of the sentence embedding; Σ captures the relationship between K different sentence embeddings and is expressed as: The differential entropy is simplified to Among them, λ i represents the i-th eigenvalue of the covariance matrix ∑; G is a constant; and q is any of the instructions.
[0010] Optionally, clustering the replies in the reply set to obtain semantic clusters includes:
[0011] The replies are clustered using a pre-trained natural language inference model to obtain different semantic clusters.
[0012] Optionally, after performing semantic equivalence detection on the target reply corresponding to any of the instructions and the replies in the reply set, determining the target semantic cluster to which the target reply belongs includes: reasoning on the target reply and the replies in each semantic cluster based on the natural language inference model to determine whether the target reply and the replies in the semantic cluster are semantically equivalent; if they are equivalent, determining the semantic cluster to which the target reply belongs based on majority voting.
[0013] Optionally, determining the second score based on the number of replies generated in the target semantic cluster and the number of all generated replies includes: taking the ratio of the number of replies generated in the target semantic cluster to the number of all generated replies as the second score.
[0014] Optionally, filtering out instruction compliance data corresponding to any instruction based on the semantic consistency detection score and the second score includes: determining a final scoring result based on the ratio of the first score to the second score, wherein the scoring result is used as a basis for filtering the instruction compliance data of any instruction.
[0015] Optionally, the final scoring results calculated based on the reply sets corresponding to any instruction extracted at different times are sorted, and the top N reply sets with the highest sorting are used as instruction compliance data.
[0016] According to a second aspect of the present invention, there is provided an instruction-following data screening device for the large language model hallucination phenomenon, comprising a semantic consistency detection unit for extracting, for any instruction, a reply corresponding to the any instruction to obtain a reply set, wherein the reply corresponding to the any instruction is generated by a large language model based on the any instruction; detecting the semantic consistency of the replies in the reply set to obtain a semantic consistency detection score in the embedding space; a semantic equivalence detection unit for clustering the replies in the reply set to obtain semantic clusters, and determining a target semantic cluster to which the target reply belongs after performing a semantic equivalence detection on a target reply corresponding to the any instruction and the replies in the reply set; determining a second score based on the number of replies generated in the target semantic cluster and the number of all replies generated; and a screening unit for screening out instruction-following data corresponding to any instruction based on the semantic consistency detection score and the second score.
[0017] According to a third aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute any one of the methods described in the first aspect.
[0018] This embodiment provides a method and apparatus for filtering instruction compliance data for the hallucination phenomenon of a large language model, wherein the method includes extracting the replies corresponding to any instruction to obtain a reply set, wherein the replies corresponding to any instruction are generated by a large language model based on the instruction; detecting the semantic consistency of the replies in the reply set to obtain a semantic consistency detection score in the embedding space; clustering the replies in the reply set to obtain semantic clusters, and performing semantic equivalence detection on the target reply corresponding to any instruction and the replies in the reply set to determine the target semantic cluster to which the target reply belongs; determining a second score based on the number of replies generated in the target semantic cluster and the number of all replies generated; and filtering instruction compliance data corresponding to any instruction based on the semantic consistency detection score and the second score. The instruction compliance data is filtered through internal state consistency detection and semantic equivalence detection, and the filtered high-quality compliance data can be used for training a large-scale language model, so that the trained large-scale language model significantly reduces hallucinations while maintaining a strong instruction compliance capability. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 This is a flow chart of a method for screening instruction-compliant data for large language model hallucination phenomena according to an embodiment of the present invention;
[0021] Figure 2 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0022] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0023] It should be noted that the terms "first," "second," and the like in the specification and claims of the present invention and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate for the embodiments of the present invention described herein. In addition, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatuses.
[0024] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0025] According to an embodiment of the present invention, a method for filtering instruction-following data for the large language model hallucination phenomenon is provided. Figure 1 As shown, it includes the following steps 101 to 103:
[0026] Step 101: For any instruction, extract the reply corresponding to the instruction to obtain a reply set, wherein the reply corresponding to the instruction is generated by a large language model based on the instruction; detect the semantic consistency of the replies in the reply set to obtain a semantic consistency detection score in the embedding space.
[0027] In this step, semantic consistency is measured in the dense sentence embedding space via internal consistency test (ICP).
[0028] As an optional implementation method of this embodiment, for any instruction, extract the reply corresponding to the any instruction; detect the semantic consistency of the reply, and obtain the semantic consistency detection score in the embedding space, including: for any instruction r, extract the reply [r′_1, r′_2, ..., r′_K] corresponding to the any instruction to obtain a reply set, where K represents the quantity, and [r′_1, r′_2, ..., r′_K] is generated based on a large language model; determine the final sentence embedding for the reply, where the final sentence embedding E = [e_1, e_2, ..., e_K] is determined for the internal state of the K generated replies and the last token of the last layer of the large language model; and after processing the final sentence embedding into a multivariate Gaussian distribution, determine the differential entropy, where the differential entropy is used as a semantic consistency detection score to measure the semantic consistency in the continuous embedding space.
[0029] In this optional implementation, for a given instruction data s = (q, r), q represents the instruction and r represents the corresponding target reply. For the instruction q, multiple replies [r′_1, r′_2, ..., r′_K] are extracted, where K represents the number of samples. These replies are generated by a basic LLM. For the generated K replies [r′_1, ..., r′_K], the internal state of the last token of each reply in the last layer is used as the final sentence embedding E = [e_1, e_2, ..., e_K}, because it effectively captures the semantics of the sentence. Differential entropy (DE) is further used to measure the semantic consistency in the continuous embedding space:
[0030] As an optional implementation of this embodiment, after the final sentence embedding is processed into a multivariate Gaussian distribution, determining the differential entropy includes: for the sentence embedding E, processing it into a multivariate Gaussian distribution E~N(μ, ∑); based on Determine the differential entropy, where det(∑) represents the determinant of the covariance matrix ∑ and d is the dimension of the sentence embedding; ∑ captures the relationship between K different sentence embeddings and is expressed as: Differential entropy simplifies to Among them, λ i represents the i-th eigenvalue of the covariance matrix ∑; G is a constant; and q is any of the instructions.
[0031] In this optional implementation, for the sentence embedding E, the aforementioned DE is processed into a multivariate Gaussian distribution E~N(μ,∑), so the differential entropy can be expressed as:
[0032]
[0033] Where det(∑) represents the determinant of the covariance matrix ∑, and d is the dimension of the sentence embedding (e.g., 4096 for LLaMA-3-8B). ∑ captures the relationship between K different sentence embeddings and can be expressed as The formula simplifies to Among them, λ i It represents the i-th eigenvalue of the covariance matrix Σ, which can be easily calculated by singular value decomposition. G is a constant. Finally, the internal state consistency is measured by this formula. For a given instruction q in the data S, it is defined as F ins (q).
[0034] Step 102: Cluster the replies in the reply set to obtain semantic clusters, and after performing semantic equivalence detection on the target reply corresponding to any instruction and the replies in the reply set, determine the target semantic cluster to which the target reply belongs; determine the second score based on the number of replies generated in the target semantic cluster and the number of all replies generated.
[0035] In this step, we analyze the LLM's familiarity with the target response at the lexical level. We first attempt to cluster the generated responses that express the same content. This is because instruction data responses are usually free-form, and multiple responses may express the same meaning in different ways.
[0036] As an optional implementation of this embodiment, clustering the replies in the reply set to obtain semantic clusters includes: clustering the replies using a pre-trained natural language inference model to obtain different semantic clusters.
[0037] In this optional implementation, a trained Natural Language Inference (NLI) model can be used to test each pair of generated responses to cluster them. NLI models are well-suited for identifying semantic equivalence, because if one sentence can be inferred from another, then the two generated responses mean the same thing. Using a small NLI model as a clustering model, two responses that can imply each other are considered semantically equivalent.
[0038] As an optional implementation of this embodiment, after performing semantic equivalence detection on the target reply corresponding to any one of the instructions and the replies in the reply set, determining the target semantic cluster to which the target reply belongs includes: reasoning on the target reply and the replies in each semantic cluster based on the natural language inference model to determine whether the target reply and the replies in the semantic cluster are semantically equivalent;
[0039] If they are equal, the semantic cluster to which the target reply belongs is determined based on majority voting.
[0040] In this optional implementation, the generated replies are clustered using the NLI model to obtain multiple semantic clusters. Subsequently, the NLI model is again applied to each generated reply to determine if the generated reply and the target reply are semantically equivalent. A majority vote can be used to determine the semantic cluster to which the target reply belongs.
[0041] Step 103: Filter out instruction compliance data corresponding to any instruction based on the semantic consistency detection score and the second score.
[0042] In this step, the data is finally sorted from large to small according to the weighted scores, and a subset of a specified proportion (for example, the top 10%) can be selected as the final training data.
[0043] As an optional implementation method of this embodiment, determining the second score based on the number of replies generated in the target semantic cluster and the number of all replies generated includes: taking the ratio of the number of replies generated in the target semantic cluster to the number of all replies generated as the second score.
[0044] In this optional implementation, the ratio of the number of generated replies contained in the semantic cluster to the total number of generated replies is used as the final score F res (r).
[0045] As an optional implementation of this embodiment, screening out instruction compliance data corresponding to any instruction based on the semantic consistency detection score and the second score includes: determining a final scoring result based on the ratio of the first score to the second score, wherein the scoring result is used as a basis for screening the instruction compliance data of the any instruction.
[0046] As an optional implementation method of this embodiment, the final scoring results calculated based on the reply sets corresponding to any instruction extracted at different times are sorted, and the top N reply sets with the highest ranking are used as instruction compliance data.
[0047] In the above optional implementation, both scores are considered and the weighted score is used as the final scoring result: Using this score as the final sorting criterion, the dataset to be screened is sorted from largest to smallest, and a subset of a specified proportion (e.g., the top 10%) can be selected as the final training data.
[0048] This embodiment screens instruction-following data through internal state consistency testing and semantic equivalence testing. The high-quality instruction-following data screened by this method can be used to train large-scale language models, significantly reducing hallucinations while maintaining strong instruction-following capabilities.
[0049] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0050] According to an embodiment of the present invention, there is also provided an instruction compliance data screening device for the large language model hallucination phenomenon, comprising a semantic consistency detection unit, for extracting, for any instruction, a reply corresponding to the any instruction to obtain a reply set, wherein the reply corresponding to the any instruction is generated by a large language model based on the any instruction; detecting the semantic consistency of the replies in the reply set to obtain a semantic consistency detection score in the embedding space; a semantic equivalence detection unit, for clustering the replies in the reply set to obtain semantic clusters, and performing semantic equivalence detection on the target reply corresponding to the any instruction and the replies in the reply set to determine the target semantic cluster to which the target reply belongs; determining a second score based on the number of replies generated in the target semantic cluster and the number of all replies generated; and a screening unit, for screening out instruction compliance data corresponding to any instruction based on the semantic consistency detection score and the second score.
[0051] As an optional implementation method of this embodiment, for any instruction, extract the reply corresponding to the any instruction; detect the semantic consistency of the reply, and obtain the semantic consistency detection score in the embedding space, including: for any instruction r, extract the reply [r′_1, r′_2, ..., r′_K] corresponding to the any instruction to obtain a reply set, where K represents the quantity, and [r′_1, r′_2, ..., r′_K] is generated based on a large language model; determine the final sentence embedding for the reply, where the final sentence embedding E = [e_1, e_2, ..., e_K] is determined for the internal state of the K generated replies and the last token of the last layer of the large language model; and after processing the final sentence embedding into a multivariate Gaussian distribution, determine the differential entropy, where the differential entropy is used as a semantic consistency detection score to measure the semantic consistency in the continuous embedding space.
[0052] As an optional implementation of this embodiment, after processing the final sentence embedding into a multivariate Gaussian distribution, determining the differential entropy includes: for the sentence embedding E, processing it into a multivariate Gaussian distribution E~N(μ,Σ); based on Determine the differential entropy, where det(Σ) represents the determinant of the covariance matrix Σ, d is the dimension of the sentence embedding; Σ captures the relationship between K different sentence embeddings, expressed as: The differential entropy is simplified to Among them, λ i represents the i-th eigenvalue of the covariance matrix Σ; G is a constant; and q is any of the aforementioned instructions.
[0053] As an optional implementation of this embodiment, clustering the replies in the reply set to obtain semantic clusters includes: clustering the replies using a pre-trained natural language inference model to obtain different semantic clusters.
[0054] As an optional implementation method of this embodiment, after performing semantic equivalence detection on the target reply corresponding to any instruction and the replies in the reply set, determining the target semantic cluster to which the target reply belongs includes: reasoning on the target reply and the replies in each semantic cluster based on the natural language inference model to determine whether the target reply and the replies in the semantic cluster are semantically equivalent; if they are equivalent, determining the semantic cluster to which the target reply belongs based on majority voting.
[0055] As an optional implementation method of this embodiment, determining the second score based on the number of replies generated in the target semantic cluster and the number of all replies generated includes: taking the ratio of the number of replies generated in the target semantic cluster to the number of all replies generated as the second score.
[0056] As an optional implementation method of this embodiment, filtering out instruction compliance data corresponding to any instruction based on the semantic consistency detection score and the second score includes: determining a final scoring result based on the ratio of the first score to the second score, wherein the scoring result is used as a basis for filtering the instruction compliance data of any instruction.
[0057] As an optional implementation method of this embodiment, the final scoring results calculated based on the reply sets corresponding to any instruction extracted at different times are sorted, and the top N reply sets with the highest ranking are used as instruction compliance data.
[0058] According to an embodiment of the present invention, the present invention also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can implement the method described in any of the above embodiments when executing.
[0059] According to an embodiment of the present invention, the present invention further provides a readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to implement the method described in any of the above embodiments when executed.
[0060] According to an embodiment of the present invention, the present invention further provides a computer program product, which can implement the method described in any of the above embodiments when executed by a processor.
[0061] Figure 2A schematic block diagram of an example electronic device 300 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices.
[0062] like Figure 2 As shown, the electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. Various programs and data required for the operation of the electronic device 300 can also be stored in the RAM 303. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0063] Multiple components in the electronic device 300 are connected to the I / O interface 305, including an input unit 306, such as a keyboard, a mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; a storage unit 308, such as a magnetic disk, an optical disk, etc.; and a communication unit 309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the electronic device 300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0064] The computing unit 301 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 301 performs the various methods and processes described above, such as the object matching method. For example, in some embodiments, the object matching method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 300 via the ROM 302 and / or the communication unit 309. When the computer program is loaded into the RAM 303 and executed by the computing unit 301, one or more steps of the method described above can be performed.
[0065] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0066] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0067] In the context of the present invention, machine-readable medium can be a tangible medium that can contain or store a program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
Claims
1. A method for filtering instruction-following data for large language model hallucination phenomenon, characterized in that: include: For any instruction, extract the reply corresponding to the instruction to obtain a reply set, wherein the reply corresponding to the instruction is generated by the large language model based on the instruction; Detecting the semantic consistency of the replies in the reply set to obtain a semantic consistency detection score in the embedding space; Clustering the replies in the reply set to obtain semantic clusters, and performing semantic equivalence testing between the target reply corresponding to any one of the instructions and the replies in the reply set to determine the target semantic cluster to which the target reply belongs; and determining a second score based on the number of replies generated in the target semantic cluster and the number of all replies generated; Filtering instruction compliance data corresponding to any one of the instructions based on the semantic consistency detection score and the second score; Filtering instruction compliance data corresponding to any instruction based on the semantic consistency detection score and the second score includes: determining a final scoring result based on a ratio of the first score to the second score, wherein the scoring result is used as a basis for filtering the instruction compliance data of any instruction; The final scoring results calculated based on the response sets corresponding to any instruction extracted at different times are sorted, and the top N response sets with the highest ranking are used as instruction compliance data.
2. The instruction-following data screening method for large language model hallucination phenomenon according to claim 1, characterized in that: For any instruction, extract the corresponding reply of the instruction; Detecting the semantic consistency of the reply and obtaining a semantic consistency detection score in the embedding space includes: For any command r, extract the corresponding reply of any command , get the reply set, where K represents the number, Generated based on large language models; Determine the final sentence embedding for the reply, wherein the final sentence embedding is determined for the K generated replies and the internal state of the last token of the last layer of the large language model ; After processing the final sentence embedding into a multivariate Gaussian distribution, a differential entropy is determined, wherein the differential entropy is used as a semantic consistency detection score to measure semantic consistency in a continuous embedding space.
3. The method for filtering instruction-following data for large language model hallucination according to claim 2, characterized in that: After processing the final sentence embedding into a multivariate Gaussian distribution, determining the differential entropy includes: For sentence embedding , which is processed into a multivariate Gaussian distribution ; based on Determine the differential entropy, where Represents the covariance matrix The determinant of is the dimension of sentence embedding; Captured The relationship between different sentence embeddings is expressed as: ; The differential entropy is simplified to ,in, Represents the covariance matrix No. eigenvalues; is a constant; For any of the above instructions.
4. The method for filtering instruction-following data for large language model hallucination according to claim 1, characterized in that: Clustering the replies in the reply set to obtain semantic clusters includes: The replies are clustered using a pre-trained natural language inference model to obtain different semantic clusters.
5. The method for filtering instruction-following data for large language model hallucination according to claim 4, characterized in that: After performing semantic equivalence detection on the target reply corresponding to any one of the instructions and the replies in the reply set, determining the target semantic cluster to which the target reply belongs includes: Reasoning the target reply and replies in each semantic cluster based on the natural language inference model to determine whether the target reply and replies in the semantic cluster are semantically equivalent; If they are equal, the semantic cluster to which the target reply belongs is determined based on majority voting.
6. The method for filtering instruction-following data for large language model hallucination according to claim 1, characterized in that: Determining the second score based on the number of replies generated in the target semantic cluster and the number of all generated replies includes taking the ratio of the number of replies generated in the target semantic cluster to the number of all generated replies as the second score.
7. A device for filtering instruction-following data for large language model hallucination phenomenon, characterized in that: include: A semantic consistency detection unit is configured to extract, for any instruction, a reply corresponding to the instruction to obtain a reply set, wherein the reply corresponding to the instruction is generated by a large language model based on the instruction; detect the semantic consistency of the replies in the reply set to obtain a semantic consistency detection score in the embedding space; a semantic equivalence detection unit, configured to cluster the replies in the reply set to obtain semantic clusters, and determine a target semantic cluster to which the target reply belongs after performing semantic equivalence detection on the target reply corresponding to any one of the instructions and the replies in the reply set; and determine a second score based on the number of replies generated in the target semantic cluster and the number of all replies generated; A screening unit for screening out instruction compliance data corresponding to any instruction based on the semantic consistency detection score and the second score; Filtering instruction compliance data corresponding to any instruction based on the semantic consistency detection score and the second score includes: determining a final scoring result based on a ratio of the first score to the second score, wherein the scoring result is used as a basis for filtering the instruction compliance data of any instruction; The final scoring results calculated based on the response sets corresponding to any instruction extracted at different times are sorted, and the top N response sets with the highest ranking are used as instruction compliance data.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Chat content sorting method and device, robot equipment and readable storage medium
CN115964543A
Text generation method and device, computer program product, electronic equipment and medium
CN118396123A