Methods and devices for detecting and correcting LLM's hallucination
The method addresses the issue of hallucinations in LLMs by analyzing internal vectors to calculate a hallucination score, enhancing output accuracy and reducing factual errors in generated text.
Patent Information
- Application Number
- PCT/US2025/011944
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-16
- Filing Date
- 2025-01-16
- Publication Date
- 2025-07-24
AI Technical Summary
Large language models (LLMs) often generate human-like outputs that contain subtle or obvious mistakes, known as hallucinations, which existing methods struggle to detect and correct effectively.
A method for detecting hallucinations in LLMs by collecting internal vectors from transformer blocks, calculating a hallucination score based on these vectors, and improving output quality through metrics such as target token ranking and clustering.
Enhances the accuracy of LLM outputs by quantifying the degree of factual correctness and reducing the likelihood of hallucinations, thereby improving the reliability of generated text.
Smart Images

Figure US2025011944_24072025_PF_FP_ABST
Abstract
Description
METHODS AND DEVICES FOR DETECTING AND CORRECTING LLM’S HALLUCINATIONCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 621,550 filed on January 16, 2024. The entire disclosure of the aforementioned application is incorporated herein by reference in its entirety for all purposes.TECHNICAL FIELD
[0002] The present disclosure generally relates to the field of artificial neural networks (ANNs), and more specifically, to large language models (LLMs).BACKGROUND
[0003] Large Language Models (LLMs), such as ChatGPT, are computer algorithms that learn from a vast amount of text data and are capable of generating seemingly intelligent text outputs from human inputs. The core of the algorithms is artificial neural networks (ANNs), which is a way for computers to simulate how neurons in animal brains connect and communicate with each other to process information. Despite the human-like outputs LLMs generate today, they often make subtle or obvious mistakes that an equally intelligent person rarely makes. This phenomenon is called hallucination by the LLM research community. New methods are desired to detect and correct LLM’s hallucination.SUMMARY
[0004] According to a first aspect of the present application, a method for detecting hallucination in a large language neural network is provided. The method may include receiving an input sequence of a plurality of input tokens, where the large language neural network includes a plurality of transformer blocks, each transformer block includes a plurality of transformer layers, a number of the plurality of input tokens is N, and a number of the plurality of transformer layers is L, where N and L are respectively positive integers. Additionally, the method may include: obtaining a plurality of internal vectors according to N and L, storing the plurality of internal vectors, and obtaining a hallucination score of an output token for the large language neural network according to the plurality of internal vectors that are stored.
[0005] According to a second aspect of the present application, an apparatus for detecting hallucination in a large language neural network is provided. The apparatus may include one or more processors and a memory coupled to the one or more processors and configured to storeinstructions executable by the one or more processors. Furthermore, the one or more processors, upon execution of the instructions, are configured to perform acts according to the first aspect.
[0006] According to a second aspect of the present application, a non-transitory computer- readable storage medium is provided. The non-transitory computer-readable storage medium stores computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform acts according to the first aspect.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure.
[0008] FIG. 1 illustrates an example of a transformer block including transformer layers in accordance with some implementations of the present disclosure.
[0009] FIG. 2 is a diagram that shows attention block allows the computation to involve other tokens in accordance with some implementations of the present disclosure.
[0010] FIG. 3 is a diagram illustrating a computing environment coupled with a user interface in accordance with some implementations of the present disclosure.
[0011] FIG. 4 is a flowchart illustrating a method for detecting hallucination in a large language neural network or an LLM according to an example of the present disclosure.
[0012] FIG. 5 is a flowchart illustrating a method for detecting hallucination in a large language neural network or an LLM according to an example of the present disclosure.DETAILED DESCRIPTION
[0013] Reference will now be made in detail to specific implementations, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous nonlimiting specific details are set forth in order to assist in understanding the subject matter presented herein. But it will be apparent to one of ordinary skill in the art that various alternatives may be used without departing from the scope of claims and the subject matter may be practiced without these specific details. For example, it will be apparent to one of ordinary skill in the art that the subject matter presented herein can be implemented on many types of electronic devices with digital video capabilities.
[0014] Terms used in the disclosure are only adopted for the purpose of describing specific embodiments and not intended to limit the disclosure. “A / an,” “said,” and “the” in a singular formin the disclosure and the appended claims are also intended to include a plural form, unless other meanings are clearly denoted throughout the disclosure. It is also to be understood that term “and / or” used in the disclosure refers to and includes one or any or all possible combinations of multiple associated items that are listed.
[0015] Reference throughout this specification to “one embodiment,” “an embodiment,” “an example,” “some embodiments,” “some examples,” or similar language means that a particular feature, structure, or characteristic described is included in at least one embodiment or example. Features, structures, elements, or characteristics described in connection with one or some embodiments are also applicable to other embodiments, unless expressly specified otherwise.
[0016] Throughout the disclosure, the terms “first,” “second,” “third,” etc. are all used as nomenclature only for references to relevant elements, e.g., devices, components, compositions, steps, etc., without implying any spatial or chronological orders, unless expressly specified otherwise. For example, a “first device” and a “second device” may refer to two separately formed devices, or two parts, components, or operational states of a same device, and may be named arbitrarily.
[0017] The terms “module,” “sub-module,” “circuit,” “sub-circuit,” “circuitry,” “sub-circuitry,” “unit,” or “sub-unit” may include memory (shared, dedicated, or group) that stores code or instructions that can be executed by one or more processors. A module may include one or more circuits with or without stored code or instructions. The module or circuit may include one or more components that are directly or indirectly connected. These components may or may not be physically attached to, or located adjacent to, one another.
[0018] As used herein, the term “if’ or “when” may be understood to mean “upon” or “in response to” depending on the context. These terms, if appear in a claim, may not indicate that the relevant limitations or features are conditional or optional. For example, a method may comprise steps of: i) when or if condition X is present, function or action X’ is performed, and ii) when or if condition Y is present, function or action Y’ is performed. The method may be implemented with both the capability of performing function or action X’, and the capability of performing function or action Y’. Thus, the functions X’ and Y’ may both be performed, at different times, on multiple executions of the method.
[0019] A unit or module may be implemented purely by software, purely by hardware, or by a combination of hardware and software. In a pure software implementation, for example, the unitor module may include functionally related software components, that are directly or indirectly linked together, so as to perform a particular function.
[0020] LLMs are advanced natural language processing models that use deep learning techniques, particularly neural networks, to understand and generate human-like text. Examples of LLMs include the GPT (Generative Pre-trained Transformer) series developed by OpenAI. These models are characterized by their large number of parameters and their ability to capture complex language patterns.
[0021] LLMs include transformer blocks which may include stacked self-attention and point-wise, fully connected layers, feedforward network (FFN) layers and different layer normalization (LayerNorm) structures. FIG. 1 illustrates an example of a transformer block including transformer layers in accordance with some implementations of the present disclosure.
[0022] As shown in FIG. 1, xi and x represent 2 floating point vectors. This diagram shows how the output x / +i is computed from the input xi. In some examples, all vectors in the present disclosure may have a fixed dimension D and the fixed dimension D may have a range of 4K-12K. The diagram in FIG. 1 may include multiple neural network blocks including LayerNorm 101, Attention 102, Addition 103, LayerNorm 104, FFN 105, and Addition 106. LayerNorm 101 or 104 performs layer normalization by normalizing each individual sample across all features. Attention 102 may perform attention function by mapping a query and a set of key -value pairs to an output, where the query, keys, values, and output are all vectors. FFN 105 may include FFN layers that perform linear transformation and non-linear activation and transform input data into a more abstract representation. Addition 102 or 106 may include an additional layer that combines multiple inputs by performing element-wise addition.
[0023] As shown in FIG. 1, LayerNorm 101 receives the input xi, normalizes the input xi, and generates an output xo. Attention 102 receives the output xo from LayerNorm 101 and generates an output X], Furthermore, Addition 103 receives both the input xi and the output xo and generates an output X2. LayerNorm 104 receives the input X2, normalizes the input x . and generates an output xs. FFN 105 receives the output xj from LayerNorm 104 and generates an output x . Addition 106 receives both the output X2 from Addition 103 and the output X4 from FFN 105 and generates an output xs as the output x / +i. xo, xy, X2, xj, x#, xj may all considered internal vectors.
[0024] The diagram illustrated in FIG. 1 is equivalent for the following pseudo code:• Xo = layernorm(Xi)• Xi = attention(Xo)• X2= X1 + X1• X3 = layernorm(X2)• X4= FFN(X3)• x5= x2+ x4
[0025] The above computation is repeated for each input token. More discussion about token is provided below.
[0026] FIG. 2 illustrates an example diagram that shows attention block, such as Attention 102, allows the computation to involve other tokens. As shown in FIG. 2, middle-layer MLPs process subject tokens correspond to factual recall, whereas late-layer attention modules read this information to predict a specific word sequence. Each row in FIG. 2 represents computation for a token and each column represents a transformer block and its computation for the given token.
[0027] Specifically, multiple blocks 200 in FIG. 2 represent embedding layers, each of which may include a learnable lookup table that converts a token ID into a floating point vector. Blocks 202 represent attention layers, blocks 203 represent feedforward Multi-Layer Perceptron (MLP) layers that are at a range of middle layers. Block 210 represents an output layer that generate a prediction for the multiple blocks 200. An output layer is a matrix multiply layer that contains learnable parameters. It converts an input vector into a probability distribution over all the tokens. Usually, the token that has the maximum probability is chosen to be the output token from the input vector. Block 201 represents a hidden state.
[0028] In one example, the prompt input in FIG. 2 may be “The Space Needle is in downtown,” for which the expected completion is “Seattle.” Specifically, blocks 200 sequentially from top to bottom represent embedding layers related to “The” “Space” “Needle” “is” “in” “downtown” respectively. Block 210 generates a prediction of “Seattle” for these inputs.
[0029] The Transformer blocks are chained together in a layer by layer fashion into a so-called deep learning model. It is common for a LLM to have 30-100 Transformer blocks that are exactly identical in terms of computation. However, they usually have distinct learnable parameters in each layer (i.e., network weights, or weights). In addition to the Transformer blocks, LLMs also contains 4 other important parts: tokenizer, embedding layer, output layer, and next token prediction training method.
[0030] Tokenizer
[0031] Tokenizer is an algorithm that breaks the input text into tokens, which are a predefined set of strings. The tokens may include English words. Some tokens may also include numbers, subwords, symbols or foreign language characters, etc. Each token is represented by an integer, called token ID. LLMs usually have 30K-50K tokens.
[0032] Embedding Layer
[0033] Embedding layer is a learnable lookup table that converts a token ID into a floating point vector. As described above, blocks 200 in FIG. 2 include embedding layers.
[0034] Output layer
[0035] Output layer is a matrix multiply layer that contains learnable parameters. It converts an input vector into a probability distribution over all the tokens. Usually, the token that has the maximum probability is chosen to be the output token from the input vector. As described above, block 210 in FIG. 2 includes the output layer.
[0036] Next token prediction training method
[0037] Next token prediction training method is an objective function used to train an LLM. For any given input, after tokenization, the LLMs see a sequence of tokens. LLMs are trained to minimize the error in predicting the next token coming after this input sequence. This unidirectional nature means 2 things. First, once a token is generated from an input sequence, it is added to the end of the sequence and used to generate the next token.
[0038] Second, when generating a token, the model can only refer back to the previous tokens before it, even though the subsequent tokens might be already available. For example, given a sequence <Si, S2, Ss>, a token Ti is generated by LLMs from <Si>, where Ti and S2 may or may not be the same, especially when Si, S2 are given by a human. In a similar manner, the sequence <Si, S2> would generate a new token T2. In doing so, T2 is selected from a set of likely output tokens. The strategy in selecting T2 from the likely set affects future outputs after it.
[0039] Collect Internal Vectors
[0040] The present disclosure provides a method for extracting internal state from the LLMs. In the method, internal vectors will be collected and saved for further analysis.
[0041] In some examples, 2 * N * L internal vectors may be collected from the LLM for a given input sequence of N input tokens and a LLM with L transformer layers. In one example, 2 vectors may be collected per layer, per token. In another example, as shown in FIG. 1, the internal vectors of X2 and x may be collected.
[0042] Store collected vectors with metadata
[0043] Given the collected internal vectors according to the method above, the collected vectors with their metadata may be stored for further analysis.
[0044] In some examples, for each collected vector, one or more of the following metadata will be stored including: layer number from which the vector is collected; token position, token id, token string for the token from which the vector is collected; sublayer number, for example, the sublayer number may be 2 or 5, depending on whether the collected vector is X2 or xj; and / or a unique tag that numbers the entire prompt token sequence. The unique tag may be recorded with each collected vector.
[0045] Collect Parameters of Embedding Layer and / or Output Layer
[0046] In some examples, parameters of embedding layer and / or output layer may be collected in the form of vectors. For both of embedding layer and output layer, the parameters are usually stored as a V-by-D matrix, where V is the number of unique tokens in the model’s vocabulary and D is the vector dimension. The V-by-D matrix may be broken or divided into V vectors and stored with associated token metadata including token integer number and token string. Two sets of vectors may be obtained respectively. One set is input embeddings for embedding layer and the other set is output embeddings for output layer.
[0047] After collecting the 2 sets of vectors, a vector index will be built, using methods such as FAISS. This index is later used to query which input or output embedding token is closest from a given vector of interest.
[0048] Computing Metrics from the Collected Vectors
[0049] The present disclosure provides several methods for computing meaningful and more analytic metrics from the collected vectors.
[0050] In some examples, the nearest token for each collected vector may be used. For examples, for each collected vector, the two sets of input embeddings and output embeddings may be queried and the nearest token in the vector space may be obtained.
[0051] In some other examples, target token ranking may be used. For each collected vector V, the output embeddings may be queried and the ranking of a given target token based on relative distance between V and all embedding vectors may be computed. The target token may be either the final predicted token or the “correct” target token based on the definition of “correct.”
[0052] Hallucination Score
[0053] The present disclosure provides a method to compute a “hallucination score” from the above computed metrics. A hallucination score or a hallucination value may be used to quantify the degree to which a generated output of information is factually incorrect, nonsensical, or not grounded in reality. The hallucination score may be calculated in one of following methods.
[0054] In some examples, the hallucination score is calculated as a sum of the rankings of the final predicted token at different layers from layer X to layer Y. The higher the score, the lower the quality of the output and more likely it is a hallucination.
[0055] In some examples, the hallucination score is a first layer prediction score computed by finding the first layer where the final predicted token appears in the top N nearest vectors from the embedding space. The earlier the first layer prediction, the lower the chance of hallucination.
[0056] In some examples, the hallucination score is a predicted token clustering score is computed by collecting the nearest tokens from layer X to layer Y and compute the cluster size for these tokens. The smaller the cluster, the lower the chance of hallucination.
[0057] Based on the hallucination score above, the present disclosure further provides methods to improve the LLM’s generated outputs.
[0058] In some examples, “sampling” is used to try different output tokens and rank their hallucination scores and pick the best output.
[0059] In one example, during token generation, output tokens may be selected from a set of possible tokens which have higher probability than the rest.
[0060] In another example, different output tokens may be selected and the hallucination scores may be ranked among different selections.
[0061] In another example, based on the hallucination scores, the generated outputs with the least chance of hallucination may be selected.
[0062] In some other examples, for the same question provided to a LLM, the input may vary in different ways. For example, the input may be in different order of words and sentences in the inputs, different tokenization for the same inputs, different selection of relevant information from a big pool of candidate sentences (RAG: retrieval augmented generation), layer-dropping: randomly skip computation from LLM’s layers. In one example, hallucination scores among different prompts may be calculated based on the target token ranking as described above.
[0063] In another example, the selected prompt may be used as an input to the LLM to improve its generation quality.
[0064] FIG. 3 shows a computing environment or a computing device 310 coupled with a user interface 360. The computing environment 310 can be part of a data processing server. In some embodiments, the computing device 310 can perform any of various methods or processes (such as encoding / decoding methods or processes) as described hereinbefore in accordance with various examples of the present disclosure. The computing environment 310 may include one or more processors 320, a memory 340, and an I / O interface 350.
[0065] The one or more processors 320 typically controls overall operations of the computing environment 310, such as the operations associated with the display, data acquisition, data communications, and image processing. The one or more processors 320 may include one or more processors to execute instructions to perform all or some of the steps in the above-described methods. Moreover, the one or more processors 320 may include one or more modules that facilitate the interaction between the one or more processors 320 and other components. The processor may be a Central Processing Unit (CPU), a microprocessor, a single chip machine, a GPU, or the like.
[0066] The memory 340 is configured to store various types of data to support the operation of the computing environment 310. Memory 340 may include predetermine software 342. Examples of such data include instructions for any applications or methods operated on the computing environment 310, video datasets, image data, etc. The memory 340 may be implemented by using any type of volatile or non-volatile memory devices, or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic or optical disk.
[0067] The VO interface 350 provides an interface between the one or more processors 320 and peripheral interface modules, such as a keyboard, a click wheel, buttons, and the like. The buttons may include but are not limited to, a home button, a start scan button, and a stop scan button. The VO interface 350 can be coupled with an encoder and decoder.
[0068] In some embodiments, there is also provided a non-transitory computer-readable storage medium including a plurality of programs, such as included in the memory 340, executable by the one or more processors 320 in the computing environment 310, for performing the abovedescribed methods. For example, the non-transitory computer-readable storage medium may be aROM, a RAM, a CD-ROM, a magnetic tape, a floppy disc, an optical data storage device or the like.
[0069] The non-transitory computer-readable storage medium has stored therein a plurality of programs for execution by a computing device having one or more processors, where the plurality of programs when executed by the one or more processors, cause the computing device to perform the above-described method for motion prediction.
[0070] In some embodiments, the computing environment 310 may be implemented with one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), graphical processing units (GPUs), controllers, micro-controllers, microprocessors, or other electronic components, for performing the above methods.
[0071] FIG. 4 is a flowchart illustrating a method for detecting hallucination in a large language neural network, i.e., an LLM, according to an example of the present disclosure.
[0072] In step 401, the one or more processors 320 may receive an input sequence of a plurality of input tokens, where the large language neural network includes a plurality of transformer blocks, each transformer block includes a plurality of transformer layers.
[0073] In step 402, the one or more processors 320 may obtain a plurality of internal vectors according to N and L, where a number of the plurality of input tokens is N, a number of the plurality of transformer layers is L, and N and L are respectively positive integers.
[0074] In some examples, the plurality of internal vectors obtained or collected in step 402 may be obtained or collected by different ways or methods including: obtaining two internal vectors for each input token and each transformer layer; or obtaining a third internal vector and a fifth internal vector for each input token and each transformer layer, where the third internal vector and the fifth internal vector may be respectively vectors X2 or xs as shown in FIG. 1. Specifically, the first normalization layer (LayerNorm 101) receives a first floating point vector xi associated with an input token and outputs a first internal vector xo, the attention block (Attention 102) receives the first internal vector xo and outputs a second internal vector xy, the first addition layer (Addition 103) receives both the first floating point vector xi and the second internal vector xi, and outputs a third internal vector X2, the second normalization layer (LayerNorm 104) receives the third internal vector X2 and outputs a fourth internal vector j, the FFN layer (FFN 105) receives the fourth internal vector xj and outputs a fifth internal vector X4, the second addition layer (Addition 106)receives both the third internal vector x and the fifth internal vector X4, and outputs a sixth internal vector xj, i.e., a second floating pointer vector x / +y.
[0075] In step 403, the one or more processors 320 may store the plurality of internal vectors.
[0076] In some examples, the one or more processors 320 may further store the plurality of internal vectors and corresponding metadata, where the corresponding metadata includes a layer number of a layer at which an internal vector is obtained, a token position, a token ID, a token string for an input token from which an internal vector is obtained, a sublayer number, a unique tag. For example, the entire prompt token sequence with a unique tag and record this unique tag may be recorded together with each collected vector.
[0077] In step 404, the one or more processors 320 may obtain a hallucination score of an output token for the large language neural network according to the plurality of internal vectors that are stored.
[0078] In some examples, as shown in FIG. 5, before step 404, the one or more processors 320 may obtain input embeddings and output embeddings in step 405 and store the input embeddings and the output embeddings in step 406, and at step 404, the one or more processors 320 may obtain the hallucination score of the output token for the large language neural network according to the plurality of internal vectors and further according to the input embeddings and the output embeddings that are stored, as shown in FIG. 5. The order of steps in FIG. 5 is for illustration only, but not limited to the specific order in FIG. 5.
[0079] At step 407, the one or more processors 320 may generate vector indexes for the input embeddings and the output embeddings.
[0080] Furthermore, at step 408, the one or more processors 320 may obtain metrics according to the plurality of internal vectors.
[0081] In some examples, the metrics obtained in step 408 may be obtained in different ways. In one example, the one or more processors 320 may query the input embeddings and the output embeddings according to the vector indexes and obtain a nearest token from a given vector for each internal vector that is stored.
[0082] In another example, the one or more processors 320 may query the output embeddings for each internal vector and calculate rankings of a given target token based on relative distances between each internal vector and all the output embeddings.
[0083] At step 409, the one or more processors 320 may calculate the hallucination score according to the metrics.
[0084] In some examples, the hallucination score obtained in step 409 may be calculated in different ways. In one example, the one or more processors 320 may calculate a sum of rankings of the output token at different layers in the large language neural network and obtain the hallucination score of the output token for the large language neural network according to the sum of rankings. In this example, the higher the score, the lower the quality of the output and more likely it is a hallucination.
[0085] In another example, the one or more processors 320 may calculate a first layer prediction score according to a first layer, where at the first layer, the output token appears in top N nearest vectors in an embedding space including input and output embeddings, and where N is a positive integer. In this example, the earlier the first layer prediction, the lower the chance of hallucination.
[0086] In another example, the one or more processors 320 may calculate a predicted token clustering score by obtaining nearest tokens from a first layer to a second layer and calculating the cluster size for the nearest tokens. In this example, the smaller the cluster, the lower the chance of hallucination.
[0087] In some examples, the one or more processors 320 may further obtain multiple output tokens, calculate a hallucination score for each output token, obtain a ranking for each output token by ranking hallucination scores of the multiple output tokens, and determined a ranking for each output token by ranking hallucination scores of the multiple output tokens.
[0088] In some examples, the one or more processors 320 may further obtain multiple output tokens, calculate a hallucination score for each output token, and determine the output token with a least chance of hallucination based on hallucination scores for the multiple output tokens.
[0089] In some examples, the output token is selected from a set of possible tokens with probability higher than a threshold. For example, output tokens may be selected from a set of possible tokens which have higher probability than the rest.
[0090] In some examples, the input sequence of the plurality of input tokens varies in at least one of following manners: in different order of words and sentences in different input sequences; in different tokenization for a same input sequence; or in different selection of relevant information from a pool of candidate sentences.
[0091] In some examples, the one or more processors 320 may further receive multiple prompts, calculate hallucination score for each prompt, and select one prompt according hallucination scores of the multiple prompts as the input sequence. Using the selected prompt as an input to the LLM to improve its generation quality
[0092] In some examples, there is also provided a computing device including one or more processors (for example, the processor 320); and the non-transitory computer-readable storage medium or the memory 330 having stored therein a plurality of programs executable by the one or more processors, wherein the one or more processors, upon execution of the plurality of programs, are configured to perform the above-described methods.
[0093] In some examples, there is also provided a computer program product including a plurality of programs, for example, in the memory 330, executable by the processor 320 in the computing environment 310, for performing the above-described methods. For example, the computer program product may include the non-transitory computer-readable storage medium.
[0094] In some examples, the computing environment 310 may be implemented with one or more ASICs, DSPs, Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), FPGAs, GPUs, controllers, micro-controllers, microprocessors, or other electronic components, for performing the above methods.
[0095] Further embodiments also include various subsets of the above examples or embodiments combined or otherwise re-arranged in various other examples or embodiments.
[0096] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted over, as one or more instructions or code, a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another, e.g., according to a communication protocol. In this manner, computer-readable media generally may correspond to (1) tangible computer-readable storage media which is non-transitory or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code and / or data structures for implementation of the implementations described in the present application. A computer program product mayinclude a computer-readable medium.
[0097] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limited to the present disclosure. Many modifications, variations, and alternative implementations will be apparent to those of ordinary skill in the art having the benefit of the teachings presented in the foregoing descriptions and the associated drawings.
[0098] Unless specifically stated otherwise, an order of steps of the method according to the present disclosure is only intended to be illustrative, and the steps of the method according to the present disclosure are not limited to the order specifically described above, but may be changed according to practical conditions. In addition, at least one of the steps of the method according to the present disclosure may be adjusted, combined or deleted according to practical requirements.
[0099] The examples were chosen and described in order to explain the principles of the disclosure and to enable others skilled in the art to understand the disclosure for various implementations and to best utilize the underlying principles and various implementations with various modifications as are suited to the particular use contemplated. Therefore, it is to be understood that the scope of the disclosure is not to be limited to the specific examples of the implementations disclosed and that modifications and other implementations are intended to be included within the scope of the present disclosure.
Claims
WHAT IS CLAIMED IS:
1. A method for detecting hallucination in a large language neural network, comprising: receiving an input sequence of a plurality of input tokens, wherein the large language neural network comprises a plurality of transformer blocks, each transformer block comprises a plurality of transformer layers, a number of the plurality of input tokens is N, and a number of the plurality of transformer layers is L, wherein N and L are respectively positive integers; obtaining a plurality of internal vectors according to N and L; storing the plurality of internal vectors; and obtaining a hallucination score of an output token for the large language neural network according to the plurality of internal vectors that are stored.
2. The method of claim 1, wherein obtaining the plurality of internal vectors according to N and L comprises: obtaining two internal vectors for each input token and each transformer layer.
3. The method of claim 2, wherein one transformer block comprises a first normalization layer, an attention block, a first addition layer, a second normalization layer, a feedforward network (FFN) layer, and a second addition layer; wherein the first normalization layer receives a first floating point vector associated with an input token and outputs a first internal vector, wherein the attention block receives the first internal vector and outputs a second internal vector, wherein the first addition layer receives both the first floating point vector and the second internal vector, and outputs a third internal vector, wherein the second normalization layer receives the third internal vector and outputs a fourth internal vector, wherein the FFN layer receives the fourth internal vector and outputs a fifth internal vector, wherein the second addition layer receives both the third internal vector and the fifth internal vector, and outputs a sixth internal vector,wherein obtaining two internal vectors for each input token and each transformer layer further comprises: obtaining the third internal vector and the fifth internal vector for each input token and each transformer layer.
4. The method of claim 3, wherein storing the plurality of internal vectors comprises: storing the plurality of internal vectors and corresponding metadata, wherein the corresponding metadata comprises a layer number of a layer at which an internal vector is obtained, a token position, a token ID, a token string for an input token from which an internal vector is obtained, a sublayer number, a unique tag.
5. The method of claim 1, further comprising: obtaining input embeddings and output embeddings; and storing the input embeddings and the output embeddings; and wherein obtaining the hallucination score of the output token for the large language neural network according to the plurality of internal vectors that are stored further comprises: obtaining the hallucination score of the output token for the large language neural network according to the plurality of internal vectors, the input embeddings, and the output embeddings that are stored.
6. The method of claim 5, further comprising: generating vector indexes for the input embeddings and the output embeddings.
7. The method of claim 6, further comprising: obtaining metrics according to the plurality of internal vectors; and calculating the hallucination score according to the metrics.
8. The method of claim 7, wherein obtaining the metrics according to the plurality of internal vectors comprises: querying the input embeddings and the output embeddings according to the vector indexes; andobtaining a nearest token from a given vector for each internal vector that is stored.
9. The method of claim 7, wherein obtaining the metrics according to the plurality of internal vectors comprises: querying the output embeddings for each internal vector; and calculating rankings of a given target token based on relative distances between each internal vector and all the output embeddings.
10. The method of claim 7, wherein calculating the hallucination score according to the metrics comprises: calculating a sum of rankings of the output token at different layers in the large language neural network; and obtaining the hallucination score of the output token for the large language neural network according to the sum of rankings.
11. The method of claim 7, wherein calculating the hallucination score according to the metrics comprises: calculating a first layer prediction score according to a first layer, wherein at the first layer, the output token appears in top one or more nearest vectors in an embedding space.
12. The method of claim 7, wherein calculating the hallucination score according to the metrics comprises: calculating a predicted token clustering score by obtaining nearest tokens from a first layer to a second layer and calculating the cluster size for the nearest tokens.
13. The method of claim 7, further comprising: obtaining multiple output tokens; calculating a hallucination score for each output token; obtaining a ranking for each output token by ranking hallucination scores of the multiple output tokens; and determining a best output token according the ranking.
14. The method of claim 7, further comprising: obtaining multiple output tokens; calculating a hallucination score for each output token; and determining the output token with a least chance of hallucination based on hallucination scores for the multiple output tokens.
15. The method of claim 1, wherein the input sequence of the plurality of input tokens varies in at least one of following manners: in different order of words and sentences in different input sequences; in different tokenization for a same input sequence; or in different selection of relevant information from a pool of candidate sentences.
16. The method of claim 1, further comprising: receiving multiple prompts; calculating a hallucination score for each prompt; and selecting one prompt according hallucination scores of the multiple prompts as the input sequence.
17. An apparatus for detecting hallucination in a large language neural network, comprising: one or more processors; and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, wherein the one or more processors, upon execution of the instructions, are configured to perform acts comprising: receiving an input sequence of a plurality of input tokens, wherein the large language neural network comprises a plurality of transformer blocks, each transformer block comprises a plurality of transformer layers, a number of the plurality of input tokens is N, and a number of the plurality of transformer layers is L, wherein N and L are respectively positive integers; obtaining a plurality of internal vectors according to N and L; storing the plurality of internal vectors; andobtaining a hallucination score of an output token for the large language neural network according to the plurality of internal vectors that are stored.
18. The apparatus of claim 17, wherein the one or more processors, upon execution of the instructions, are configured to perform acts further comprising: obtaining input embeddings and output embeddings; and storing the input embeddings and the output embeddings; and wherein obtaining the hallucination score of the output token for the large language neural network according to the plurality of internal vectors that are stored further comprises: obtaining the hallucination score of the output token for the large language neural network according to the plurality of internal vectors, the input embeddings, and the output embeddings that are stored.
19. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform acts comprising: receiving an input sequence of a plurality of input tokens, wherein the large language neural network comprises a plurality of transformer blocks, each transformer block comprises a plurality of transformer layers, a number of the plurality of input tokens is N, and a number of the plurality of transformer layers is L, wherein N and L are respectively positive integers; obtaining a plurality of internal vectors according to N and L; storing the plurality of internal vectors; and obtaining a hallucination score of an output token for the large language neural network according to the plurality of internal vectors that are stored.
20. The non-transitory computer-readable storage medium of claim 19, wherein the one or more computer processors are caused to perform acts further comprising: obtaining input embeddings and output embeddings; and storing the input embeddings and the output embeddings; and wherein obtaining the hallucination score of the output token for the large language neural network according to the plurality of internal vectors that are stored further comprises:obtaining the hallucination score of the output token for the large language neural network according to the plurality of internal vectors, the input embeddings, and the output embeddings that are stored.
Citation Information
Patent Citations
System, method, and computer program for transformer neural networks
US20220051080A1
Cited By
Large visual language model illusion mitigation method and device
CN120781883A