Large model illusion detection method and device, storage medium and electronic equipment
By inserting perturbed text into the generative large model and detecting its output, the problem of hallucination detection of large models is solved, and the stability and robustness of the model is quantified, which improves the reliability of the application.
Patent Information
- Application Number
- CN202510089598.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to effectively detect and identify hallucinatory problems in generative large model output, resulting in limited application of models in scenarios with high authenticity requirements.
By inserting perturbing characters into the text to be detected, multiple perturbing texts are generated, and these perturbing texts are input in parallel to the target big model, the characterization vectors output from each layer are extracted, and the consistency value is centralized to calculate the consistency value, and finally the illusion detection result of the model is judged based on the stability score.
Accurate illusion detection of large model output is realized, and by quantifying the stability and robustness of the model, application reliability in scenarios with high authenticity requirements is improved.
Smart Images

Figure CN119990124A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a large model hallucination detection method, device, storage medium and electronic equipment. Background Art
[0002] Generative large models (also called large models) demonstrate contextual semantic understanding and text generation capabilities close to those of humans, and have performed very well in various natural language understanding and natural language generation tasks. However, the hallucination problem is an obstacle that hinders the large-scale application of large models. Large model hallucination refers to the content generated by large models that is unreal or inconsistent with the input, which reduces the credibility of the large model output and hinders the application of large models in various scenarios with high authenticity requirements.
[0003] Currently, there is an urgent need to provide a solution for hallucination detection on large model outputs when applying large models, so as to determine the authenticity of the large model outputs. Summary of the invention
[0004] The embodiments of this specification provide a large model hallucination detection method, which can efficiently and comprehensively identify the sensitivity of the large model to the disturbance input by associating the hallucination with the input disturbance stability, and quantify the stability of the large model, thereby realizing the accurate detection of the large model hallucination. The method includes:
[0005] Acquire a text to be detected, and insert disturbed characters into the text to be detected to generate multiple disturbed texts;
[0006] Inputting the plurality of disturbance texts into the target large model in parallel, obtaining a representation vector outputted by each disturbance text at each layer of the target large model, and forming a vector set of the corresponding layer from the representation vectors outputted by each layer;
[0007] Centralize each of the vector sets to obtain a consistency value corresponding to each of the vector sets, wherein the consistency value is used to measure the correlation between different representation vectors in each of the vector sets;
[0008] A stability score of the target large model under input disturbance is calculated based on each of the consistency values, and a hallucination detection result of the target large model is determined based on the stability score.
[0009] Furthermore, in some implementations, inserting disturbed characters into the text to be detected to generate a plurality of disturbed texts includes:
[0010] Dividing the text to be detected into multiple words;
[0011] When generating each of the plurality of disturbed texts, at least two target words are determined according to a preset probability, and a disturbed character is inserted between the two target words to generate different disturbed texts.
[0012] Furthermore, in some implementations, determining at least two target words according to a preset probability and inserting a disturbance character between the two target words includes:
[0013] For any two adjacent words in the text to be detected, generate a corresponding random number;
[0014] Comparing the random number with the preset probability, and determining at least two target words according to the comparison result;
[0015] At least one disturbance character is randomly selected from a preset semantic-free character library, and the disturbance character is inserted between the two target words.
[0016] Furthermore, in some implementations, comparing the random number with the preset probability and determining at least two target words according to the comparison result includes:
[0017] If the random number is less than or equal to the preset probability, the two words corresponding to the random number are determined to be target words.
[0018] Furthermore, in some implementations, the step of inputting the plurality of disturbance texts into the target large model in parallel to obtain a representation vector outputted by each disturbance text at each layer of the target large model includes:
[0019] Input the multiple disturbance texts into the target large model in parallel, and after the model inference is completed, obtain the output vector of all words in each disturbance text at each layer of the target large model;
[0020] The output vectors of all words in each of the perturbed texts at each layer of the target large model are averaged to obtain the representation vectors output by each of the perturbed texts at each layer of the target large model.
[0021] Furthermore, in some implementations, the representation vectors output by each layer form a vector set of the corresponding layer, including:
[0022] The representation vectors output by each layer are concatenated to obtain the vector set of the corresponding layer.
[0023] Furthermore, in some implementations, centralizing each of the vector sets to obtain a consistency value corresponding to each of the vector sets includes:
[0024] Transpose the matrix formed by each of the vector sets to obtain a corresponding transposed matrix;
[0025] Calculate a target symmetric matrix corresponding to each of the vector sets using a preset central matrix, a matrix represented by each of the vector sets, and a transposed matrix corresponding to each of the vector sets;
[0026] The consistency value corresponding to each of the vector sets is determined according to the eigenvalues in each of the target symmetric matrices.
[0027] Further, in some implementations, determining the consistency value corresponding to each of the vector sets according to the eigenvalues in each of the target symmetric matrices includes:
[0028] Performing eigenvalue decomposition on each of the target symmetric matrices to obtain all eigenvalues of each of the target symmetric matrices;
[0029] The consistency value corresponding to each of the vector sets is calculated based on the target eigenvalue among all the eigenvalues.
[0030] Furthermore, in some embodiments, the method further comprises:
[0031] Determine an identity matrix and an all-1 matrix having the same dimension as the matrix represented by each of the vector sets;
[0032] The central matrix is constructed according to the dimensions of the matrices represented by the vector sets, the identity matrix and the all-1 matrix.
[0033] Further, in some embodiments, the calculating, according to each of the consistency values, a stability score of the target large model under input disturbance comprises:
[0034] The consistency values are weighted and summed to obtain the stability score of the target large model under input disturbance.
[0035] Further, in some embodiments, determining the hallucination detection result of the target macro model according to the stability score includes:
[0036] The stability score is compared with a preset stability threshold, and if the stability score is less than the preset stability threshold, it is determined that the output of the target large model under the input disturbance has hallucinations.
[0037] The embodiment of this specification also provides a large model hallucination detection device, the device comprising:
[0038] A disturbed text generation module, used for acquiring a text to be detected, and inserting disturbed characters into the text to be detected to generate a plurality of disturbed texts;
[0039] A vector set determination module, used for inputting the plurality of disturbance texts into the target large model in parallel, obtaining a representation vector outputted by each disturbance text at each layer of the target large model, and forming a vector set of the corresponding layer from the representation vectors outputted by each layer;
[0040] A consistency evaluation module, used for centralizing each of the vector sets to obtain a consistency value corresponding to each of the vector sets, wherein the consistency value is used to measure the correlation between different representation vectors in each of the vector sets;
[0041] The model hallucination detection module is used to calculate the stability score of the target large model under input disturbance according to each of the consistency values, and to determine the hallucination detection result of the target large model according to the stability score.
[0042] The embodiments of the present specification also provide a storage medium, wherein the storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the steps of the above method.
[0043] An embodiment of the present specification also provides an electronic device, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the above method.
[0044] The embodiments of the present specification also provide a computer program product, wherein the computer program product stores at least one instruction, and the at least one instruction is suitable for being loaded by a processor and executing the above method steps.
[0045] In an embodiment of the present specification, first, a text to be detected is obtained, and a disturbance character is inserted into the text to be detected to generate multiple disturbance texts. Then, multiple disturbance texts are input into the target large model in parallel to obtain the representation vector of each disturbance text output at each layer of the target large model, and each representation vector output at each layer forms a vector set of the corresponding layer. Further, each vector set is centrally processed to obtain a consistency value corresponding to each vector set, wherein the consistency value is used to measure the correlation between different representation vectors in each vector set. Finally, the stability score of the target large model under the input disturbance is calculated based on each consistency value, and the hallucination detection result of the target large model is obtained based on the stability score. On the one hand, by inserting perturbation characters into the text to be detected to generate multiple perturbation texts, the target large model is exposed to diversified input scenarios, and the response behavior of the large model to subtle input changes is tested, which can effectively reveal the robustness characteristics of the internal representation of the large model; and by inputting multiple perturbation texts into the target large model in parallel, extracting the representation vectors of the output of each layer and forming a corresponding vector set, the semantic representation of the model at different levels can be fully covered, thereby ensuring the globality and hierarchy of the detection; on the other hand, by centralizing the vector set and calculating the consistency value, the correlation between different representation vectors in the vector set is measured, which can effectively avoid the limitations of single-layer or single input results and improve the reliability and judgment accuracy of the detection results; on the other hand, by calculating the stability score based on the consistency value and judging the hallucination detection results, a progressive analysis from representation stability to global robustness is achieved, which helps to more accurately distinguish the normal semantic reasoning ability and hallucination generation mode of the large model. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 A schematic diagram of a system architecture for applying the large model hallucination detection method provided in the embodiments of this specification.
[0047] Figure 2 A schematic flow chart of a large model hallucination detection method provided in an embodiment of this specification.
[0048] Figure 3 A schematic diagram of a process for generating a disturbance text provided in an embodiment of this specification.
[0049] Figure 4 A schematic diagram of a flow chart for implementing input disturbance provided in an embodiment of this specification.
[0050] Figure 5 A schematic diagram of a process for determining a representation vector of a disturbed text in a target large model provided in an embodiment of this specification.
[0051] Figure 6 A schematic diagram of a process of evaluating the consistency of a vector set provided in an embodiment of this specification.
[0052] Figure 7 A schematic diagram of the principle of a large model hallucination detection method provided in an embodiment of this specification.
[0053] Figure 8 A schematic diagram of the structure of a large model hallucination detection device provided in an embodiment of this specification.
[0054] Fig. 9 A schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION
[0055] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.
[0056] Figure 1 A schematic diagram of the architecture of the large model hallucination detection method provided by the embodiments of this specification is shown.
[0057] like Figure 1 As shown, the system architecture 100 may include one or more of terminal devices such as a smart phone 101, a portable computer 102, a desktop computer 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal device and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc. The terminal device may be various electronic devices with data processing functions and capable of running a target large model, and the electronic device may have a display screen, which is used to display the text to be detected, the disturbed text, and the hallucination detection results of the target large model input into the target large model, etc.
[0058] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. According to the implementation requirements, there may be any number of terminal devices, networks and servers. For example, the server 105 may be a server cluster composed of multiple servers.
[0059] See also Figure 2 , which is a flow chart of a large model hallucination detection method provided in the embodiment of this specification. In the embodiment of this specification, the large model hallucination detection method is applied to a large model hallucination detection device or an electronic device equipped with a large model hallucination detection device. Figure 2The process shown is described in detail, and the large model hallucination detection method may specifically include the following steps:
[0060] S202, obtaining a text to be detected, and inserting a disturbed character into the text to be detected to generate a plurality of disturbed texts;
[0061] In one or more embodiments of this specification, the text to be detected refers to the input text that needs to be hallucinated, which can be user input in natural language generation tasks, source text for summary generation, queries for question-answering tasks, etc. Disturbed characters refer to characters that do not contain semantic information, such as special symbols such as "#, ^, $", which do not affect the semantics of the text to be detected, and can ensure that the generated text is semantically consistent with the original text, but there are slight changes in the structure. The embodiments of this specification do not limit this. Accordingly, the disturbed text refers to a set of text copies with slight differences generated by slightly disturbing the original input text, which is used to test the representation stability of the target large model when responding to input changes. For example, multiple disturbed texts may refer to multiple texts generated by inserting disturbed characters at different positions in the text to be detected.
[0062] See also Figure 3 The process of generating multiple disturbance texts in this specification may specifically include step S302 and step S304:
[0063] In step S302, the text to be detected is divided into a plurality of words;
[0064] Optionally, the text to be detected can be divided into multiple words by a word segmentation algorithm. For example, the text to be detected "large model generates text content" is divided into three words: large model, generation, and text content. Among them, the word segmentation algorithm can be a rule-based word segmentation, such as a forward maximum matching method, a reverse maximum matching method, etc., or a statistical word segmentation, such as an N-gram model, etc., or a machine learning-based word segmentation, such as using a hidden Markov model, a conditional random field, etc. to estimate the word segmentation results. This specification embodiment does not limit this, and a suitable word segmentation algorithm can be selected according to the specific application scenario.
[0065] After word segmentation is completed, all gaps between adjacent words are potential locations for inserting perturbation characters, such as between the large model and the generation, and between the generation and the text content.
[0066] In step S304, when generating each of the plurality of disturbed texts, at least two target words are determined according to a preset probability, and a disturbed character is inserted between the two target words to generate different disturbed texts.
[0067] In one or more embodiments of this specification, see Figure 4, step S304 may further include steps S402 to S406:
[0068] S402, for any two adjacent words in the text to be detected, generating a corresponding random number;
[0069] For the insertion positions corresponding to any two adjacent words in the text to be detected, a random number of 0≤r≤1 is generated to decide whether to insert a meaningless character at the position.
[0070] For example, for the large model and generating two adjacent words, the generated random number r=0.2, and for generating two adjacent words with text content, the generated random number r=0.4.
[0071] S404, comparing the random number with the preset probability, and determining at least two target words according to the comparison result;
[0072] According to the preset probability, it can be determined whether to insert a disturbance character for all insertion positions between adjacent words one by one. If the preset probability is larger, it means that there are more positions for inserting characters and the degree of disturbance is higher. If the preset probability is smaller, it means that there are fewer positions for inserting characters and the degree of disturbance is lower.
[0073] Optionally, if the random number is less than or equal to a preset probability, the two words corresponding to the random number are determined to be target words. If the random number is greater than the preset probability, the position is skipped and no disturbance character is inserted. Of course, the comparison rule for determining the insertion position can also be flexibly set, and this specification embodiment does not limit this.
[0074] For example, the preset probability p=0.3 means that a random number is independently generated at each insertion position, and a perturbation character is inserted with a probability of 30%, and the position is skipped with a probability of 70%. Based on this, if the preset probability p=0.3 is greater than the random number r=0.2 generated for the large model and the two adjacent words generated, the large model and the two words generated can be determined as the target words, which means that the perturbation character can be inserted between the large model and the two adjacent words generated later.
[0075] By randomly determining the insertion position, the sentence generated by each perturbation is unique, the semantics remain relatively stable, but there are random changes in the structure, providing rich input samples for subsequent model stability evaluation.
[0076] S406: randomly selecting at least one disturbance character from a preset non-semantic character library, and inserting the disturbance character between the two target words.
[0077] Exemplarily, the preset non-semantic character library is ["#","^","$","&","@"]. For the selected insertion position, a disturbance character is randomly selected from the non-semantic character library and inserted into the position, ensuring that the inserted character will not affect the overall semantics of the text, but will cause slight disturbance to the text structure.
[0078] For example, for the target words big model and generation, randomly select the perturbation character "#" and insert it between the big model and generation, and the corresponding perturbation text is: "big model#generated text content".
[0079] It can be understood that, by repeating the operation of randomly selecting disturbance characters for different insertion positions, a plurality of different disturbance texts can be generated.
[0080] The embodiments of this specification make full use of the random perturbation characteristics of text, providing diversity and representativeness for subsequent model input testing, thereby effectively evaluating the response and robustness of the target model to input perturbations. The multiple perturbation texts generated can cover various variations of the original text in the input dimension, and test the robustness of the target model to input changes. Character insertion ensures that the semantics remain consistent, but there are sufficient differences in the structure to avoid excessive perturbations to the model, resulting in uncontrollable output changes. By setting the insertion probability and the character library, the balance between the randomness and controllability of the insertion behavior can be achieved, and a representative set of perturbation copies can be generated, which lays a data foundation for subsequent characterization vector extraction and stability analysis, and is an important prerequisite for detecting the sensitivity of large models to input perturbations.
[0081] S204, inputting the plurality of disturbance texts into the target large model in parallel, obtaining a representation vector output by each disturbance text at each layer of the target large model, and forming a vector set of the corresponding layer from the representation vectors output by each layer;
[0082] Among them, the target large model refers to a model that has been trained to process natural language input and generate relevant representations or outputs, such as a language generation model, an encoding-decoding model, a natural language understanding model, and a multimodal large model, etc. The specific type and specific structure of the target large model are not limited in this specification embodiment. Each perturbation text generates a global representation vector at each layer through the reasoning process of the target large model. The global representation vector is obtained by averaging the output vectors of all words output at each layer. The global representation vectors of multiple perturbation texts are combined to form a vector set of the corresponding layer.
[0083] It should be noted that for inputs with more serious hallucinations, the target large model will be very sensitive to such input differences during the calculation process. Some subtle character differences may cause large deviations in the calculation process of the target large model, resulting in large differences in the output vectors of each layer of the target large model.
[0084] Multiple perturbed texts are used as parallel inputs to the target large model. During the inference process, the target large model generates output representations layer by layer, including the representation of each word in the input text at each layer. It should be noted that each layer of the target large model includes an input layer, a hidden layer, and an output layer.
[0085] In one or more embodiments of this specification, see Figure 5 , step S204 may specifically include the following steps S502 and S504:
[0086] S502, inputting the plurality of disturbance texts into the target large model in parallel, and obtaining output vectors of all words in each disturbance text at each layer of the target large model after the model inference is completed;
[0087] The target large model processes the input perturbation text layer by layer through a multi-layer neural network. For example, the target large model can be a network model based on the Transformer architecture, which generates an output vector corresponding to each token (word or character) in each perturbation text.
[0088] By obtaining the vectors of all words output at each layer, the response of the target large model to input disturbances can be fully described. Moreover, the changes in the representation vectors of different words can reflect the local performance of the model's sensitivity to disturbances, which is helpful for subsequent stability analysis and consistency calculation.
[0089] S504, averaging the output vectors of all words in each of the perturbed texts at each layer of the target large model to obtain a representation vector outputted by each of the perturbed texts at each layer of the target large model.
[0090] For example, for K perturbation texts, there will be K different representation vectors in a certain layer of the target large model, and each representation vector is obtained by averaging the output vectors of all words in each perturbation text in the corresponding layer.
[0091] After obtaining the representation vectors of each perturbed text output at each layer of the target large model, the representation vectors output at each layer can be concatenated to obtain a vector set of the corresponding layer.
[0092] For example, the vector set of each perturbation text at a certain layer of the target model is recorded as:
[0093]
[0094] In formula (1), the matrix Z is a d×K matrix, which means that the matrix Z consists of K representation vectors, such as Z 1 is the representation vector of the first perturbation text output at a certain layer of the target large model, Z 2is the representation vector of the second perturbation text output at the same layer of the target model, each vector Z i The dimension of (i=1,2,…K) is d, represents a d×K real matrix space. Correspondingly, all elements in the matrix Z are real numbers.
[0095] In the embodiments of this specification, not only the representation of a single text is obtained, but also the response of the model is observed from multiple perspectives through the diversity of the perturbed text. Moreover, not only the final output is captured, but also the multi-level representations within the model are covered, and its sensitivity to input changes is deeply analyzed. By generating a vector set, the representations of different perturbed texts at each layer of the model can be effectively organized and compared to form a core data structure for subsequent analysis. As the basis for the subsequent calculation of the consistency value, the vector set provides comprehensive data support for evaluating the robustness of the model under input perturbations.
[0096] S206, performing centralization processing on each of the vector sets to obtain a consistency value corresponding to each of the vector sets, wherein the consistency value is used to measure the correlation between different representation vectors in each of the vector sets;
[0097] Among them, centralization processing refers to removing the mean of the vector set so that the distribution of the vector set is concentrated at zero and the offset effect is eliminated. The consistency value is used to measure the correlation between different vectors in the vector set. Optionally, the consistency value can reflect the similarity between different vectors in the vector set. If the consistency value is high, it means that the model is stable under perturbation input and the changes between the representation vectors are small. On the contrary, it means that the model is sensitive to input perturbations and the changes between the representation vectors are large. The consistency value is a scalar in the interval [0,1]. The higher the value, the stronger the representation consistency of the vector set and the higher the representation stability of the model.
[0098] In one embodiment of this specification, please refer to Figure 6 , the process of centralizing each vector set may include the following steps:
[0099] S602, performing transposition processing on the matrix formed by each of the vector sets to obtain a corresponding transposed matrix;
[0100] For example, the matrix Z shown in formula (1) is transposed to obtain the corresponding transposed matrix Z T .
[0101] S604, calculating a target symmetric matrix corresponding to each of the vector sets using a preset central matrix, a matrix represented by each of the vector sets, and a transposed matrix corresponding to each of the vector sets;
[0102] Optionally, first determine the identity matrix and the all-one matrix with the same dimension as the matrix represented by each vector set, and then construct the central matrix according to the dimension of the matrix represented by each vector set, the identity matrix and the all-one matrix, that is:
[0103]
[0104] Among them, J d is the center matrix, I d is the unit matrix, d is the dimension of the matrix represented by each vector set, 1 K is a matrix of all 1s, is the transpose of the all-one matrix.
[0105] Furthermore, the target symmetric matrix corresponding to each vector set can be calculated using the central matrix shown in formula (2), the matrix representing each vector set, and the transposed matrix corresponding to each vector set.
[0106] For example, taking the vector set Z of a certain layer as an example, we have:
[0107] M=Z T ·J d ·Z (3)
[0108] Among them, M is the target symmetric matrix corresponding to the vector set, Z T is the transposed matrix corresponding to the vector set, J d is the center matrix.
[0109] S606: Determine the consistency value corresponding to each of the vector sets according to the eigenvalues in each of the target symmetric matrices.
[0110] Among them, the eigenvalue can measure the strength of the correlation between different dimensions in the target symmetric matrix. When the input disturbance is large, the representation vector of the target large model may change drastically in different dimensions, resulting in a more dispersed distribution of the eigenvalue. Therefore, the concentration or distribution characteristics of the eigenvalue can reflect the sensitivity of the target large model to the disturbance input.
[0111] In one or more embodiments of the present specification, each target symmetric matrix can be subjected to eigenvalue decomposition to obtain all eigenvalues of each target symmetric matrix. Then, the consistency value corresponding to each vector set is calculated based on the target eigenvalue among all eigenvalues. Among them, the target eigenvalue can be the maximum eigenvalue or the average eigenvalue. Of course, the target eigenvalue can also be determined by other methods such as normalized distribution, which is not limited in the embodiments of the present specification. Accordingly, the maximum eigenvalue in the target symmetric matrix can be used as the consistency value corresponding to the vector set, or the average eigenvalue of the target symmetric matrix can be calculated and used as the consistency value corresponding to the vector set.
[0112] In the embodiments of this specification, the calculated consistency value can effectively quantify the similarity between different representation vectors, and provide hierarchical data for the subsequent calculation of the overall stability score of the model. Importantly, the consistency value can reveal the model's sensitivity to input perturbations and its ability to maintain semantics.
[0113] S208, calculating a stability score of the target large model under input disturbance according to each of the consistency values, and determining a hallucination detection result of the target large model according to the stability score.
[0114] In one or more embodiments of the present specification, the stability score is a comprehensive quantitative indicator of the consistency values of all layers of the target large model, which is used to evaluate the overall stability of the model under input disturbances.
[0115] Optionally, the consistency values corresponding to each vector set may be weighted and summed to obtain the stability score of the target large model under input disturbance, that is:
[0116]
[0117] Among them, S is the stability score, w l is the weight of the target large model at layer l, n is the consistency value corresponding to the vector set at layer l of the target large model, and L is the total number of network layers of the target large model. The stability score S is a scalar in the range of [0,1]. The higher the value, the more stable the target large model is under perturbation input and the smaller the representation change. On the contrary, it means that the target large model is sensitive to input perturbation and the representation change is large.
[0118] The stability score is compared with a preset stability threshold. If the stability score is less than the preset stability threshold, it is determined that the output of the target large model under the input disturbance has hallucinations.
[0119] The embodiments of this specification do not limit the specific value of the stability threshold T, and a reasonable threshold can be selected according to actual needs, such as T = 0.6. If the stability score S = 0.5, it can be determined that the output of the target large model under input disturbance has hallucinations, and if the stability score S = 0.8, it can be determined that the target large model is stable under input disturbances, and the output of the model can be considered to be credible, and no hallucinations are found.
[0120] First, the stability score combines the consistency values of all layers into one indicator, which can fully reflect the overall stability of the model under input disturbances. Second, by comparing the stability score with the threshold, the model's hallucination behavior can be automatically and accurately determined. Finally, the detection results can help locate the model's weaknesses in specific layers or specific inputs, providing a basis for model structure optimization and training improvements.
[0121] In the embodiments of this specification, by calculating the stability score of the target large model under input disturbances and judging the hallucination detection results in combination with a preset threshold, not only the representation stability of the model is quantified, but also accurate detection of hallucination outputs is achieved, providing effective support for high-trust applications.
[0122] In one or more embodiments of this specification, reference is made to Figure 7 , which is a schematic diagram of the principle of a large model hallucination detection method provided in an embodiment of this specification.
[0123] exist Figure 7 In the example, for the text to be detected "the large model can generate high-quality text content", at least one disturbance character is randomly selected from the interference character library 701 each time, such as selecting the first disturbance character 702 for the first time and the second disturbance character 703 for the second time, and repeated multiple times according to actual needs to select multiple disturbance characters, and the disturbance characters selected each time can be the same or different. It should be noted that when the selected disturbance characters are the same, the same disturbance characters need to be inserted into different positions of the text to be detected to ensure that the inserted characters will not affect the overall semantics of the text, but will bring slight disturbances to the text structure.
[0124] In the embodiment of the present specification, before inserting the disturbance character into the text to be detected, the text to be detected needs to be segmented into multiple words. After the word segmentation is completed, all the gaps between adjacent words are potential insertion positions of the disturbance character. Then, at least two target words are determined from the text to be detected according to a preset probability, and the disturbance character is inserted between the two target words to generate multiple different disturbance texts 704.
[0125] Among them, for the insertion position corresponding to any two adjacent words in the text to be detected, such as the gap between "large model" and "can", a random number of 0≤r≤1 is generated to determine whether to insert a meaningless character at this position. Combined with the preset probability, all insertion positions between adjacent words can be judged one by one whether to insert a disturbance character. For example, if the random number is less than or equal to the preset probability, the two words corresponding to the random number are determined to be target words, which means that a disturbance character can be inserted between the two adjacent target words in the future, thereby determining the disturbance insertion position in the text to be detected.
[0126] Furthermore, the generated multiple disturbance texts 704 are input into the target large model 705 in parallel, and the target large model 705 processes each disturbance text 704 layer by layer through a multi-layer neural network. After the model reasoning is completed, the output result 706 corresponding to each disturbance text 704 is obtained. It should be noted that in the embodiment of this specification, it is necessary to obtain the output vector 707 of each layer of the target large model 705 including the input layer, the middle layer and the output layer of all words in each disturbance text 704 after the model reasoning is completed. By obtaining the vectors of all words output in each layer, the response of the target large model to the input disturbance can be fully described. Moreover, the change of the representation vector of different words can reflect the local performance of the model's sensitivity to disturbance, which is helpful for subsequent stability analysis and consistency calculation.
[0127] Finally, the hallucination detection result 708 can be calculated based on the output vector 707. Specifically, the output vectors 707 of all words in each perturbation text in each layer of the target large model can be averaged to obtain the representation vector of each perturbation text output in each layer of the target large model, and the representation vectors output from each layer can be spliced to obtain the vector set of the corresponding layer. Further, each vector set is centralized to obtain a consistency value corresponding to each vector set, where the consistency value is used to measure the correlation between different representation vectors in each vector set. The process of calculating the consistency value corresponding to each vector set can be referred to steps S602 to S606, which will not be repeated here.
[0128] After calculating the consistency values corresponding to each vector set, the stability score of the target macro model under the input disturbance is calculated according to each consistency value, and the hallucination detection result of the target macro model is obtained according to the stability score 708. If the stability score is compared with a preset stability threshold, if the stability score is less than the preset stability threshold, it is determined that the output of the target macro model under the input disturbance has hallucinations.
[0129] In the embodiments of this specification, the stability of the network calculation process is focused on, and the stability of the semantic representation of the input sentence and the stability of the output sentence generation process are taken into account at the same time. The entire application process not only obtains the representation of a single text, but also observes the response of the model from multiple perspectives through the diversity of the perturbed text. Moreover, it not only captures the final output, but also covers the multi-level representations within the model, and deeply analyzes its sensitivity to input changes. By generating a vector set, the representations of different perturbed texts at each layer of the model can also be effectively organized and compared to form a core data structure for subsequent analysis. As the basis for the subsequent calculation of the consistency value, the vector set provides comprehensive data support for evaluating the robustness of the model under input perturbations.
[0130] See also Figure 8, is a schematic diagram of the structure of a large model hallucination detection device provided in an embodiment of this specification. Figure 8 As shown, the large model hallucination detection device 1 can be implemented as all or part of an electronic device through software, hardware or a combination of both. According to some embodiments, the large model hallucination detection device 1 includes a disturbance text generation module 11, a vector set determination module 12, a consistency evaluation module 13 and a model hallucination detection module 14, specifically including:
[0131] The disturbed text generation module 11 is used to obtain the text to be detected and insert disturbed characters into the text to be detected to generate multiple disturbed texts;
[0132] A vector set determination module 12 is used to input the plurality of disturbance texts into the target large model in parallel, obtain a representation vector output by each disturbance text at each layer of the target large model, and form a vector set of the corresponding layer from each representation vector output by each layer;
[0133] A consistency evaluation module 13, used for centralizing each of the vector sets to obtain a consistency value corresponding to each of the vector sets, wherein the consistency value is used to measure the correlation between different representation vectors in each of the vector sets;
[0134] The model hallucination detection module 14 is used to calculate the stability score of the target large model under input disturbance according to each of the consistency values, and determine the hallucination detection result of the target large model according to the stability score.
[0135] Optionally, when inserting disturbed characters into the text to be detected to generate a plurality of disturbed texts, the disturbed text generating module 11 is specifically used to:
[0136] Dividing the text to be detected into multiple words;
[0137] When generating each of the plurality of disturbed texts, at least two target words are determined according to a preset probability, and a disturbed character is inserted between the two target words to generate different disturbed texts.
[0138] Optionally, when the perturbation text generating module 11 determines at least two target words according to a preset probability and inserts a perturbation character between the two target words, it is specifically configured to:
[0139] For any two adjacent words in the text to be detected, generate a corresponding random number;
[0140] Comparing the random number with the preset probability, and determining at least two target words according to the comparison result;
[0141] At least one disturbance character is randomly selected from a preset semantic-free character library, and the disturbance character is inserted between the two target words.
[0142] Optionally, when the perturbation text generation module 11 compares the random number with the preset probability and determines at least two target words according to the comparison result, it is specifically used to:
[0143] If the random number is less than or equal to the preset probability, the two words corresponding to the random number are determined to be target words.
[0144] Optionally, when the vector set determination module 12 inputs the plurality of disturbance texts into the target large model in parallel to obtain the representation vector outputted by each disturbance text at each layer of the target large model, it is specifically used to:
[0145] Input the multiple disturbance texts into the target large model in parallel, and after the model inference is completed, obtain the output vector of all words in each disturbance text at each layer of the target large model;
[0146] The output vectors of all words in each of the perturbed texts at each layer of the target large model are averaged to obtain the representation vectors output by each of the perturbed texts at each layer of the target large model.
[0147] Optionally, when executing the process of forming a vector set of a corresponding layer from the representation vectors output by each layer, the vector set determination module 12 is specifically configured to:
[0148] The representation vectors output by each layer are concatenated to obtain the vector set of the corresponding layer.
[0149] Optionally, when the consistency evaluation module 13 performs centralization processing on each of the vector sets to obtain a consistency value corresponding to each of the vector sets, it is specifically used to:
[0150] Transpose the matrix formed by each of the vector sets to obtain a corresponding transposed matrix;
[0151] Calculate a target symmetric matrix corresponding to each of the vector sets using a preset central matrix, a matrix represented by each of the vector sets, and a transposed matrix corresponding to each of the vector sets;
[0152] The consistency value corresponding to each of the vector sets is determined according to the eigenvalues in each of the target symmetric matrices.
[0153] Optionally, when the consistency evaluation module 13 determines the consistency value corresponding to each of the vector sets according to the eigenvalues in each of the target symmetric matrices, it is specifically used to:
[0154] Performing eigenvalue decomposition on each of the target symmetric matrices to obtain all eigenvalues of each of the target symmetric matrices;
[0155] The consistency value corresponding to each of the vector sets is calculated based on the target eigenvalue among all the eigenvalues.
[0156] Optionally, the consistency evaluation module 13 is further used for:
[0157] Determine an identity matrix and an all-1 matrix having the same dimension as the matrix represented by each of the vector sets;
[0158] The central matrix is constructed according to the dimensions of the matrices represented by the vector sets, the identity matrix and the all-1 matrix.
[0159] Optionally, when the model hallucination detection module 14 calculates the stability score of the target large model under the input disturbance according to each of the consistency values, it is specifically used to:
[0160] The consistency values are weighted and summed to obtain the stability score of the target large model under input disturbance.
[0161] Optionally, when the model hallucination detection module 14 obtains the hallucination detection result of the target large model according to the stability score, it is specifically used to:
[0162] The stability score is compared with a preset stability threshold, and if the stability score is less than the preset stability threshold, it is determined that the output of the target large model under the input disturbance has hallucinations.
[0163] The above device embodiments correspond to the method embodiments. For specific descriptions, please refer to the description of the method embodiments, which will not be repeated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, please refer to the corresponding method embodiments.
[0164] The present specification also provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor as described above. Figures 2 to 7 The method of the embodiment shown in the figure can be specifically executed by referring to Figures 2 to 7 The specific description of the illustrated embodiment will not be repeated here.
[0165] The present specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor as described above. Figures 2 to 7 The method of the embodiment shown in the figure can be specifically executed by referring to Figures 2 to 7 The specific description of the illustrated embodiment will not be repeated here.
[0166] The embodiments of this specification also provide Fig. 9 The structural diagram of the electronic device shown in FIG. Fig. 9 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above-mentioned voice activity detection method.
[0167] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the executor of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0168] In the 1990s, improvements to a technology could be clearly distinguished as hardware improvements (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0169] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.
[0170] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0171] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0172] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0173] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0174] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0176] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0177] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0178] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0179] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0180] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0181] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0182] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0183] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.
Claims
1. A large model hallucination detection method, the method comprising: Acquire a text to be detected, and insert disturbed characters into the text to be detected to generate multiple disturbed texts; Inputting the plurality of disturbance texts into the target large model in parallel, obtaining a representation vector outputted by each disturbance text at each layer of the target large model, and forming a vector set of the corresponding layer from the representation vectors outputted by each layer; Centralize each of the vector sets to obtain a consistency value corresponding to each of the vector sets, wherein the consistency value is used to measure the correlation between different representation vectors in each of the vector sets; A stability score of the target large model under input disturbance is calculated based on each of the consistency values, and a hallucination detection result of the target large model is determined based on the stability score.
2. The large model hallucination detection method according to claim 1, wherein inserting disturbed characters into the text to be detected to generate a plurality of disturbed texts comprises: Dividing the text to be detected into multiple words; When generating each of the plurality of disturbed texts, at least two target words are determined according to a preset probability, and a disturbed character is inserted between the two target words to generate different disturbed texts.
3. The large model hallucination detection method according to claim 2, wherein determining at least two target words according to a preset probability and inserting a disturbance character between the two target words comprises: For any two adjacent words in the text to be detected, generate a corresponding random number; Comparing the random number with the preset probability, and determining at least two target words according to the comparison result; At least one disturbance character is randomly selected from a preset semantic-free character library, and the disturbance character is inserted between the two target words.
4. The large model hallucination detection method according to claim 3, wherein the step of comparing the random number with the preset probability and determining at least two target words according to the comparison result comprises: If the random number is less than or equal to the preset probability, the two words corresponding to the random number are determined to be target words.
5. The large model hallucination detection method according to claim 1, wherein the inputting the plurality of disturbance texts into the target large model in parallel to obtain the representation vector of each disturbance text output at each layer of the target large model comprises: Input the multiple disturbance texts into the target large model in parallel, and after the model inference is completed, obtain the output vector of all words in each disturbance text at each layer of the target large model; The output vectors of all words in each of the perturbed texts at each layer of the target large model are averaged to obtain the representation vectors output by each of the perturbed texts at each layer of the target large model.
6. The large model hallucination detection method according to claim 1, wherein the vector set of the corresponding layer composed of the representation vectors output by each layer comprises: The representation vectors output by each layer are concatenated to obtain the vector set of the corresponding layer.
7. The large model hallucination detection method according to claim 1, wherein the centralizing each of the vector sets to obtain a consistency value corresponding to each of the vector sets comprises: Transpose the matrix formed by each of the vector sets to obtain a corresponding transposed matrix; Calculate a target symmetric matrix corresponding to each of the vector sets using a preset central matrix, a matrix represented by each of the vector sets, and a transposed matrix corresponding to each of the vector sets; The consistency value corresponding to each of the vector sets is determined according to the eigenvalues in each of the target symmetric matrices.
8. The large model hallucination detection method according to claim 7, wherein the step of determining the consistency value corresponding to each of the vector sets according to the eigenvalues in each of the target symmetric matrices comprises: Performing eigenvalue decomposition on each of the target symmetric matrices to obtain all eigenvalues of each of the target symmetric matrices; The consistency value corresponding to each of the vector sets is calculated based on the target eigenvalue among all the eigenvalues.
9. The large model hallucination detection method according to claim 7, further comprising: Determine an identity matrix and an all-1 matrix having the same dimension as the matrix represented by each of the vector sets; The central matrix is constructed according to the dimensions of the matrices represented by the vector sets, the identity matrix and the all-1 matrix.
10. The large model hallucination detection method according to claim 1, wherein the step of calculating the stability score of the target large model under input disturbance according to each of the consistency values comprises: The consistency values are weighted and summed to obtain the stability score of the target large model under input disturbance.
11. The large model hallucination detection method according to claim 1, wherein the step of obtaining the hallucination detection result of the target large model according to the stability score comprises: The stability score is compared with a preset stability threshold, and if the stability score is less than the preset stability threshold, it is determined that the output of the target large model under the input disturbance has hallucinations.
12. A large model hallucination detection device, the device comprising: A disturbed text generation module, used for acquiring a text to be detected, and inserting disturbed characters into the text to be detected to generate a plurality of disturbed texts; A vector set determination module, used for inputting the plurality of disturbance texts into the target large model in parallel, obtaining a representation vector outputted by each disturbance text at each layer of the target large model, and forming a vector set of the corresponding layer from the representation vectors outputted by each layer; A consistency evaluation module, used for centralizing each of the vector sets to obtain a consistency value corresponding to each of the vector sets, wherein the consistency value is used to measure the correlation between different representation vectors in each of the vector sets; The model hallucination detection module is used to calculate the stability score of the target large model under input disturbance according to each of the consistency values, and to determine the hallucination detection result of the target large model according to the stability score.
13. A storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.
14. An electronic device comprising: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the method as claimed in any one of claims 1 to 11.
15. A computer program product having at least one instruction stored thereon, wherein when the at least one instruction is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.
Citation Information
Cited By
Big language model illusion detection method and system
CN121167438A