Model illusion detection method, system and related device
By obtaining the reply text of the trained and untrained language model, extracting key information and combining multiple output hierarchical feature vectors, the accuracy problem of hallucination detection of large-model text is solved, and more efficient hallucination detection is achieved.
Patent Information
- Application Number
- CN202411822235.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is difficult to detect text illusions of large models quickly and at low cost, affecting the accuracy of specific application scenarios.
By obtaining the problem text input to the reply text of the trained and untrained language model, key information is extracted, and the hallucination detection results are determined based on the feature vectors of multiple output levels, and the information vectors output by the comprehensive model are used for judgment.
It improves the accuracy and efficiency of text hallucination detection, enhances the distinction between hallucinations, reduces the impact of outliers and noise, and improves the stability and generalization ability of the model.
Smart Images

Figure CN119961389A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a model hallucination detection method, system and related devices. Background Art
[0002] With the development of artificial intelligence technology, big models are increasingly appearing in various fields. Big models trained with a large amount of data can significantly help users answer various types of questions. However, the hallucination problem of big models has correspondingly become a difficult problem in the field of big model-related research and has received extensive research attention. The so-called model hallucination refers to the fact that when a big model answers a user's question, the answer will be contrary to the facts. The reason for this is most likely that the big model is interfered by erroneous data when receiving large-scale training data.
[0003] Although academia and industry have been actively seeking effective detection and correction strategies in recent years, current solutions still cannot complete the problem of hallucination detection at a fast speed and low cost. Since the hallucination problem has a very serious impact on the audience of certain specific application scenarios, how to improve the accuracy of text hallucination detection has become an urgent problem to be solved. Summary of the invention
[0004] The main technical problem solved by the present application is to provide a model hallucination detection method, system and related devices, which can improve the accuracy of text hallucination detection.
[0005] In order to solve the above technical problems, a technical solution adopted in the present application is: to provide a model hallucination detection method, including: obtaining a first reply text obtained by inputting a question text into a trained first language model, and a second reply text obtained by inputting the question text into an untrained second language model; obtaining key information based on the question text and the first reply text, and obtaining feature vectors of multiple output levels corresponding to the key information based on reference vectors of multiple output levels obtained during the output of the key information and the second reply text; obtaining an information vector corresponding to the key information based on the feature vectors of multiple output levels, and determining a hallucination detection result of the first reply text based on the information vector.
[0006] In order to solve the above technical problems, another technical solution adopted in the present application is: to provide a model hallucination detection system, including: an acquisition module, used to obtain a first reply text obtained by inputting a question text into a trained first language model, and a second reply text obtained by inputting the question text into an untrained second language model; a generation module, connected to the acquisition module, used to obtain key information based on the question text and the first reply text, and obtain feature vectors of multiple output levels corresponding to the key information based on reference vectors of multiple output levels obtained during the output of the key information and the second reply text; a determination module, used to obtain an information vector corresponding to the key information based on the feature vectors of multiple output levels, and determine the hallucination detection result of the first reply text based on the information vector.
[0007] To solve the above technical problems, another technical solution adopted in the present application is: to provide an electronic device, comprising: a memory and a processor coupled to each other, the memory storing program instructions, and the processor being used to execute the program instructions to implement the method mentioned in the above technical solution.
[0008] In order to solve the above technical problems, another technical solution adopted in the present application is: providing a computer-readable storage medium storing program instructions that can be executed by a processor, wherein the program instructions are used to implement the method mentioned in the above technical solution.
[0009] The beneficial effect of the present application is as follows: different from the prior art, the model hallucination detection method proposed in the present application obtains a first reply text obtained from a first language model trained with a large amount of training data for a question text input, and obtains a second reply text obtained from a second language model that has not been trained for a question text input, and after obtaining key information based on the question text and the first reply text, obtains feature vectors of multiple output levels corresponding to the key information based on reference vectors of multiple output levels obtained in the process of outputting the key information and the second reply text, obtains information vectors corresponding to the key information based on the feature vectors of the multiple output levels, and determines the hallucination detection result of the first reply text output by the first language model based on the information vector, and when obtaining relevant features such as the information vector corresponding to the key information, by combining the outputs of the trained first language model and the untrained second language model, the judgment on whether a hallucination occurs is more discriminatory, thereby improving the accuracy of text hallucination detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:
[0011] Figure 1 It is a flow chart of an implementation method of a model hallucination detection method of the present application;
[0012] Figure 2 It is a flow chart of another implementation method of the model hallucination detection method of the present application;
[0013] Figure 3 yes Figure 2 Step S206 corresponds to a flowchart of an implementation method;
[0014] Figure 4 It is a structural schematic diagram of an implementation method of the model hallucination detection system of the present application;
[0015] Figure 5 It is a structural schematic diagram of an embodiment of the electronic device of the present application;
[0016] Figure 6 It is a structural schematic diagram of an implementation method of a computer-readable storage medium of the present application. DETAILED DESCRIPTION
[0017] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of them, and different implementation methods can be adaptively combined. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0018] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. In addition, "many" in this article means two or more than two.
[0019] The model hallucination detection method proposed in the present application is implemented by a smart terminal, which can be a smart device that at least integrates a corresponding detection function, or an application on a smart device. The smart device can be a mobile phone, a tablet computer, or a personal computer.
[0020] It should be noted that the model hallucination detection method proposed in this application is mainly used to detect the reply text generated by the intelligent analysis model in the medical field. Of course, the method can also be used to detect the reply text generated by the intelligent analysis model in other fields, and this application does not limit this. For the sake of convenience, the following embodiments of this application are introduced in scenarios related to medical care.
[0021] See also Figure 1 , Figure 1 : is a flow chart of an embodiment of a method for detecting model hallucinations of the present application, the method comprising:
[0022] S101: Obtain a first reply text obtained by inputting a question text into a trained first language model, and a second reply text obtained by inputting the question text into an untrained second language model.
[0023] Specifically, a first reply text is obtained by inputting the question text into a first language model that has been trained with a large amount of training data, and a second reply text is obtained by inputting the question text into an untrained second language model.
[0024] In one application, voice data is obtained and recognized to obtain question text, a first reply text is obtained by inputting the question text into a first language model that has been trained with a large amount of training data, and a second reply text is obtained by inputting the question text into an untrained second language model.
[0025] In another application, image data is obtained and scanned and recognized to obtain question text, obtain first reply text obtained by inputting the question text into a first language model trained with a large amount of training data, and obtain second reply text obtained by inputting the question text into an untrained second language model.
[0026] Optionally, the first language model and the second language model are large language models, which may include but are not limited to deep neural networks (DNNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory networks (LSTM), and generative pre-trained Transformer models, etc. No specific restrictions are made on the specific construction and deployment of the large language model.
[0027] S102: Obtain key information based on the question text and the first reply text, and obtain feature vectors of multiple output levels corresponding to the key information based on reference vectors of multiple output levels obtained during the output of the key information and the second reply text.
[0028] Specifically, after key information is obtained based on the question text and the first reply text, feature vectors of multiple output levels corresponding to the key information are obtained based on reference vectors of multiple output levels obtained during the output of the key information and the second reply text.
[0029] In one application method, key information is obtained by inputting the question text and the first reply text into a preset question template, and based on the reference vectors of multiple output levels obtained during the output of the key information and the second reply text, feature vectors of multiple output levels corresponding to the key information are obtained.
[0030] In another application method, the question text and the first reply text are input into a pre-trained intelligent agent to obtain key information of the intelligent agent's reply, and based on the reference vectors of multiple output levels obtained during the output of the key information and the second reply text, feature vectors of multiple output levels corresponding to the key information are obtained.
[0031] In some application scenarios, the multiple output layers include a self-attention layer, a neural network layer, and an output layer arranged in sequence.
[0032] S103: Based on the feature vectors of multiple output levels, an information vector corresponding to the key information is obtained, and based on the information vector, a hallucination detection result of the first reply text is determined.
[0033] Specifically, based on the feature vectors of multiple output levels, an information vector corresponding to the key information is obtained, and based on the information vector, a hallucination detection result of the first reply text output by the first language model is determined.
[0034] In one application, different weights are given to the feature vectors of each output level according to their importance, and then the feature vectors are weighted and summed to obtain an information vector corresponding to the fused key information. Based on the information vector, the hallucination detection result of the first reply text output by the first language model is determined.
[0035] In another application, after dimensionality reduction processing is performed on the feature vector of each output level, the main components are extracted as information vectors corresponding to the key information, and based on the information vectors, the hallucination detection result of the first reply text output by the first language model is determined.
[0036] The model hallucination detection method proposed in the present application obtains a first reply text obtained from a first language model that has been trained with a large amount of training data for a question text input, and obtains a second reply text obtained from a second language model that has not been trained for the question text input. After obtaining key information based on the question text and the first reply text, characteristic vectors of multiple output levels corresponding to the key information are obtained based on reference vectors of multiple output levels obtained during the output of the key information and the second reply text. Based on the characteristic vectors of the multiple output levels, an information vector corresponding to the key information is obtained. Based on the information vector, a hallucination detection result of the first reply text output by the first language model is determined. When obtaining relevant features such as the information vector corresponding to the key information, the outputs of the trained first language model and the untrained second language model are combined to make the judgment on whether a hallucination occurs more discriminative, thereby improving the accuracy of text hallucination detection.
[0037] See also Figure 2 , Figure 2 : is a flow chart of another embodiment of the model hallucination detection method of the present application. The method comprises:
[0038] S201: Obtain a first reply text obtained by inputting a question text into a trained first language model, and a second reply text obtained by inputting the question text into an untrained second language model.
[0039] Specifically, a first reply text is obtained by inputting the question text into a first language model that has been trained with a large amount of training data, and a second reply text is obtained by inputting the question text into an untrained second language model.
[0040] S202: Based on the question text and the first reply text, key information consisting of at least one keyword is obtained.
[0041] Specifically, at least one keyword in the question text and the first reply text is extracted to obtain corresponding key information. By extracting the keyword as key information, the content understanding and analysis capabilities can be enhanced, thereby improving the efficiency of subsequent text hallucination detection.
[0042] In an application scenario, the question text QU and the first reply text AU are input into a preset question template to obtain at least one keyword KUx extracted by the preset question template, thereby obtaining corresponding key information.
[0043] In a specific application scenario, a medical-related question text QU input by a user is obtained: "What medicine can I take for viral colds?", and a first reply text AU is obtained through a first language model LM1 trained with a large amount of medical field corpus: "Can I take amoxicillin for viral colds?" The preset question template is: "The following is a pair of questions and answers. Please extract all medical-related keywords in the answers, including but not limited to medical terms, indicator numbers and other information. When outputting, please output in list form for easy processing." The key information obtained by inputting the question text QU and the first reply text AU into the preset question template is "[viral colds, amoxicillin]".
[0044] S203: Obtain a first self-attention score vector, a first decoding feature vector, and a first output probability vector obtained by matching keywords at each output level during the first language model output process, and obtain a second self-attention score vector, a second decoding feature vector, and a second output probability vector obtained by matching keywords at each output level during the second language model output process.
[0045] Specifically, the first self-attention score vector, the first decoding feature vector and the first output probability vector obtained by matching the keywords at each output level in the first language model output process are obtained, and then it is determined whether there is a corresponding keyword in the second reply text. If so, the second self-attention score vector, the second decoding feature vector and the second output probability vector obtained by matching the corresponding keywords at each output level in the second language model output process are obtained.
[0046] It can be understood that when there is no corresponding keyword in the second reply text, the second self-attention score vector, the second decoding feature vector and the second output probability vector obtained by matching the corresponding keyword at each output level in the second language model output process are all empty sets.
[0047] In one application scenario, the multiple output layers include a self-attention layer, a neural network layer, and an output layer arranged in sequence.
[0048] In other application scenarios, the multiple output levels also include other network layers, and this application does not impose any specific restrictions on this.
[0049] S204: Based on the first self-attention score vector and the second self-attention score vector, obtain a first output vector corresponding to the keyword, based on the first decoding feature vector and the second decoding feature vector, obtain a second output vector corresponding to the keyword, based on the first output probability vector and the second output probability vector, obtain a third output vector corresponding to the keyword.
[0050] Specifically, based on the first self-attention score vector and the second self-attention score vector, the first output vector corresponding to the keyword is calculated, based on the first decoding feature vector and the second decoding feature vector, the second output vector corresponding to the keyword is calculated, and based on the first output probability vector and the second output probability vector, the third output vector corresponding to the keyword is calculated.
[0051] In one application scenario, the first self-attention score vector and the second self-attention score vector are averaged to obtain the first output vector corresponding to the keyword, the first decoding feature vector and the second decoding feature vector are averaged to obtain the second output vector corresponding to the keyword, and the first output probability vector and the second output probability vector are averaged to obtain the third output vector corresponding to the keyword. In other application scenarios, the first output vector, the second output vector, and the third output vector corresponding to the keyword can also be calculated in other ways, and this application does not impose specific restrictions on this.
[0052] S205: Based on the first output vector, the second output vector and the third output vector corresponding to all the keywords, feature vectors of multiple output levels corresponding to the key information are obtained.
[0053] Specifically, based on the first output vector, the second output vector and the third output vector corresponding to all the keywords, feature vectors of multiple output levels corresponding to the key information are calculated.
[0054] In one application scenario, step S205 specifically includes: fusing the first output vectors corresponding to all keywords to obtain the feature vector corresponding to the self-attention layer, fusing the second output vectors corresponding to all keywords to obtain the feature vector corresponding to the neural network layer, and fusing the third output vectors corresponding to all keywords to obtain the feature vector corresponding to the output layer.
[0055] Specifically, the first output vectors corresponding to all keywords are fused to calculate the feature vector corresponding to the self-attention layer, the second output vectors corresponding to all keywords are fused to calculate the feature vector corresponding to the neural network layer, and the third output vectors corresponding to all keywords are fused to calculate the feature vector corresponding to the output layer. By fusing the vectors, the model can capture data characteristics in more dimensions and improve model performance. In addition, by fusing output vectors at different levels, the generalization ability of the model can be significantly enhanced.
[0056] In a specific application scenario, the average of the first output vectors corresponding to all keywords KUx is taken to obtain the feature vector FU1 corresponding to the self-attention layer, whose vector size is the hidden neuron dimension × 1, and the number of hidden neurons is 768, so the feature vector FU1 is a 768×1 vector matrix. The average of the second output vectors corresponding to all keywords KUx is taken to obtain the feature vector FU2 corresponding to the neural network layer, whose vector size is the hidden neuron dimension × 1, and the number of hidden neurons is 768, so the feature vector FU2 is a 768×1 vector matrix. The average of the third output vectors corresponding to all keywords Kux is taken to obtain the feature vector FU3 corresponding to the output layer, whose vector size is the hidden neuron dimension × 1, and the number of hidden neurons is 30522, so the feature vector FU3 is a 30522×1 vector matrix.
[0057] S206: Based on the feature vectors of the multiple output levels, an information vector corresponding to the key information is obtained, and based on the information vector, a hallucination detection result of the first reply text is determined.
[0058] Specifically, based on the feature vectors of multiple output levels, an information vector corresponding to the key information is obtained, and based on the information vector, a hallucination detection result of the first reply text output by the first language model is determined.
[0059] In one implementation scenario, the hallucination detection result is obtained using a hallucination detection model, which includes a transformation network and an output network. The hallucination detection model is trained using a training sample set, which includes a sample text, a first training text obtained by inputting the sample text into a trained first language model and its corresponding training label, and a second training text obtained by inputting the sample text into an untrained second speech model.
[0060] In one embodiment, see Figure 3 , Figure 3 yes Figure 2 Step S206 in the flowchart corresponds to an implementation method. Step S206 specifically includes:
[0061] S301: Input the feature vectors of multiple output layers into the transformation network of the hallucination detection model to obtain the information vector corresponding to the key information.
[0062] Specifically, feature vectors of multiple output levels are input into a transformation network of a hallucination detection model to obtain information vectors corresponding to key information output by the transformation network.
[0063] In one application scenario, the transformation network includes a linear layer that matches each output level respectively, and step S301 includes: inputting the feature vector of each output level into the linear layer that matches the corresponding output level in the transformation network to obtain a transformation vector corresponding to the feature vector of each output level; based on the transformation vectors corresponding to all output levels, obtaining an information vector corresponding to the key information.
[0064] Specifically, the feature vector of each output level is transformed through a linear layer matching the corresponding output level in the transformation network to obtain a transformation vector corresponding to the feature vector of each output level after the transformation. Based on the transformation vectors corresponding to all output levels, the information vector corresponding to the final key information is obtained. Through feature transformation, more representative features can be extracted, thereby improving the performance and generalization ability of the model, and reducing the impact of outliers and noise on the model, thereby improving the stability and reliability of the model.
[0065] In a specific application scenario, the feature vector FU1 corresponding to the self-attention layer is passed through the linear layer L1 in the transformation network to obtain the transformed transformation vector FU1', the feature vector FU2 corresponding to the neural network layer is passed through the linear layer L2 in the transformation network to obtain the transformed transformation vector FU2', the feature vector FU3 corresponding to the output layer is passed through the linear layer L3 in the transformation network to obtain the transformed transformation vector FU3', and then the transformation vectors FU1', FU2' and FU3' are added to obtain the final information vector FU4.
[0066] S302: Input the information vector into the output network of the hallucination detection model to obtain the hallucination detection result of the first reply text.
[0067] Specifically, the information vector is input into the output network in the hallucination detection model to obtain the hallucination detection result of the first reply text output by the output network.
[0068] In one application scenario, the output network includes a linear layer and a probability layer, and step S302 includes: inputting the information vector into the linear layer of the output network to obtain the vector feature of the information vector; inputting the vector feature into the probability layer of the output network to obtain the hallucination probability value, and based on the hallucination probability value and its corresponding probability threshold, obtaining the hallucination detection result of the first reply text.
[0069] Specifically, the information vector is input into the linear layer of the output network to obtain the vector feature of the information vector output by the linear layer, the vector feature is input into the probability layer of the output network to obtain the hallucination probability value output by the probability layer, the obtained hallucination probability value is compared with a preset probability threshold, and the hallucination detection result of the first reply text is obtained. In this way, large-model hallucination detection can be completed by a small-scale neural network composed of only several linear layers, without the need for a large amount of labeled data, thereby improving the efficiency of text hallucination detection.
[0070] In a specific application scenario, the information vector FU4 is input into the linear layer L4 with an output dimension of 2 in the output network to obtain the vector feature of the information vector output by the linear layer L4, and the vector feature is input into the probability layer of the output network. The hallucination probability value PU is calculated by the Softmax function in the probability layer. If the hallucination probability value PU is lower than the preset probability threshold of 0.5, it is considered that the first reply text has a hallucination risk and is not accepted. In other application scenarios, the value of the probability threshold is set according to the specific situation, and this application does not impose specific restrictions on this.
[0071] Optionally, the training process of the hallucination detection model includes: obtaining a number of medical-related sample texts QT, and obtaining a first training text AT1 through a first language model LM1 that has been trained with a large amount of medical field corpus, and manually marking whether it is wrong, recorded as the manually marked answer LT, and at the same time obtaining a second training text AT2 obtained by the sample text QT through a second language model LM2 that has not been trained with medical corpus. For example, the sample text QT is: "What BMI is greater than what is considered obese?", the first training text AT1 is: "BMI of 30 and above is obese.", the second training text AT2 is: "BMI of 35 and above is obese.", and the manually marked training label LT is: "Error (BMI greater than 28 is obese)".
[0072] Furthermore, the sample text QT "BM1 greater than what is considered obese?" and the first training text AT1 "BMI of 30 and above is considered obese." are input into the preset question template "The following is a pair of questions and answers. Please extract all medical-related keywords in the answers, including but not limited to medical terms, indicator numbers and other information. When outputting, please output in list form for easy processing." The key information obtained is "[BMI, 30, obesity]", where keyword KT1 is BMI, keyword KT2 is 30, and keyword KT3 is obesity. The first self-attention score vector, the first decoding feature vector and the first output probability vector of each keyword KTx in the output process of the first language model LM1 are obtained. Then, it is determined whether the keyword KTx exists in the first In the second training text AT2 output by the second language model LM2, if it exists, the second self-attention score vector, the second decoding feature vector and the second output probability vector of the keyword KTx in the output process of the second language model LM2 are obtained, and the first self-attention score vector and the second self-attention score vector are averaged to obtain the first output vector, the first decoding feature vector and the second decoding feature vector are averaged to obtain the second output vector, and the first output probability vector and the second output probability vector are averaged to obtain the third output vector. If it does not exist, no operation is performed, that is, the first self-attention score vector is directly used as the first output vector, the first decoding feature vector is directly used as the second output vector, and the first output probability vector is used as the third output vector. Finally, the average of the first output vectors corresponding to all keywords KTx is taken as the feature vector F1 corresponding to the self-attention layer, the average of the second output vectors corresponding to all keywords KTx is taken as the feature vector F2 corresponding to the neural network layer, and the average of the third output vectors corresponding to all keywords KTx is taken as the feature vector F3 corresponding to the output layer.
[0073] Furthermore, the feature vector F1 is passed through a linear layer L1 whose output dimension is the number of hidden neurons to obtain a transformed transformation vector F1', the feature vector F2 is passed through a linear layer L2 whose output dimension is the number of hidden neurons to obtain a transformed transformation vector F2', and the feature vector F3 is passed through a linear layer L3 whose output dimension is the number of hidden neurons to obtain a transformed transformation vector F3'. Subsequently, the transformed vectors F1', F2' and F3' are added to obtain a final information vector F4, wherein the input dimension and output dimension of the linear layer through which F1 and F2 pass are both 768, and the input dimension of the linear layer through which F3 passes is 30522, and the output dimension is 768.
[0074] Further, the information vector F4 is input into the linear layer L4 with an output dimension of 2 to obtain a vector feature, and the vector feature is input into the probability layer, and the predicted probability value PT of the first training text AT1 is calculated by the Softmax function in the probability layer. Subsequently, according to the manually annotated answer LT, if it is wrong, the corresponding training label is 0, otherwise it is 1, and supervised training is performed according to the training label and the predicted probability value PT through the cross entropy loss function, and during training, only the linear layers L1, L2, L3 and L4 will be trained. For example, the predicted probability value PT is 0.3, and the corresponding training label is 0, then the cross entropy loss is calculated accordingly. The specific calculation process is the existing technology and will not be repeated here. The parameters of the linear layers L1-L4 are adjusted by the calculated losses.
[0075] See also Figure 4 , Figure 4 The schematic diagram of the structure of an embodiment of the model hallucination detection system of the present application is shown in FIG.
[0076] Specifically, the acquisition module 400 is used to acquire a first reply text obtained by inputting a question text into a trained first language model, and a second reply text obtained by inputting a question text into an untrained second language model.
[0077] The generation module 401 is connected to the acquisition module 400, and is used to obtain key information based on the question text and the first reply text, and to obtain feature vectors of multiple output levels corresponding to the key information based on the reference vectors of multiple output levels obtained during the output process of the key information and the second reply text.
[0078] The determination module 402 is connected to the generation module 401, and is used to obtain an information vector corresponding to the key information based on the feature vectors of multiple output levels, and determine the hallucination detection result of the first reply text based on the information vector.
[0079] Optionally, the hallucination detection result is obtained using a hallucination detection model, the hallucination detection model includes a transformation network and an output network, the hallucination detection model is trained using a training sample set, the training sample set includes a sample text, a first training text obtained by inputting the sample text into a trained first language model and its corresponding training label, and a second training text obtained by inputting the sample text into an untrained second speech model, and the determination module 402 is also used to input feature vectors of multiple output levels into the transformation network of the hallucination detection model to obtain an information vector corresponding to the key information; and input the information vector into the output network of the hallucination detection model to obtain the hallucination detection result of the first reply text.
[0080] Optionally, the transformation network includes a linear layer that matches each output level respectively, and the determination module 402 is also used to input the feature vector of each output level into the linear layer that matches the corresponding output level in the transformation network to obtain the transformation vector corresponding to the feature vector of each output level; based on the transformation vectors corresponding to all output levels, the information vector corresponding to the key information is obtained.
[0081] Optionally, the output network includes a linear layer and a probability layer, and the determination module 402 is also used to input the information vector into the linear layer of the output network to obtain the vector feature of the information vector; input the vector feature into the probability layer of the output network to obtain the hallucination probability value, and based on the hallucination probability value and its corresponding probability threshold, obtain the hallucination detection result of the first reply text.
[0082] Optionally, the multiple output layers include a self-attention layer, a neural network layer and an output layer arranged in sequence.
[0083] Optionally, the generation module 401 is also used to obtain key information consisting of at least one keyword based on the question text and the first reply text; obtain the first self-attention score vector, the first decoding feature vector and the first output probability vector obtained by matching the keyword at each output level in the first language model output process, and obtain the second self-attention score vector, the second decoding feature vector and the second output probability vector obtained by matching the keyword at each output level in the second language model output process; obtain the first output vector corresponding to the keyword based on the first self-attention score vector and the second self-attention score vector, obtain the second output vector corresponding to the keyword based on the first decoding feature vector and the second decoding feature vector, and obtain the third output vector corresponding to the keyword based on the first output probability vector and the second output probability vector; obtain the feature vectors of multiple output levels corresponding to the key information based on the first output vectors, the second output vectors and the third output vectors corresponding to all the keywords.
[0084] Optionally, the generation module 401 is also used to fuse the first output vectors corresponding to all keywords to obtain the feature vector corresponding to the self-attention layer, fuse the second output vectors corresponding to all keywords to obtain the feature vector corresponding to the neural network layer, and fuse the third output vectors corresponding to all keywords to obtain the feature vector corresponding to the output layer.
[0085] See also Figure 5 , Figure 5: It is a structural diagram of an embodiment of an electronic device of the present application. The electronic device 50 includes: a memory 501 and a processor 502 coupled to each other. The memory 501 stores program instructions (not shown), and the processor 502 is used to execute the program instructions to implement the method mentioned in any of the above embodiments. For the description of the relevant content, please refer to the detailed description of the above method embodiment, which will not be repeated here. Specifically, the electronic device 50 includes but is not limited to: a desktop computer, a laptop computer, a tablet computer, a server, etc., which is not limited here. In addition, the processor 502 can also be called a CPU (Center Processing Unit). The processor 502 may be an integrated circuit chip with signal processing capabilities. The processor 502 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 502 can be implemented by an integrated circuit chip.
[0086] See also Figure 6 , Figure 6 The computer-readable storage medium 60 stores program instructions 600 that can be executed by a processor, and the program instructions 600 are used to implement the method mentioned in any of the above embodiments.
[0087] In the several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation described above is only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0088] It should be noted that the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present implementation scheme.
[0089] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0090] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program codes.
[0091] The above description is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A model hallucination detection method, characterized in that: include: Obtaining a first reply text obtained by inputting the question text into a trained first language model, and a second reply text obtained by inputting the question text into an untrained second language model; Based on the question text and the first reply text, key information is obtained, and based on reference vectors of multiple output levels obtained during the output of the key information and the second reply text, feature vectors of multiple output levels corresponding to the key information are obtained; Based on the feature vectors of multiple output levels, an information vector corresponding to the key information is obtained, and based on the information vector, a hallucination detection result of the first reply text is determined.
2. The method according to claim 1, characterized in that The hallucination detection result is obtained by using a hallucination detection model, the hallucination detection model includes a transformation network and an output network, the hallucination detection model is trained by using a training sample set, the training sample set includes a sample text, a first training text obtained by inputting the sample text into a trained first language model and its corresponding training label, and a second training text obtained by inputting the sample text into an untrained second speech model; The step of obtaining an information vector corresponding to the key information based on the feature vectors of the multiple output levels, and determining a hallucination detection result of the first reply text based on the information vector, comprises: Inputting the feature vectors of multiple output levels into the transformation network of the hallucination detection model to obtain the information vector corresponding to the key information; The information vector is input into an output network of the hallucination detection model to obtain a hallucination detection result of the first reply text.
3. The method according to claim 2, characterized in that The transformation network includes a linear layer matched to each output level respectively, and the feature vectors of multiple output levels are input into the transformation network of the hallucination detection model to obtain the information vector corresponding to the key information, including: Inputting the feature vector of each output level into the linear layer matched with the corresponding output level in the transformation network to obtain the transformation vector corresponding to the feature vector of each output level; Based on the transformation vectors corresponding to all the output levels, an information vector corresponding to the key information is obtained.
4. The method according to claim 2, characterized in that: The output network includes a linear layer and a probability layer, and the step of inputting the information vector into the output network of the hallucination detection model to obtain a hallucination detection result of the first reply text includes: Inputting the information vector into the linear layer of the output network to obtain the vector feature of the information vector; The vector feature is input into the probability layer of the output network to obtain a hallucination probability value, and based on the hallucination probability value and its corresponding probability threshold, a hallucination detection result of the first reply text is obtained.
5. The method according to any one of claims 1 to 4, characterized in that: The multiple output layers include a self-attention layer, a neural network layer and an output layer arranged in sequence.
6. The method according to claim 5, characterized in that The step of obtaining key information based on the question text and the first reply text, and obtaining feature vectors of multiple output levels corresponding to the key information based on reference vectors of multiple output levels obtained during the output of the key information and the second reply text, includes: Based on the question text and the first reply text, obtaining key information consisting of at least one keyword; Obtaining a first self-attention score vector, a first decoding feature vector, and a first output probability vector obtained by matching the keyword at each output level during the output process of the first language model, and obtaining a second self-attention score vector, a second decoding feature vector, and a second output probability vector obtained by matching the keyword at each output level during the output process of the second language model; Based on the first self-attention score vector and the second self-attention score vector, obtain a first output vector corresponding to the keyword, based on the first decoding feature vector and the second decoding feature vector, obtain a second output vector corresponding to the keyword, based on the first output probability vector and the second output probability vector, obtain a third output vector corresponding to the keyword; Based on the first output vector, the second output vector and the third output vector corresponding to all the keywords, feature vectors of multiple output levels corresponding to the key information are obtained.
7. The method according to claim 6, characterized in that The step of obtaining feature vectors of multiple output levels corresponding to the key information based on the first output vector, the second output vector, and the third output vector corresponding to all the keywords includes: The first output vectors corresponding to all keywords are fused to obtain the feature vector corresponding to the self-attention layer, the second output vectors corresponding to all keywords are fused to obtain the feature vector corresponding to the neural network layer, and the third output vectors corresponding to all keywords are fused to obtain the feature vector corresponding to the output layer.
8. A model hallucination detection system, characterized in that: include: An acquisition module, used to acquire a first reply text obtained by inputting a question text into a trained first language model, and a second reply text obtained by inputting the question text into an untrained second language model; A generating module connected to the acquiring module, configured to obtain key information based on the question text and the first reply text, and obtain feature vectors of multiple output levels corresponding to the key information based on reference vectors of multiple output levels obtained during the output of the key information and the second reply text; A determination module is connected to the generation module and is used to obtain an information vector corresponding to the key information based on the feature vectors of multiple output levels, and determine a hallucination detection result of the first reply text based on the information vector.
9. An electronic device, characterized in that: The method comprises a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the method according to any one of claims 1 to 7.
Citation Information
Cited By
Knowledge-enhanced medical illusion static detection and correction method and system
CN120911443A
A knowledge-enhanced static detection and correction method and system for medical hallucinations
CN120911443B