This invention provides a hallucination detection method for a large
language model, comprising: extracting a set of first intrinsic features from the original response and constructing a first intrinsic uncertainty representation; inputting the input to a response hallucination
detector and outputting a first hallucination probability distribution; guiding the large
language model to generate a self-judgment result from the original response; extracting second intrinsic features from the self-judgment
generation process and constructing a second intrinsic uncertainty representation; inputting the input to a judgment hallucination
detector and outputting a second hallucination probability distribution; constructing a
logical relationship between the hallucination probability distributions based on the symbolic
semantics of the self-judgment result and quantifying it as a logical loss; constructing a total
loss function by combining the classification losses of the response hallucination
detector and the judgment hallucination detector respectively, and performing joint optimization training. This invention also provides a hallucination detection device, storage medium, and electronic device for a large
language model. Therefore, by constructing a dual-view detection framework and a
mutual learning mechanism of logical constraints, this invention can achieve more accurate and robust hallucination detection.