The invention discloses a large
language model watermark embedding and detecting method based on attention head
perception, and relates to the technical field of
natural language processing and
digital watermarking. The method comprises the following four core steps of: firstly, screening an attention head set which plays a key role in factual information coding through a'pure-
pollution-
recovery 'three-stage causal intervention experiment; secondly, a binary linear fact
detector is trained based on the set, and accurate distinguishing of token factuality and non-factuality is achieved; then, in a model reasoning stage, a dynamic green
list is only generated for non-factual tokens,
watermark hidden embedding is completed through logits bias intervention, and bias intensity is adaptively adjusted along with text fact density; and finally, in a detection stage, reproducing an
embedding process to screen non-factual tokens, and quantifying an abnormal degree by counting a matching ratio and a z
score to realize accurate detection of the
watermark. The text factual accuracy is prevented from being damaged through the attention head
perception mechanism, the
watermark robustness and concealment are improved through the dynamic green
list design, the method is suitable for scenes such as copyright tracing and authenticity
verification of the text generated by the large
language model, and the
engineering feasibility is high.