Large language model generation content detection method based on sentence semantics watermark injection

Through the watermark injection method based on sentence semantics, sentences generated by large language models are semantically marked, which solves the problems of fragile watermark capabilities and poor generalization capabilities in the prior art, and achieves fast and accurate detection and content traceability.

CN119939544AActive Publication Date: 2025-05-06HEFEI UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510087710.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The existing large language model generation content detection methods have problems such as fragile watermarking ability and poor generalization ability, and cannot effectively resist sentence-level attacks and distribution external generalization.

Method used

Using a watermark injection method based on sentence semantics, the sentences generated by the large language model are semantically marked, thereby injecting hidden watermark marks into the text.

Benefits of technology

It realizes the rapid and accurate detection of content generated by large language models under various text attack conditions, enhances the concealment and tamper resistance of watermarks, reduces the error detection rate and missed detection rate, and ensures the traceability of content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939544A_ABST
    Figure CN119939544A_ABST
Patent Text Reader

Abstract

The invention discloses a large language model generation content detection method based on sentence semantics watermark injection, and relates to the field of information processing, and the method comprises the following steps: 1, constructing and training a watermark model used for determining sentence marks; 2, acquiring a prompt text for generating a watermark text; 3, marking sentences generated by the large language model to inject watermarks; and 4, extracting a mark of each sentence of the text, and verifying whether the text is generated by the large language model or not. According to the method, the sentences are marked based on the sentence semantics, and the watermark marks are injected into the generated contents by screening the sentences generated by the large language model, so that the invisible watermark marks are added into the generated contents of the large language model, and the generated contents of the large language model are further detected and prevented from being abused.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information processing, and in particular to a method for detecting content generated by a large language model. Background Art

[0002] Large language models have demonstrated impressive generative capabilities and have been widely used in various applications, such as ChatGPT, Copilot, Claude, etc. With the deep collaboration of large language models in content generation, the potential risks (such as misleading information, copyright issues) when using generated content have also become critical. Text watermarking technology can be used for both information identification and copyright tracking, and has become a hot issue in the application field of large language models.

[0003] Text watermarking for large language models aims to embed implicitly identifiable information into generated content, where the design of "red" and "green" groups is a common example. Although greater progress has been made, existing methods still suffer from fragile watermarking capabilities and poor generalization capabilities. Specifically, existing vocabulary-level algorithms inject watermarks by constructing vocabulary tags. Since attacks on the sentence level of watermarked text will replace text vocabulary and modify the text structure, methods that inject watermarks at the vocabulary level cannot resist attacks at the sentence level. Existing sentence-level algorithm methods usually construct sentence tags based on the distance between the generated sentence and a pre-defined sentence tag anchor. Since the pre-defined sentence tag anchor requires the distribution of the content generated by the large language model to be obtained in advance, existing sentence-level algorithms have the problem of weak generalization capabilities outside the distribution. Summary of the invention

[0004] The present invention aims to solve the deficiencies of the above-mentioned prior art and proposes a method for detecting content generated by a large language model based on sentence semantic injection watermark, so as to inject watermarks into the large language model by selecting sentences, thereby adding invisible watermarks to the large language model, thereby being able to quickly and accurately test the content generated by the large language model under various text attack conditions and effectively avoid the harm caused by the content generated by the large language model.

[0005] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme:

[0006] The invention discloses a method for detecting content generated by a large language model based on sentence semantics watermark injection, which comprises the following steps:

[0007] S1. Build and train a watermark model for determining sentence tags ;

[0008] S11. Get sentence dataset ,in, Represents the sentence in the sentence dataset D Sentences, , Represents the sentence in the sentence dataset D Sentences, , express The number of sentences in

[0009] Will and The input is processed into the embedding model and the Embed and Embed ;

[0010] S12. Based on sentence semantics, construct a watermark model for determining sentence tags , including: encoder , code table and decoder , and and Process and obtain The reconstruction representation and The reconstruction representation ;

[0011] S13. Constructing watermark model The total loss function ;

[0012] S14, using the back propagation algorithm to train the watermark model , and calculate the overall loss function To update the model parameters until the overall loss function Until convergence, the trained watermark model is obtained. ;

[0013] S2. Get the prompt text to be embedded in the watermark ,in, represents the i-th sentence in the prompt text, , Indicates the prompt text The number of sentences in

[0014] S3. Mark the sentences generated by the large language model for use in prompt text Inject watermark into

[0015] S31. Define the current loop variable as t and initialize it ;

[0016] Get the tth sentence in the prompt text and use it as the sentence for the t-1th loop, denoted as ;

[0017] S32. Using the trained watermark model generate Mark ;

[0018] S33, will Set as a random seed, so that the code table is The numbers of the n potential representations are divided into class A tags and class B tags;

[0019] S34, use the large language model to analyze the text of the first t-1 cycles Process and generate sentences for the tth cycle ;

[0020] S35. Using the trained watermark model generate Mark ;

[0021] S36, Judgment Whether it belongs to the A-type mark, if it does, return to S34 to execute sequentially, otherwise, execute S37;

[0022] S37, Order Assign to After that, return to S32 and execute sequentially until So that the A-type marker is injected as a watermark into the sentences generated by the large language model, where T represents the number of sentences expected to be generated;

[0023] S4, extract the token of each sentence in the text and verify whether the text is generated by the large language model;

[0024] S41. For a given new text ,make ;

[0025] S42, according to the process of S32 and S35 Process and judge Is the token of the t-th sentence in the A-type token? If so, the counter is incremented by 1; otherwise, it is not counted;

[0026] S43, Order Assign to After that, return to S42 and execute sequentially until So far, the final count value is obtained;

[0027] S44, if the count value exceeds Half of the total number of sentences in Generate text for a large language model, otherwise, It’s not the big language models that generate text.

[0028] The method for generating content detection based on a large language model with sentence semantic watermark injection according to the present invention is also characterized in that the encoder in S12 and decoder All of them are composed of linear layers, and the code table By n d-dimensional potential representation Composition; among them, represents the kth potential representation, d represents the dimension of each potential representation, and n represents the code table The number of potential representations in ;

[0029] S121, will and Input to the encoder Processed in Potential representation of and Potential representation of ;

[0030] S122, using formula (1) to identify the code table Zhongyu The closest potential representation and The closest potential representation , and the corresponding numbers i and j are respectively Sentences The mark and Sentences Marking;

[0031] (1)

[0032] S123, will and Input to the decoder Processed in The reconstruction representation and The reconstruction representation .

[0033] Further, the S13 includes:

[0034] S131. Use formula (2) to construct a semantic loss function that captures the text embedding semantics :

[0035] (2)

[0036] In formula (2), represents the weight coefficient, Indicates stopping the gradient operator;

[0037] S132. Use formula (3) to construct a semantic consistency loss function :

[0038] (3)

[0039] In formula (3), represents the threshold value, Represents similarity calculation, Indicates taking the maximum value, represents the indicator function;

[0040] S133. Use formula (4) to construct the total loss function :

[0041] (4)

[0042] In formula (4), Represents the weight coefficient.

[0043] An electronic device of the present invention includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the method, and the processor is configured to execute the program stored in the memory.

[0044] The present invention provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the method when executed by a processor.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] 1. The present invention injects watermarks based on sentence semantics and marks sentences using sentence semantic features without significantly modifying the surface form of the text (such as vocabulary). This method not only hides the watermark in the text, but also avoids the destruction of the naturalness of the language, ensuring that the generated content still has high readability and fluency. In addition, the semantic-based watermark mechanism has high robustness and can retain watermark information after the text has been edited, transcoded or other common transformations, thereby enhancing the practicality and stability of the watermark.

[0047] 2. The present invention uses a watermark model to semantically mark sentences, embeds watermarks in the text generation stage, and uses the uniqueness and uniqueness of the mark to verify in the detection stage. This marking mechanism based on sentence semantics has high accuracy and can significantly reduce the false detection rate and missed detection rate. Therefore, the present invention is suitable for large language model generation scenarios, and can quickly identify and confirm whether there is content generated by a large language model, providing strong technical support for content management and copyright protection.

[0048] 3. The present invention achieves content traceability by injecting hidden watermarks into generated content, thereby effectively reducing the risk of abuse. By detecting watermarks, it is possible to identify the source of content generation, clarify the attribution of responsibility, and provide a basis for identifying potential abuse. The detection mechanism of the present invention not only achieves accurate identification of abuse at the technical level, but also enhances the ability to supervise content generated by large language models at the social level, providing a guarantee for building a trustworthy generated content ecosystem. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 The present invention is a flow chart of a method for generating content detection based on a large language model that injects watermarks based on sentence semantics. DETAILED DESCRIPTION

[0050] In this embodiment, a method for detecting content generated by a large language model based on sentence semantics watermarking is to mark sentences based on text semantics, and to inject watermarks by selecting sentences generated by a large language model, thereby adding invisible watermarks to the content generated by the large language model, thereby verifying the source of the detected text and preventing the content generated by the large language model from being abused, and includes: constructing and training a watermark model for determining sentence marks, obtaining prompt text for generating watermark text, marking sentences generated by the large language model to inject watermarks, extracting marks of each sentence in the text, and verifying whether the text is generated by the large language model; specifically, Figure 1 As shown, the method includes:

[0051] S1. Build and train a watermark model for determining sentence tags ;

[0052] S11. Get sentence dataset ,in, Represents the sentence in the sentence dataset D Sentences, , Represents the sentence in the sentence dataset D Sentences, , express The number of sentences in

[0053] Will and The input is processed into an embedding model, where the embedding model can use an existing pre-trained model, such as BGE-M3, to obtain Embed and Embed .

[0054] S12. Based on sentence semantics, construct a watermark model for determining sentence tags , including: encoder , code table and decoder , and and Process and obtain The reconstruction representation and The reconstruction representation ;

[0055] Among them, the encoder and decoder All of them are composed of linear layers, and the code table By n d-dimensional potential representation Composition; among them, represents the kth potential representation, d represents the dimension of each potential representation, and n represents the code table The number of potential representations in , n and d can be set to 64 and 1000 respectively;

[0056] S121, will and Input to the encoder Processed in Potential representation of and Potential representation of ;

[0057] S122, using formula (1) to identify the code table Zhongyu The closest potential representation and The closest potential representation , and the corresponding numbers i and j are respectively Sentences The mark and Sentences Marking;

[0058] (1)

[0059] S123, will and Input to the decoder Processed in The reconstruction representation and The reconstruction representation .

[0060] S13. Constructing watermark model The total loss function ;

[0061] S131. Use formula (2) to construct a semantic loss function that captures the text embedding semantics :

[0062] (2)

[0063] In formula (2), Represents the weight coefficient, which can be set to 0.25. Indicates stopping the gradient operator;

[0064] S132. Use formula (3) to construct a semantic consistency loss function :

[0065] (3)

[0066] In formula (3), represents the threshold value, Indicates similarity calculation, such as cosine similarity, Indicates taking the maximum value, represents the indicator function;

[0067] S133. Use formula (4) to construct the total loss function :

[0068] (4)

[0069] In formula (4), Represents the weight coefficient.

[0070] S14, using the back propagation algorithm to train the watermark model , and calculate the overall loss function To update the model parameters until the overall loss function Until convergence, the trained watermark model is obtained. .

[0071] S2, obtaining a prompt text for generating a watermark text;

[0072] Get the hint text to be embedded in the watermark ,in represents the i-th sentence in the prompt text, , Indicates the prompt text The number of sentences in .

[0073] S3, marking the sentences generated by the large language model to inject watermarks;

[0074] S31. Define the current loop variable as t and initialize it ;

[0075] Get the tth sentence in the prompt text and use it as the sentence for the t-1th loop, denoted as ;

[0076] S32. Using the trained watermark model generate Mark ;

[0077] S33, will Set as a random seed, so that the code table is The numbers of the n potential representations are divided into class A tags and class B tags, where class A tags can be represented as "green" tags and class B tags can be represented as "red" tags.

[0078] S34, use the large language model to analyze the text of the first t-1 cycles Process and generate sentences for the tth cycle ;

[0079] S35. Using the trained watermark model generate Mark ;

[0080] S36, Judgment Whether it belongs to the A-type mark, if it does, return to S34 to execute sequentially, otherwise, execute S37;

[0081] S37, Order Assign to After that, return to S32 and execute sequentially until So far, T represents the number of sentences expected to be generated, so that the class A marker is injected as a watermark into the text generation process of the large language model.

[0082] S4, extract the token of each sentence in the text and verify whether the text is generated by the large language model;

[0083] S41. For a given new text ,make ;

[0084] S42, according to the process of S32 and S35 Processing, judging Is the token of the t-th sentence in the A-type token? If so, the counter is incremented by 1; otherwise, it is not counted;

[0085] S43, Order Assign to After that, return to S42 and execute sequentially until So far, the final count value is obtained;

[0086] S44, if the count value exceeds Half of the total number of sentences in Generate text for a large language model, otherwise, It’s not the big language models that generate text.

[0087] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.

[0088] In this embodiment, a computer-readable storage medium stores a computer program on the computer-readable storage medium, and the computer program executes the steps of the above method when executed by a processor.

[0089] In summary, the present invention provides a method for detecting content generated by a large language model based on sentence semantic watermark injection, which aims to solve the problem that the generated content in the prior art is difficult to track and detect. By constructing a watermark model, the candidate sentences generated by the large language model are semantically marked, thereby injecting a hidden watermark mark into the text. The present invention uses a watermark injection method at the semantic level to ensure that the watermark has no effect on the naturalness and fluency of the text, while enhancing the concealment and anti-tampering ability of the watermark. The method completes watermark embedding in the text generation stage, and realizes content source verification by extracting sentence tags in the detection stage, and has the characteristics of high efficiency, accuracy and reliability. The watermark embedding mechanism of the present invention is more flexible, can adapt to a variety of editing and transformation scenarios, and significantly reduces the false detection rate and missed detection rate. In addition, the present invention can effectively prevent content abuse and improve the security and credibility of generated content by giving traceability to the content generated by the large language model. This technology is not only of great significance for maintaining intellectual property rights and regulating the use of generated content, but also provides a new solution for building a credible artificial intelligence generated content ecosystem.

Claims

1. A method for detecting content generated by a large language model based on sentence semantic watermark injection, characterized in that: The following steps are involved: S1. Build and train a watermark model for determining sentence tags ; S2. Get the prompt text to be embedded in the watermark ,in, represents the i-th sentence in the prompt text, , Indicates the prompt text The number of sentences in S3, based on watermark model , marking sentences generated by the large language model for use in prompt text Inject watermark into S4. For a given new text ,extract , and verify whether the text is generated by the large language model.

2. The method for detecting content generated by a large language model based on sentence semantic watermark injection according to claim 1 is characterized in that: The S1 includes: S11. Get sentence dataset ,in, Represents the sentence in the sentence dataset D Sentences, , Represents the sentence in the sentence dataset D Sentences, , express The number of sentences in Will and The input is processed into the embedding model and the Embed and Embed ; S12. Based on sentence semantics, construct a watermark model for determining sentence tags , including: encoder , code table and decoder , and and Process and obtain The reconstruction representation and The reconstruction representation ; S13. Constructing watermark model The total loss function ; S14, using the back propagation algorithm to train the watermark model , and calculate the overall loss function To update the model parameters until the overall loss function Until convergence, the trained watermark model is obtained. .

3. The method for detecting content generated by a large language model based on sentence semantic watermark injection according to claim 2 is characterized in that: The encoder in the S12 and decoder All of them are composed of linear layers, and the code table By n d-dimensional potential representation Composition; among them, represents the kth potential representation, d represents the dimension of each potential representation, and n represents the code table The number of potential representations in ; S121, will and Input to the encoder Processed in Potential representation of and Potential representation of ; S122, using formula (1) to identify the code table Zhongyu The closest potential representation and The closest potential representation , and the corresponding numbers i and j are respectively Sentences The mark and Sentences Marking; (1) S123, will and Input to the decoder Processed in The reconstruction representation and The reconstruction representation .

4. The method for detecting content generated by a large language model based on sentence semantic watermark injection according to claim 3 is characterized in that: The S13 includes: S131. Use formula (2) to construct a semantic loss function that captures the text embedding semantics : (2) In formula (2), represents the weight coefficient, Indicates stopping the gradient operator; S132. Use formula (3) to construct a semantic consistency loss function : (3) In formula (3), represents the threshold value, Represents similarity calculation, Indicates taking the maximum value, represents the indicator function; S133. Use formula (4) to construct the total loss function : (4) In formula (4), Represents the weight coefficient.

5. The method for detecting content generated by a large language model based on sentence semantic watermark injection according to claim 4 is characterized in that: The S3 includes: S31. Define the current loop variable as t and initialize it ; Get the tth sentence in the prompt text and use it as the sentence for the t-1th loop, denoted as ; S32. Using the trained watermark model generate Mark ; S33, will Set as a random seed, so that the code table is The numbers of the n potential representations are divided into class A tags and class B tags; S34, use the large language model to analyze the text of the first t-1 cycles Process and generate sentences for the tth cycle ; S35. Using the trained watermark model generate Mark ; S36, Judgment Whether it belongs to the A-type mark, if it does, return to S34 to execute sequentially, otherwise, execute S37; S37, Order Assign to After that, return to S32 and execute sequentially until So far, the A-type marker is injected as a watermark into the sentences generated by the large language model, where T represents the number of sentences expected to be generated.

6. The method for detecting content generated by a large language model based on sentence semantic watermark injection according to claim 5 is characterized in that: The S4 includes: S41. Order ; S42, according to the process of S32 and S35 Process and judge Is the token of the t-th sentence in the A-type token? If so, the counter is incremented by 1; otherwise, it is not counted; S43, Order Assign to After that, return to S42 and execute sequentially until So far, the final count value is obtained; S44, if the count value exceeds Half of the total number of sentences in Generate text for a large language model, otherwise, It’s not the big language models that generate text.

7. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the large language model generated content detection method described in any one of claims 1-6, and the processor is configured to execute the program stored in the memory.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for detecting content generated by a large language model described in any one of claims 1 to 6 are executed.

Citation Information

Patent Citations

  • Machine learning based models for labelling text data

    CA3237882A1

  • Method and device for embedding and detecting watermark in document by using computer system

    CN101957810A

  • Database watermark embedding method and system oriented to text type data and storage medium

    CN118153007A

  • Sentence semantics-based watermarking method for large language model transfer attack

    CN118821086A