Watermark addition method, watermark detection method and watermark addition model training method

By dynamically adjusting word classification and probability distribution, the problem of unstable effects of existing watermark addition methods in multiple text situations is solved, and stable watermark addition is achieved and the accuracy of watermark detection is improved.

CN119962541BActive Publication Date: 2025-07-22IFLYTEK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510437524.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-22
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The existing watermark addition method uses a fixed word list ratio when driving the model, and cannot adapt to text in various situations, resulting in unstable watermark addition effect and affecting subsequent watermark detection effects.

Method used

By obtaining the historical word elements and current word elements probability distributions of the text generation model, the split model and deviation model are used to determine the word elements classification parameters and probability deviation parameters, dynamically adjust the category and probability distribution of word elements, and introduce a watermark addition module for updates to generate a stable watermark addition effect.

Benefits of technology

It realizes stable watermark addition in various situations, improves the accuracy and reliability of watermark detection, and avoids interference with the text generation model by fixed proportion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962541B_ABST
    Figure CN119962541B_ABST
Patent Text Reader

Abstract

The present invention provides a watermark addition method, a watermark detection method, and a watermark addition model training method, which relate to the field of computer vision technology. A splitting model is introduced to determine the token classification parameters according to historical tokens, so that the proportion of tokens of the same category in the token dictionary corresponding to different outputs of the text generation model is different, and it can be applied to the generated text in various situations, avoiding forcibly setting a fixed proportion to damage the accuracy and usability of the generated content of the text generation model, making the watermark addition effect stable, and thus ensuring the subsequent watermark detection effect. Moreover, a deviation model is also introduced to determine the probability deviation parameters of different token categories in the token dictionary according to historical tokens, so that the watermark addition module can update the first token probability distribution by combining the probability deviation parameters and the token classification parameters, which can change the probability values of tokens of different token categories in the token dictionary being selected, and further improve the subsequent watermark detection effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular, to a watermark addition method, a watermark detection method, and a watermark addition model training method. Background Art

[0002] With the emergence and wide application of large language models (LLMs), such as text generation, translation, summarization, etc., the abuse of LLMs has also been escalating. For example, the generation of fake news and web content, users may use LLMs to cheat in academic writing and programming assignments, and the text generated by LLMs on social media platforms may be used for social engineering, etc. Moreover, the quality of the text generated by LLMs varies. Therefore, it has become increasingly important to track the usage of the text generated by LLMs and determine whether the text belongs to the text generated by LLMs.

[0003] In the prior art, to track the text generated by LLMs, watermarks are usually added to the text generated by LLMs. The methods of adding watermarks usually include data-driven, model-driven, and post-processing methods, etc. Data-driven usually relies on background insertion, adding a small number of watermarked samples to the dataset of LLMs, allowing the model to implicitly learn, and setting relevant loss functions. Model-driven manipulates the distribution of logits or token sampling of the output layer of the LLM during the inference process, divides the tokens into two lists, and marks them as red and green respectively, and encodes the watermark information through the statistical characteristics of the words in the green list. Post-processing refers to the technology of embedding watermarks by processing the text after the LLM outputs the text. This method usually works in a pipeline as an independent module together with the output of the generation model.

[0004] To determine whether the text belongs to the text generated by LLMs, it is necessary to detect whether there is a watermark in the text. The existing watermark detection method corresponding to the model-driven watermark addition method is a watermark detection method based on interpretable p-values, which identifies the watermark by statistically analyzing the red and green marks in the text, thereby calculating the significance of the p-value.

[0005] However, in the existing watermark addition methods, when using the model-driven approach, the proportion of the words in the pre-set green list and red list in the model token list is applied. The fixed proportion cannot be suitable for the text in various situations, resulting in unstable watermark addition effects and also affecting the subsequent watermark detection effects. Summary of the Invention

[0006] The present invention provides a watermark addition method, a watermark detection method, and a watermark addition model training method to solve the defects existing in the related technologies.

[0007] The present invention provides a watermark adding method, including:

[0008] Obtaining a historical token output by a text generation model and a first token probability distribution output by a text generation unit in the text generation model at the current time;

[0009] Inputting the historical token into a splitting model and a deviation model of a watermark adding model in the text generation model to obtain a token classification parameter output by the splitting model and a probability deviation parameter of different token categories in a token dictionary output by the deviation model; the token classification parameter is used to classify each token in the token dictionary;

[0010] Inputting the first token probability distribution, the token classification parameter, and the probability deviation parameter into a watermark adding module of the watermark adding model to obtain a second token probability distribution with a watermark added output by the watermark adding module;

[0011] Determining a current token output by the text generation model at the current time based on the second token probability distribution.

[0012] According to the watermark adding method provided by the present invention, the watermark adding module is specifically configured to:

[0013] Determining a current category of each token in the token dictionary based on the token classification parameter;

[0014] Updating the first token probability distribution based on the probability deviation parameter and the current category of each token to obtain the second token probability distribution.

[0015] According to the watermark adding method provided by the present invention, the token classification parameter includes a random seed and distribution parameters of a preset distribution;

[0016] Correspondingly, the watermark adding module is specifically configured to:

[0017] Based on the random seed, rearranging the positions of each token in the token dictionary and sampling probability values equal in number to the tokens in the token dictionary from the preset distribution;

[0018] One-to-one corresponding the rearranged positions of each token in the token dictionary to the probability values, and determining the current category of each token in the token dictionary based on the probability values corresponding to the rearranged positions of each token in the token dictionary.

[0019] According to the watermark adding method provided by the present invention, each token in the token dictionary includes a watermark token and a non-watermark token;

[0020] Accordingly, the probability deviation parameter includes the gain parameter corresponding to the watermark token and the loss parameter of the non-watermark token.

[0021] The present invention also provides a watermark detection method, including:

[0022] Determine the window to be detected of the text to be detected and the full text including the window to be detected and its previous text in the text to be detected, and obtain each token segment with gradually increasing length in the full text;

[0023] For any token segment, based on the text generation unit in the text generation model, determine the third token probability distribution of the last token in the token segment, and apply the above watermark addition method to determine the fourth token probability distribution after adding the watermark and each token category in the token dictionary;

[0024] Based on the fourth token probability distribution, calculate the perplexity information entropy of the last token, and based on the perplexity information entropy, calculate the proportion weight of the last token in the hypothesis test;

[0025] Based on the number of watermark tokens in the window to be detected, each token category in the token dictionary corresponding to each token in the window to be detected, and the proportion weight, take the score of the hypothesis test as the detection index, and judge whether the window to be detected contains a watermark.

[0026] According to a watermark detection method provided by the present invention, the step of judging whether the window to be detected contains a watermark based on the number of watermark tokens in the window to be detected, each token category in the token dictionary corresponding to each token in the window to be detected, and the proportion weight, with the score of the hypothesis test as the detection index, includes:

[0027] Based on each token category in the token dictionary corresponding to each token in the window to be detected, calculate the proportion of watermark tokens in the token dictionary corresponding to each token in the window to be detected;

[0028] Based on the number of watermark tokens in the window to be detected, the proportion of watermark tokens corresponding to each token in the window to be detected, and the proportion weight, calculate the score of the hypothesis test;

[0029] If the score of the hypothesis test is greater than a preset threshold, it is determined that the window to be detected contains a watermark.

[0030] The present invention also provides a watermark addition model training method, including:

[0031] Obtain a text sample, input the text sample into an initial text generation model, obtain the sample probability distribution output by the text generation unit in the initial text generation model and the first type of token output by the initial text generation model, and determine the second type of token based on the sample probability distribution;

[0032] Based on the above watermark detection method, calculate the first score of the hypothesis test corresponding to the first type of token and the second score of the hypothesis test corresponding to the second type of token respectively;

[0033] Calculate a semantic loss based on the first type of token and the second type of token, and calculate a detection loss based on the first score and the second score;

[0034] Train the initial watermark addition model in the initial text generation model based on the semantic loss and the detection loss to obtain a trained watermark addition model.

[0035] According to a watermark addition model training method provided by the present invention, the calculating the semantic loss based on the first type of token and the second type of token includes:

[0036] Extract the first sentence vector in the first type of token and the second sentence vector in the second type of token respectively;

[0037] Calculate the semantic loss based on the first sentence vector and the second sentence vector.

[0038] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements any one of the above watermark addition methods, or watermark detection methods, or watermark addition model training methods.

[0039] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the above watermark addition methods, or watermark detection methods, or watermark addition model training methods.

[0040] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements any one of the above watermark addition methods, or watermark detection methods, or watermark addition model training methods.

[0041] The watermark addition method, watermark detection method, and watermark addition model training method provided by the present invention introduce a splitting model in the method to determine the token classification parameters according to historical tokens, so that the proportion of tokens of the same category in the token dictionary corresponding to different outputs of the text generation model is different, and it can be applicable to the generated text in various situations, avoiding forcibly setting a fixed proportion to damage the accuracy and usability of the generated content of the text generation model, making the watermark addition effect stable, and thus ensuring the subsequent watermark detection effect. Moreover, the method also introduces a deviation model to determine the probability deviation parameters of different token categories in the token dictionary according to historical tokens, so that the watermark addition module updates the first token probability distribution by combining the probability deviation parameters and the token classification parameters, which can change the probability values of tokens of different token categories being selected in the token dictionary, and further improve the subsequent watermark detection effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 is one of the flow diagrams of the watermark addition method provided by the present invention.

[0044] Figure 2 is the second flow diagram of the watermark addition method provided by the present invention.

[0045] Figure 3 is one of the flow diagrams of the watermark detection method provided by the present invention.

[0046] Figure 4 is the second flow diagram of the watermark detection method provided by the present invention.

[0047] Figure 5 is one of the flow diagrams of the watermark addition model training method provided by the present invention.

[0048] Figure 6 is the second flow diagram of the watermark addition model training method provided by the present invention.

[0049] Figure 7 is the structural diagram of the watermark addition device provided by the present invention.

[0050] Figure 8 is the structural diagram of the watermark detection device provided by the present invention.

[0051] Figure 9It is a schematic structural diagram of the watermark addition model training device provided by the present invention.

[0052] Figure 10 It is a schematic structural diagram of the electronic device provided by the present invention. Specific embodiments

[0053] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] Due to the existing watermark addition method, when the model is driven, the green list and the red list in the model token list are divided with a fixed ratio, and the division is rigid. It fails to set exclusive red and green partitions for the current output based on the output historical tokens, and thus cannot be suitable for texts in various situations. Moreover, when the model is driven, there is only a reward for the green area and no penalty for the red area, resulting in the elements in the originally high-probability red area still having a high probability, leading to unstable watermark addition effects and affecting subsequent watermark detection effects.

[0055] Based on this, an embodiment of the present invention provides a watermark addition method.

[0056] Figure 1 It is a schematic flowchart of a watermark addition method provided in an embodiment of the present invention. As Figure 1 shown, the method includes:

[0057] S11, obtaining the historical tokens output by the text generation model and the first token probability distribution of the current output of the text generation unit in the text generation model;

[0058] S12, inputting the historical tokens into the splitting model and the deviation model of the watermark addition model in the text generation model, and obtaining the token classification parameters output by the splitting model and the probability deviation parameters of different token categories in the token dictionary output by the deviation model; the token classification parameters are used to classify each token in the token dictionary;

[0059] S13, inputting the first token probability distribution, the token classification parameters and the probability deviation parameters into the watermark addition module of the watermark addition model, and obtaining the second token probability distribution with watermark added output by the watermark addition module;

[0060] S14, determining the current token of the current output of the text generation model based on the second token probability distribution.

[0061] Specifically, in the watermark addition method provided in the embodiments of the present invention, the execution subject is a watermark addition device, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., and no specific limitation is made here.

[0062] First, step S11 is executed to obtain the historical token output by the text generation model and the first token probability distribution of the current output of the text generation unit in the text generation model. The text generation model adopted can include a text generation unit and a watermark addition model connected in sequence. The text generation unit can be an LLM or other models or tools with text generation functions, and no specific limitation is made here.

[0063] It can be understood that the text generation unit takes the user interaction content and the output historical tokens as inputs and generates text content through autoregressive decoding. In each output, it includes the probability distribution of each token in the token dictionary of that time. Therefore, the first token probability distribution of the current output refers to the probability distribution of each token in the current token dictionary. The token dictionary can be the token set of the text generation unit and is a reference database for the text generation unit to compare and output the optimal probability of tokens.

[0064] It can be understood that after obtaining the first token probability distribution of the current output of the text generation unit, the token corresponding to the optimal probability in the first token probability distribution is not directly used as the current token of the current output. Instead, it is necessary to use the watermark addition model and obtain the current token of the current output of the text generation model through the operations of steps S12 - S14.

[0065] The historical tokens output by the text generation model can include one or more. Each historical token is the token probability distribution of the current output of the text generation unit and is obtained by combining the tokens output before that time and performing the operations of steps S12 - S14. In particular, the historical tokens can also be 0. At this time, the current output is the first output of the text generation model. Since there are no historical tokens in the first output and steps S12 - S14 cannot be continued, the optimal probability can be directly selected from the token probability distribution of the first output, and the token corresponding to this optimal probability is used as the token of the first output.

[0066] Then, step S12 is executed. The watermark addition model in the text generation model includes a split model (SplitModel), a bias model (BiasModel), and a watermark addition module.

[0067] The splitting model is used to parse historical tokens to obtain token classification parameters. The token classification parameters are used to classify each token in the token dictionary. Each token in the token dictionary can be divided into two categories. One category is the watermark tokens for adding watermarks. All the watermark tokens in the token dictionary can form a watermark token list, which has the same function as the existing green list but contains different content. The other category is the non-watermark tokens without added watermarks. All the non-watermark tokens in the token dictionary can form a non-watermark token list, which has the same function as the existing red list but contains different content.

[0068] Since after the text generation unit outputs the token probability distribution each time and combines the historical tokens output previously, the token output each time can be determined. This token will continue to be used as a historical token to determine the token output by the text generation model next time, resulting in a change in the input of the splitting model each time. Therefore, the token classification parameters output by the splitting model will also change. Furthermore, the categories of each token in the token dictionary will also change, that is, the categories of each token in the token dictionary corresponding to different outputs of the text generation model are not the same, and the proportion of watermark tokens and the proportion of non-watermark tokens are also not the same.

[0069] The deviation model is used to parse historical tokens to obtain probability deviation parameters for different token categories. The probability deviation parameters are used to compensate the probability values of different token categories in the first token probability distribution output this time. Since the categories of each token in the token dictionary may be the same or different, the probability deviation parameters corresponding to the tokens of the same category in the token dictionary are the same, and the probability deviation parameters corresponding to tokens of different categories are different, so as to increase the probability difference corresponding to tokens of different categories in the token dictionary. For example, the probability deviation parameters may include a first type of deviation corresponding to watermark tokens and a second type of deviation corresponding to non-watermark tokens, and the first type of deviation and the second type of deviation are different.

[0070] After that, step S13 is executed. The first token probability distribution, the token classification parameters, and the probability deviation parameters are input into the watermark addition module of the watermark addition model. Using this watermark addition module, in combination with the token classification parameters and the probability deviation parameters, the second token probability distribution with added watermarks is determined and output. Here, the watermark addition module can use the token classification parameters to determine the categories of each token in the token dictionary corresponding to the current output, and then in combination with the deviations corresponding to tokens of different categories in the probability deviation parameters, determine the second token probability distribution with added watermarks. Among them, the second token probability distribution is the probability distribution of each token in the token dictionary after adding watermarks.

[0071] It should be noted that the watermark added here is not an explicit and specific content (such as a specific string or identifier), but an implicit statistical pattern. This pattern is reflected by the selection preference of tokens. Specifically, the watermark is encoded through the statistical characteristics of the watermark token list in the token dictionary.

[0072] Finally, step S14 is executed to determine the current token output by the text generation model for the current time using the second token probability distribution. For example, the token corresponding to the optimal probability can be selected from the second token probability distribution as the current token output for the current time.

[0073] In the embodiment of the present invention, the processes of steps S11 - S14 are continuously iteratively executed. The token output by the text generation model each time is used as the historical token for the next output and also as the next input until the current token output by the text generation model for the current time is the end symbol, completing the text generation action of the text generation model, and all the generated tokens form the generated text with the added watermark.

[0074] In the watermark addition method provided in the embodiment of the present invention, first, the historical tokens generated by the text generation model and the first token probability distribution output by the text generation unit in the text generation model for the current time are obtained. The historical tokens are input into the splitting model and the deviation model of the watermark addition model in the text generation model. The token classification parameter is obtained through the splitting model, and the probability deviation parameter is obtained through the deviation model. Then, the first token probability distribution, the token classification parameter, and the probability deviation parameter are input into the watermark addition module of the watermark addition model to determine the second token probability distribution with the added watermark. Finally, the second token probability distribution is used to determine the current token output by the text generation model for the current time. The splitting model is introduced in this method to determine the token classification parameter according to the historical tokens, so that the proportion of tokens of the same category in the token dictionary corresponding to different outputs of the text generation model is different, which can be applicable to the generated text in various situations, avoiding forcibly setting a fixed proportion to damage the accuracy and usability of the generated content of the text generation model, making the watermark addition effect stable, and thus ensuring the subsequent watermark detection effect. Moreover, the deviation model is also introduced in this method to determine the probability deviation parameter of different token categories in the token dictionary according to the historical tokens, so that the watermark addition module can update the first token probability distribution by combining the probability deviation parameter and the token classification parameter, which can change the probability value of tokens of different token categories in the token dictionary being selected, further improving the subsequent watermark detection effect.

[0075] Based on the above embodiments, the watermark addition module is specifically used for:

[0076] Based on the token classification parameter, determine the current category of each token in the token dictionary;

[0077] Update the first token probability distribution based on the probability deviation parameter and the current category of each token to obtain the second token probability distribution.

[0078] Specifically, when the watermark addition module determines the second token probability distribution, it can first use the token classification parameter to determine the current category of each token in the token dictionary. Here, the token classification parameter can include the same number of random variables as the number of tokens in the token dictionary and corresponding one-to-one to each token in the token dictionary. The random variable corresponding to each token in the token dictionary can be greater than 0 and less than or equal to 1. By judging the size of the random variable corresponding to each token in the token dictionary and a specified threshold, each token in the token dictionary is classified. For example, if the random variable corresponding to a certain token is greater than the specified threshold, the current category of this token can be determined as a watermark token, otherwise the current category of this token is determined as a non-watermark token. This specified threshold can be set as needed and is not specifically limited here. By outputting different combinations of random variables each time, the current categories of each token in the token dictionary can be made random.

[0079] After that, the sum of each watermark token and the corresponding first type of deviation can be used as the current probability value of each watermark token, and the sum of each non-watermark token and the corresponding second type of deviation can be used as the current probability value of each non-watermark token. The current probability values of each watermark token and the current probability values of each non-watermark token are arranged in the order of each token in the first token probability distribution, and then the second token probability distribution can be obtained.

[0080] In the embodiment of the present invention, the watermark addition module updates the first token probability distribution by combining the probability deviation parameter and the current category of each token in the token dictionary, which can make the probability values in the second token probability distribution related to the current category of each token, and further improve the watermark addition effect.

[0081] On the basis of the above embodiment, the token classification parameter includes a random seed and distribution parameters of a preset distribution;

[0082] Correspondingly, the watermark addition module is specifically used for:

[0083] Based on the random seed, rearrange the positions of each token in the token dictionary, and sample from the preset distribution to obtain the same number of probability values as the number of tokens in the token dictionary;

[0084] Correspond the rearranged positions of each token in the token dictionary with the probability values one by one, and determine the current category of each token in the token dictionary based on the probability values corresponding to the rearranged positions of each token in the token dictionary.

[0085] Specifically, the token classification parameters in the embodiments of the present invention may include a random seed and distribution parameters of a preset distribution. The preset distribution can be set as needed. For example, it can be a normal distribution, binomial distribution, Poisson distribution, exponential distribution, t-distribution, F-distribution, chi-square distribution, hypergeometric distribution, geometric distribution, etc. Taking the normal distribution as an example, its distribution parameters may include the mean and the standard deviation .

[0086] Furthermore, when determining the current category of each token in the token dictionary, the positions of each token in the token dictionary can be rearranged first by using the random seed, that is, the positions of each token in the token dictionary are shuffled. Synchronously, probability values equal in number to the tokens in the token dictionary can be sampled from the preset distribution. For example, if the token dictionary contains N tokens, then N probability values can be sampled from the preset distribution.

[0087] After that, the rearranged positions of each token in the token dictionary are put into one-to-one correspondence with the sampled probability values, that is, a sampled probability value is reconfigured for the rearranged position of each token in the token dictionary. Furthermore, the current category of each token in the token dictionary can be determined by using the probability values corresponding to the rearranged positions of each token in the token dictionary. For example, for any token in the token dictionary, if the probability value corresponding to the rearranged position of the any token is greater than the probability threshold, the any token can be determined as a watermark token; otherwise, the any token is determined as a non-watermark token.

[0088] In the embodiments of the present invention, by introducing a random seed, the randomness of the probability values corresponding to each token in the token dictionary is improved. By introducing the distribution parameters of the preset distribution, probability values are configured for each token in the token dictionary, providing a theoretical basis for determining the current category of each token in the token dictionary.

[0089] Based on the above embodiments, each token in the token dictionary includes a watermark token and a non-watermark token; correspondingly, the probability deviation parameter includes a gain parameter corresponding to the watermark token and a loss parameter of the non-watermark token.

[0090] Specifically, in the embodiments of the present invention, to further improve the accuracy of classifying each token in the token dictionary, the probability deviation parameter may include a gain parameter corresponding to the watermark token and a loss parameter of the non-watermark token, that is, the first type of deviation corresponding to the watermark token is a gain parameter, which is a positive value, and the second type of deviation corresponding to the non-watermark token is a loss parameter, which is a negative value. In this way, the probability value of the watermark token can be increased, and the probability value of the non-watermark token can be reduced, achieving dynamic adjustment, avoiding the situation where the probability of the non-watermark token is constantly high, and improving the subsequent detection effect.

[0091] For example, for any token in the token dictionary, the probability value of the any token in the second token probability distribution can be expressed as:

[0092] ;

[0093] wherein, is the probability value of the i-th token in the token dictionary in the first token probability distribution, is the probability value of the i-th token in the token dictionary in the second token probability distribution, is the probability deviation parameter. If the i-th token in the token dictionary is a watermark token, then , is the gain parameter; if the i-th token in the token dictionary is a non-watermark token, then , is the loss parameter.

[0094] Based on the above embodiments, the watermark addition method provided in the embodiments of the present invention is applied to the text generation stage of the text generation model and is mainly in the application period of the text generation model. The core process of this watermark addition method is that the first token probability distribution of each token in the token dictionary output this time will pass through the watermark addition model, and the token with the optimal probability after passing through the watermark rule is selected as the current token output this time. After obtaining the current token output this time, the historical tokens will be updated, and the next round of autoregression will be performed. Similarly, the watermark addition model is used to obtain the current token output according to the watermark rule until the autoregression ends.

[0095] As Figure 2 shown, the complete process of this watermark addition method includes:

[0096] Obtain user interaction content;

[0097] Input the user interaction content and the output historical tokens into the text generation unit in the text generation model, and the text generation unit autoregressively obtains the first token probability distribution output this time;

[0098] Input the historical tokens into the split model and the deviation model of the watermark addition model in the text generation model respectively. Obtain the token classification parameter through the split model, and obtain the gain parameter of the watermark token and the loss parameter of the non-watermark token in the token dictionary through the deviation model;

[0099] Input the first token probability distribution, the token classification parameter, the gain parameter of the watermark token, and the loss parameter of the non-watermark token into the watermark addition module of the watermark addition model, and output the second token probability distribution with the watermark added through the watermark addition module;

[0100] Select the token corresponding to the optimal probability from the second token probability distribution as the current token output this time, and use the current token as the historical token to perform the next autoregression;

[0101] Finally, when the current token is the end token, the autoregression ends, and the generated text with the hidden watermark added is obtained as the output of the text generation model.

[0102] Based on the above embodiments, as Figure 3 shown, an embodiment of the present invention also provides a watermark detection method, including:

[0103] S21, determining a detection window of the text to be detected and the full text of the text to be detected that includes the detection window and its previous text, and obtaining each token fragment with gradually increasing length in the full text;

[0104] S22, for any token fragment, based on the text generation unit in the text generation model, determining the third token probability distribution of the last token in the any token fragment, and applying the watermark addition method provided in the above embodiments to determine the fourth token probability distribution of the added watermark and each token category in the token dictionary;

[0105] S23, calculating the perplexity information entropy of the last token based on the fourth token probability distribution, and calculating the proportion weight of the last token in the hypothesis test based on the perplexity information entropy;

[0106] S24, based on the number of watermark tokens in the detection window, each token category in the token dictionary corresponding to each token in the detection window, and the proportion weight, using the score of the hypothesis test as the detection index, to determine whether the detection window contains a watermark.

[0107] Specifically, for the watermark detection method provided in the embodiment of the present invention, the execution subject is a watermark detection device, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., and no specific limitation is made here.

[0108] First, step S21 is executed to determine the detection window of the text to be detected and the full text of the text to be detected that includes the detection window and its previous text, and obtain each token fragment with gradually increasing length in the full text. The text to be detected is the text that needs to determine whether it is generated by the text generation model. The detection window can be any length of text fragment in the text to be detected that needs to determine whether it is generated by the text generation model. For example, the detection window can be the entire text to be detected, or a certain paragraph in the text to be detected or include multiple paragraphs, and no specific limitation is made here.

[0109] After obtaining the text to be detected, preprocessing can be performed on the text to be detected, such as format conversion and token matching, so that the text to be detected after format conversion can meet the input format of the text generation model, and the tokens included therein are all within the token dictionary used by the text generation model, and the tokens not within the token dictionary are deleted.

[0110] The full text in the text to be detected refers to all the text that appears up to the detection window in the text to be detected, including all the content of the detection window and its previous text.

[0111] After that, the full text can be sliced in the way of gradually increasing the length of the token segment by a fixed step size, and each token segment with gradually increasing length in the full text can be obtained. Each token segment obtained by slicing contains a fixed number of tokens more than the previous token segment, so that each token segment contains context information.

[0112] Then step S22 is executed, and the same operation is performed on each token segment obtained by division. That is, for any token segment, the token segment is input into the text generation model, and the text generation model uses other tokens before the last token in the token segment to generate the third token probability distribution of the last token in the token segment, and the watermark addition method provided in the above embodiments is applied to determine the fourth token probability distribution of adding the watermark and each token category in the token dictionary.

[0113] Since the length of each token segment increases by one token, through the text generation model, the third token probability distribution of each token in the full text can be obtained. Furthermore, through the watermark addition method provided in the above embodiments, the fourth token probability distribution corresponding to each token in the full text and each token category in the token dictionary can be determined.

[0114] After that, step S23 is executed, and the perplexity information entropy of the last token, that is, the perplexity entropy, is calculated by using the fourth token probability distribution. The perplexity information entropy of the i-th token in the full text can be calculated by the following formula:

[0115] ;

[0116] where N is the number of tokens in the token dictionary, is the probability value of the j-th token in the token dictionary in the fourth token probability distribution.

[0117] After that, the perplexity information entropy can be utilized to calculate the proportion weight of the last token in the hypothesis test. The type of this hypothesis test can be set as needed. For example, it can be a z-test, a t-test, a chi-square test, etc. The proportion weight can be the ratio of the difference between the perplexity information entropy of the last token and the maximum value of the perplexity information entropy to the maximum value of the perplexity information entropy. The proportion weight w(i) of the i-th token in the full text in the hypothesis test can be calculated by the following formula:

[0118] ;

[0119] wherein, is the maximum value of the perplexity information entropy.

[0120] Finally, step S24 is executed. Since step S22 can determine each token category in the token dictionary corresponding to each token in the full text, and then directly determine the categories of each token in the window to be detected, and summarize the number of watermark tokens therefrom.

[0121] After that, using the number of watermark tokens in the window to be detected, each token category in the token dictionary corresponding to each token in the window to be detected, and the proportion weight, calculate the score of the hypothesis test, and use the score of the hypothesis test as the detection index to determine whether the window to be detected contains a watermark. If the score of the hypothesis test is greater than the preset threshold, it is determined that the window to be detected contains a watermark. Otherwise, it is determined that the window to be detected does not contain a watermark. Wherein, the preset threshold can be set as needed and is not specifically limited here.

[0122] In the watermark detection method provided in the embodiments of the present invention, by calculating the perplexity information entropy, the watermark detection focuses on the text of non-fixed expression types, improving the detection effect.

[0123] Based on the above embodiments, the method for determining whether the window to be detected contains a watermark by using the number of watermark tokens in the window to be detected, each token category in the token dictionary corresponding to each token in the window to be detected, and the proportion weight, with the score of the hypothesis test as the detection index, includes:

[0124] Based on each token category in the token dictionary corresponding to each token in the window to be detected, calculate the proportion of watermark tokens in the token dictionary corresponding to each token in the window to be detected;

[0125] Based on the number of watermark tokens in the window to be detected, the proportion of watermark tokens corresponding to each token in the window to be detected, and the proportion weight, calculate the score of the hypothesis test;

[0126] If the score of the hypothesis test is greater than the preset threshold, it is determined that the window to be detected contains a watermark.

[0127] Specifically, when determining whether a watermark is included in the window to be detected, the proportion of watermark tokens in the token dictionary corresponding to each token in the window to be detected can be calculated first, that is, the ratio of the number of watermark tokens in the token dictionary to the number of tokens in the token dictionary.

[0128] After that, using the number of watermark tokens in the window to be detected, the proportion of watermark tokens corresponding to each token in the window to be detected, and the proportion weight, the score of the hypothesis test is calculated. Taking the hypothesis test as a z-test as an example, the score of the hypothesis test can be calculated by the following formula:

[0129] ;

[0130] where is the number of watermark tokens in the window to be detected, is the proportion of watermark tokens corresponding to the m-th token in the window to be detected, is the proportion weight corresponding to the m-th token in the window to be detected. m1 is the serial number of the starting token of the window to be detected, and m2 is the serial number of the ending token of the window to be detected.

[0131] After that, judge the size relationship between and the preset threshold. If is greater than the preset threshold, it can be determined that the window to be detected contains a watermark. Otherwise, it is determined that the window to be detected does not contain a watermark.

[0132] In the embodiments of the present invention, the score of the hypothesis test is calculated through the proportion of watermark tokens and the proportion weight, and then it is judged whether the window to be detected contains a watermark.

[0133] Based on the above embodiments, the watermark detection method provided in the embodiments of the present invention is applied to the stage of detecting whether a hidden watermark is included in a piece of text to be detected, and mainly in the period when the official detection of the text content determines whether it is generated by a text generation model.

[0134] As Figure 4 shown, the core process of this watermark detection method includes:

[0135] By sequentially inputting each token fragment with an increasing length in the full text of the text to be detected into the text generation unit in the text generation model, the third token probability distribution of each token in the full text output by the text generation unit can be obtained. Furthermore, by applying the watermark addition methods provided in the above embodiments, the token classification parameters are obtained through the splitting model of the watermark addition model in the text generation model, the probability deviation parameters of different token categories in the token dictionary are obtained through the deviation model of the watermark addition model, and the fourth token probability distribution with the watermark added is obtained through the watermark addition module of the watermark addition model;

[0136] Thereafter, based on the fourth token probability distribution, the perplexity information entropy of each token in the full text is calculated, and using this perplexity information entropy, the proportion weight of each token in the full text in the hypothesis test is calculated;

[0137] Using the number of watermark tokens in the window to be detected, each token category and the proportion weight in the token dictionary corresponding to each token in the window to be detected, the score of the hypothesis test is calculated;

[0138] Using the score of the hypothesis test, it is determined whether the window to be detected contains a watermark.

[0139] As Figure 5 and Figure 6 shown, based on the above embodiments, an embodiment of the present invention also provides a method for training a watermark addition model, including:

[0140] S31, Obtain a text sample, input the text sample into the initial text generation model, obtain the sample probability distribution output by the text generation unit in the initial text generation model and the first type of tokens output by the initial text generation model, and based on the sample probability distribution, determine the second type of tokens;

[0141] S32, Based on the watermark detection methods provided in the above embodiments, calculate the first score of the hypothesis test corresponding to the first type of tokens and the second score of the hypothesis test corresponding to the second type of tokens respectively;

[0142] S33, Calculate the semantic loss based on the first type of tokens and the second type of tokens, and calculate the detection loss based on the first score and the second score;

[0143] S34, Based on the semantic loss and the detection loss, train the initial watermark addition model in the initial text generation model to obtain a trained watermark addition model.

[0144] Specifically, for the watermark addition model training method provided in the embodiments of the present invention, the execution subject is a watermark addition model training device, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., and no specific limitation is made here.

[0145] First, step S31 is executed to obtain a text sample. To ensure the addition effect and watermark detection effect of the trained watermark addition model, the text sample can be text content written manually.

[0146] The text sample is input into an initial text generation model. The initial text generation model can include a general text generation unit and an initial watermark addition model to be trained. Through the text generation unit, a sample probability distribution can be output. The sample probability distribution refers to the probability distribution of each token in the token dictionary. The token corresponding to the optimal probability can be selected from the sample probability distribution as the first type of token.

[0147] The sample probability distribution is input into the initial watermark addition model. After being processed by the splitting model, deviation model, and watermark addition module in the initial watermark addition model, the second type of token output by the initial watermark addition model can be obtained.

[0148] Then, step S32 is executed. Using the watermark detection method provided in the above embodiments, the first score of the hypothesis test corresponding to the first type of token is calculated respectively and the second score of the hypothesis test corresponding to the second type of token .

[0149] After that, step S33 is executed. Using the first type of token and the second type of token, the semantic loss is calculated. Here, the first sentence vector in the first type of token and the second sentence vector in the second type of token can be extracted respectively. For example, a pre-trained language model such as bert can be introduced. By inputting the first type of token into the pre-trained language model, the first sentence vector output by the pre-trained language model can be obtained. By inputting the second type of token into the pre-trained language model, the second sentence vector output by the pre-trained language model can be obtained.

[0150] Furthermore, using the first sentence vector and the second sentence vector, the semantic loss is calculated. The semantic loss can be determined by the similarity between the first sentence vector and the second sentence vector. The similarity can be cosine similarity or other similarities. For example, the semantic loss can be calculated by the following formula:

[0151] ;

[0152] where is the semantic loss, is the first type of token, is the second type of token, is a pre-trained language model, is the first sentence vector, is the second sentence vector, is the similarity between the first sentence vector and the second sentence vector.

[0153] This semantic loss can ensure that the original semantics will not be damaged after adding the watermark, giving priority to ensuring the usability of the text generation model.

[0154] Meanwhile, the detection loss can be calculated by using the first score and the second score. The detection loss can be calculated by the following formula:

[0155] ;

[0156] where, is the detection loss, is the first score, is the second score, is the first weight value.

[0157] Finally, by using the semantic loss and the detection loss, weighted summation can be performed to obtain the total loss, that is: ; where, is the total loss, is the second weight value, is the third weight value.

[0158] Using the total loss, the initial watermark addition model in the initial text generation model can be iteratively trained until the iteration reaches the specified number of times or the total loss converges, and then the trained watermark addition model can be obtained. Furthermore, the text generation model can be obtained, and then the watermark-added text content can be generated by the watermark addition method provided in the above embodiments.

[0159] The watermark addition model training method provided in the embodiments of the present invention realizes the adversarial training of watermark addition and watermark detection by means of the watermark detection method provided in the above embodiments, taking into account the semantic similarity of watermark addition and the accuracy of watermark detection.

[0160] Based on the above embodiments, the watermark addition model training method provided in the embodiments of the present invention is applied to the training adversarial stage of the watermark addition process and the watermark detection process, mainly in the R & D period before commercialization.

[0161] As Figure 7 shown, based on the above embodiments, the embodiments of the present invention provide a watermark addition device, including:

[0162] The first acquisition unit 61 is configured to acquire the historical tokens output by the text generation model and the first token probability distribution output by the text generation unit of the text generation model for the current time.

[0163] The watermark addition unit 62 is configured to input the historical tokens into the splitting model and the deviation model of the watermark addition model in the text generation model, and obtain the token classification parameters output by the splitting model and the probability deviation parameters of different token categories in the token dictionary output by the deviation model; the token classification parameters are used to classify each token in the token dictionary.

[0164] The watermark addition unit 62 is further configured to input the first token probability distribution, the token classification parameters, and the probability deviation parameters into the watermark addition module of the watermark addition model, and obtain the second token probability distribution with the watermark added output by the watermark addition module.

[0165] The output token determination unit 63 is configured to determine the current token output by the text generation model for the current time based on the second token probability distribution.

[0166] Based on the above embodiments, in the watermark addition device provided in the embodiments of the present invention, the watermark addition module is specifically configured to:

[0167] Determine the current category of each token in the token dictionary based on the token classification parameters.

[0168] Update the first token probability distribution based on the probability deviation parameters and the current categories of the tokens to obtain the second token probability distribution.

[0169] Based on the above embodiments, in the watermark addition device provided in the embodiments of the present invention, the token classification parameters include a random seed and distribution parameters of a preset distribution.

[0170] Correspondingly, the watermark addition module is specifically configured to:

[0171] Based on the random seed, rearrange the positions of each token in the token dictionary, and sample probability values equal in number to the tokens in the token dictionary from the preset distribution.

[0172] One-to-one correspond the rearranged positions of each token in the token dictionary with the probability values, and determine the current category of each token in the token dictionary based on the probability values corresponding to the rearranged positions of each token in the token dictionary.

[0173] Based on the above embodiments, in the watermark addition device provided in the embodiments of the present invention, each token in the token dictionary includes a watermark token and a non-watermark token.

[0174] Accordingly, the probability deviation parameter includes a gain parameter corresponding to the watermark token and a loss parameter of the non-watermark token.

[0175] Specifically, in the watermark adding device provided in the embodiments of the present invention, the functions of each module correspond one by one to the operation processes of each step in the above method embodiments, and the achieved effects are also the same. For details, please refer to the above embodiments, and the embodiments of the present invention will not be elaborated herein.

[0176] As Figure 8 shown, based on the above embodiments, an embodiment of the present invention provides a watermark detection device, including:

[0177] A second obtaining unit 71, configured to determine a detection window of a text to be detected and the full text including the detection window and its previous text in the text to be detected, and obtain each token segment with an increasing length in the full text;

[0178] A watermark adding unit 72, configured to, for any token segment, based on the text generation unit in the text generation model, determine the third token probability distribution of the last token in the any token segment, and apply the watermark adding method provided in the above embodiments to determine the fourth token probability distribution after adding the watermark and each token category in the token dictionary;

[0179] An occupancy weight calculation unit 73, configured to calculate the perplexity information entropy of the last token based on the fourth token probability distribution, and calculate the occupancy weight of the last token in the hypothesis test based on the perplexity information entropy;

[0180] A judgment unit 74, configured to judge whether the detection window contains a watermark based on the number of watermark tokens in the detection window, each token category in the token dictionary corresponding to each token in the detection window, and the occupancy weight, with the score of the hypothesis test as the detection index.

[0181] Based on the above embodiments, in the watermark detection device provided in the embodiments of the present invention, the judgment unit is specifically configured to:

[0182] Calculate the proportion of watermark tokens in the token dictionary corresponding to each token in the detection window based on each token category in the token dictionary corresponding to each token in the detection window;

[0183] Calculate the score of the hypothesis test based on the number of watermark tokens in the detection window, the proportion of watermark tokens corresponding to each token in the detection window, and the occupancy weight;

[0184] If the score of the hypothesis test is greater than a preset threshold, it is determined that the detection window contains a watermark.

[0185] Specifically, the functions of the modules in the watermark detection device provided in the embodiments of the present invention correspond one-to-one to the operation processes of the steps in the above method embodiments, and the achieved effects are also the same. For details, please refer to the above embodiments, and the embodiments of the present invention will not elaborate herein.

[0186] As Figure 9 shown, based on the above embodiments, the embodiments of the present invention provide a watermark addition model training device, including:

[0187] A third acquisition unit 81, configured to acquire a text sample, input the text sample into an initial text generation model, obtain a sample probability distribution output by a text generation unit in the initial text generation model and a first type of token output by the initial text generation model, and determine a second type of token based on the sample probability distribution;

[0188] A score calculation unit 82, configured to calculate a first score of the hypothesis test corresponding to the first type of token and a second score of the hypothesis test corresponding to the second type of token respectively based on the watermark detection method provided in the above embodiments;

[0189] A loss calculation unit 83, configured to calculate a semantic loss based on the first type of token and the second type of token, and calculate a detection loss based on the first score and the second score;

[0190] A training unit 84, configured to train an initial watermark addition model in the initial text generation model based on the semantic loss and the detection loss to obtain a trained watermark addition model.

[0191] Based on the above embodiments, in the watermark addition model training device provided in the embodiments of the present invention, the loss calculation unit is specifically configured to:

[0192] Extract a first sentence vector from the first type of token and a second sentence vector from the second type of token respectively;

[0193] Calculate the semantic loss based on the first sentence vector and the second sentence vector.

[0194] Specifically, the functions of the modules in the watermark addition model training device provided in the embodiments of the present invention correspond one-to-one to the operation processes of the steps in the above method embodiments, and the achieved effects are also the same. For details, please refer to the above embodiments, and the embodiments of the present invention will not elaborate herein.

[0195] Figure 10 Illustrates a schematic physical structure diagram of an electronic device, as Figure 10As shown in the figure, the electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 complete communication with each other through the communication bus 840. The processor 810 may call the logical instructions in the memory 830 to execute the watermark addition method, the watermark detection method, or the watermark addition model training method provided in the above embodiments.

[0196] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0197] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the watermark addition method, the watermark detection method, or the watermark addition model training method provided in the above embodiments.

[0198] On yet another aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the watermark addition method, the watermark detection method, or the watermark addition model training method provided in the above embodiments. This computer-readable storage medium may be either a non-transitory computer-readable storage medium or a transitory computer-readable storage medium, and no specific limitation is made here.

[0199] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0200] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A watermark adding method, characterized in that, Comprising: Obtaining historical tokens output by a text generation model and a first token probability distribution of the current output of a text generation unit in the text generation model; Inputting the historical tokens into a splitting model and a deviation model of a watermark addition model in the text generation model to obtain a token classification parameter output by the splitting model and a probability deviation parameter of different token categories in a token dictionary output by the deviation model; the token classification parameter is used to classify each token in the token dictionary; Inputting the first token probability distribution, the token classification parameter, and the probability deviation parameter into a watermark addition module of the watermark addition model to obtain a second token probability distribution with watermark added output by the watermark addition module; The token classification parameter includes a random seed and distribution parameters of a preset distribution; Correspondingly, the watermark addition module is specifically configured to: Based on the random seed, rearrange the positions of each token in the token dictionary, and sample probability values from the preset distribution that are the same in number as the tokens in the token dictionary; Put the rearranged positions of each token in the token dictionary in one-to-one correspondence with the probability values, and determine the current category of each token in the token dictionary based on the probability values corresponding to the rearranged positions of each token in the token dictionary; Based on the second token probability distribution, determine the current token output by the text generation model in the current time.

2. The watermark adding method according to claim 1, wherein The watermark addition module is further specifically configured to: Based on the probability deviation parameter and the current category of each token, update the first token probability distribution to obtain the second token probability distribution.

3. The watermark adding method according to any one of claims 1-2, characterized in that Each token in the token dictionary includes a watermark token and a non-watermark token; Correspondingly, the probability deviation parameter includes a gain parameter corresponding to the watermark token and a loss parameter of the non-watermark token.

4. A watermark detection method, characterized in that, Comprising: Determining a detection window of a text to be detected and the full text of the text to be detected including the detection window and its previous text, and obtaining each token segment with gradually increasing length in the full text; For any token segment, based on a text generation unit in a text generation model, determining a third token probability distribution of the last token in the any token segment, and applying the watermark addition method according to any one of claims 1-3 to determine a fourth token probability distribution with watermark added and each token category in the token dictionary; Based on the fourth token probability distribution, calculating the perplexity information entropy of the last token, and based on the perplexity information entropy, calculating the proportion weight of the last token in a hypothesis test; Based on the number of watermark tokens in the detection window, each token category in the token dictionary corresponding to each token in the detection window, and the proportion weight, taking the score of the hypothesis test as a detection index, determining whether the detection window contains a watermark.

5. The watermark detection method according to claim 4, characterized in that, The determining whether the detection window contains a watermark based on the number of watermark tokens in the detection window, each token category in the token dictionary corresponding to each token in the detection window, and the proportion weight, taking the score of the hypothesis test as a detection index, includes: Calculate the proportion of watermark tokens in the token dictionary corresponding to each token in the window to be detected, based on each token category in the token dictionary corresponding to each token in the window to be detected; Calculate the score of the hypothesis test based on the number of watermark tokens in the window to be detected, the proportion of watermark tokens in the token dictionary corresponding to each token in the window to be detected, and the proportion weight; If the score of the hypothesis test is greater than the preset threshold, determine that the window to be detected contains a watermark.

6. A method for training a watermark addition model, characterized in that Comprising: Obtain a text sample, input the text sample into an initial text generation model, obtain the sample probability distribution output by the text generation unit in the initial text generation model and the first type of tokens output by the initial text generation model, and determine the second type of tokens based on the sample probability distribution; Based on the watermark detection method according to any one of claims 4-5, calculate the first score of the hypothesis test corresponding to the first type of tokens and the second score of the hypothesis test corresponding to the second type of tokens respectively; Calculate a semantic loss based on the first type of tokens and the second type of tokens, and calculate a detection loss based on the first score and the second score; Train the initial watermark addition model in the initial text generation model based on the semantic loss and the detection loss to obtain a trained watermark addition model.

7. The watermark addition model training method according to claim 6, wherein The calculating the semantic loss based on the first type of tokens and the second type of tokens comprises: Extract a first sentence vector from the first type of tokens and a second sentence vector from the second type of tokens respectively; Calculate the semantic loss based on the first sentence vector and the second sentence vector.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the watermark addition method according to any one of claims 1-3, or the watermark detection method according to any one of claims 4-5, or the watermark addition model training method according to any one of claims 6-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the watermark addition method according to any one of claims 1-3, or the watermark detection method according to any one of claims 4-5, or the watermark addition model training method according to any one of claims 6-7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the watermark addition method according to any one of claims 1-3, or the watermark detection method according to any one of claims 4-5, or the watermark addition model training method according to any one of claims 6-7.

Citation Information

Patent Citations

  • Big language model watermark detection method and system capable of being publicly verified

    CN119577708A