Watermark adding method, watermark detection method and watermark adding model training method

By introducing split model and deviation model into the watermark addition method, the proportion of different word categories in the word element dictionary is dynamically adjusted, and the problem of unstable watermark addition effect in the prior art is solved, and the accuracy of watermark detection after adding text watermarks in many cases is achieved.

CN119962541AActive Publication Date: 2025-05-09IFLYTEK CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510437524.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-09
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The existing watermark addition method uses a fixed proportional distribution model word list when driving the model, which cannot adapt to text in various situations, resulting in unstable watermark addition effect and affects subsequent watermark detection effects.

Method used

A split model and deviation model are introduced, the word element classification parameters and probability deviation parameters are determined based on historical word elements, the proportion of different word element categories in the word element dictionary is dynamically adjusted, and the probability distribution of the first word element is updated through the watermark addition module to generate a more stable watermark.

Benefits of technology

It realizes text watermark addition suitable for various situations, improves the stability and consistency of watermark addition, and thus improves the accuracy of subsequent watermark detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962541A_ABST
    Figure CN119962541A_ABST
Patent Text Reader

Abstract

The invention provides a watermark adding method, a watermark detection method and a watermark adding model training method, and relates to the technical field of computer vision, a splitting model is introduced for determining lexical element classification parameters according to historical lexical elements, so that the proportions of lexical elements of the same category in lexical element dictionaries corresponding to different outputs of a text generation model are different, and the lexical element classification parameters are obtained. The method can be suitable for text generation under various conditions, the situation that the accuracy and availability of the generated content of the text generation model are damaged by forcibly setting the fixed proportion is avoided, the watermark adding effect is stable, and then the subsequent watermark detection effect is guaranteed. Moreover, a deviation model is introduced to determine probability deviation parameters of different lexical element categories in the lexical element dictionary according to historical lexical elements, so that the watermark adding module updates the first lexical element probability distribution in combination with the probability deviation parameters and lexical element classification parameters, probability values of lexical elements of different lexical element categories in the lexical element dictionary can be changed, and the lexical element classification accuracy is improved. And the subsequent watermark detection effect is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a watermark adding method, a watermark detection method and a watermark adding model training method. Background Art

[0002] With the emergence and widespread application of large language models (LLMs), such as text generation, translation, and summarization, the abuse of LLMs has also escalated, such as the generation of fake news and web content, users may use LLMs to cheat in academic writing and programming assignments, and LLM-generated texts on social media platforms may be used for social engineering and election manipulation activities, and the quality of LLM-generated texts varies. Therefore, it is becoming increasingly important to track the use of LLM-generated texts and determine whether texts are LLM-generated texts.

[0003] In the prior art, in order to track the text generated by LLM, watermarks are usually added to the text generated by LLM. Methods for adding watermarks usually include data-driven, model-driven, and post-processing methods. Data-driven methods usually rely on background insertion to add a small number of watermarked samples to the LLM dataset, allowing the model to learn implicitly and set relevant loss functions. Model-driven methods manipulate the output of the LLM output layer to output the distribution of logical values ​​(logits) or tokens (also known as tokens) during the inference process, divide the tokens into two lists, and mark them as red and green respectively, and encode the watermark information through the statistical characteristics of the vocabulary in the green list. Post-processing refers to the technology of embedding watermarks by processing the text after the LLM outputs the text. This method usually works in the pipeline as an independent module together with the output of the generative model.

[0004] In order to determine whether a text belongs to the text generated by LLM, it is necessary to detect whether there is a watermark in the text. The watermark detection method corresponding to the existing model-driven watermark addition method is a watermark detection method based on interpretable p-values, which identifies watermarks by statistically analyzing the red and green marks in the text, thereby calculating the significance of the p-value.

[0005] However, the existing watermark adding method applies the pre-set proportions of words in the green list and the red list in the model word list when the model is driven. The fixed proportions cannot be suitable for texts in various situations, resulting in unstable watermark adding effects and affecting subsequent watermark detection effects. Summary of the invention

[0006] The present invention provides a watermark adding method, a watermark detecting method and a watermark adding model training method, so as to solve the defects existing in the related technology.

[0007] The present invention provides a watermark adding method, comprising: Obtaining historical word units output by a text generation model and a probability distribution of a first word unit currently output by a text generation unit in the text generation model; Input the historical word-grams into the split model and the deviation model of the watermark adding model in the text generation model, and obtain the word-gram classification parameters output by the split model and the probability deviation parameters of different word-gram categories in the word-gram dictionary output by the deviation model; the word-gram classification parameters are used to classify each word-gram in the word-gram dictionary; Inputting the first word-unit probability distribution, the word-unit classification parameter and the probability deviation parameter into the watermark adding module of the watermark adding model, and obtaining the watermarked second word-unit probability distribution output by the watermark adding module; Based on the second word-unit probability distribution, a current word-unit currently output by the text generation model is determined.

[0008] According to a watermark adding method provided by the present invention, the watermark adding module is specifically used for: Based on the word-gram classification parameters, determining the current category of each word-gram in the word-gram dictionary; Based on the probability deviation parameter and the current category of each word unit, the first word unit probability distribution is updated to obtain the second word unit probability distribution.

[0009] According to a watermark adding method provided by the present invention, the word-element classification parameters include a random seed and a distribution parameter of a preset distribution; Accordingly, the watermark adding module is specifically used for: Based on the random seed, the positions of the word-grams in the word-gram dictionary are rearranged, and probability values ​​equal to the number of word-grams in the word-gram dictionary are obtained by sampling from the preset distribution; The rearranged position of each word in the word-gram dictionary is matched to the probability value one by one, and the current category of each word in the word-gram dictionary is determined based on the probability value corresponding to the rearranged position of each word in the word-gram dictionary.

[0010] According to a watermark adding method provided by the present invention, each word in the word-word dictionary includes a watermark word and a non-watermark word; Correspondingly, the probability deviation parameter includes a gain parameter corresponding to the watermark word element and a loss parameter of the non-watermark word element.

[0011] The present invention also provides a watermark detection method, comprising: Determine a window to be detected of the text to be detected and a full text of the text to be detected that includes the window to be detected and its preceding text, and obtain word-element fragments of gradually increasing length in the full text; For any word-meta segment, based on the text generation unit in the text generation model, determine the third word-meta probability distribution of the last word-meta in the any word-meta segment, and apply the above-mentioned watermark adding method to determine the fourth word-meta probability distribution of the watermarked word-meta and the word-meta categories in the word-meta dictionary; Based on the fourth word-gram probability distribution, calculating the perplexity information entropy of the last word-gram, and based on the perplexity information entropy, calculating the weight of the last word-gram in the hypothesis test; Based on the number of watermark word units in the window to be detected, the word unit categories in the word unit dictionary corresponding to each word unit in the window to be detected and the proportion weight, the score of the hypothesis test is used as the detection index to determine whether the window to be detected contains a watermark.

[0012] According to a watermark detection method provided by the present invention, based on the number of watermark word units in the window to be detected, the word unit categories in the word unit dictionary corresponding to each word unit in the window to be detected and the proportion weight, the score of the hypothesis test is used as the detection index to determine whether the window to be detected contains a watermark, including: Based on each word-gram category in the word-gram dictionary corresponding to each word-gram in the window to be detected, calculating the proportion of watermark word-grams in the word-gram dictionary corresponding to each word-gram in the window to be detected; Calculating the score of the hypothesis test based on the number of watermark word units in the window to be detected, the proportion of the watermark word units corresponding to each word unit in the window to be detected, and the proportion weight; If the score of the hypothesis test is greater than a preset threshold, it is determined that the window to be detected contains a watermark.

[0013] The present invention also provides a watermark adding model training method, comprising: Acquire a text sample, input the text sample into an initial text generation model, obtain a sample probability distribution output by a text generation unit in the initial text generation model and a first type of word unit output by the initial text generation model, and determine a second type of word unit based on the sample probability distribution; Based on the watermark detection method, respectively calculating a first score of the hypothesis test corresponding to the first category of words and a second score of the hypothesis test corresponding to the second category of words; Calculate semantic loss based on the first category word-gram and the second category word-gram, and calculate detection loss based on the first score and the second score; Based on the semantic loss and the detection loss, the initial watermark adding model in the initial text generation model is trained to obtain a trained watermark adding model.

[0014] According to a watermark adding model training method provided by the present invention, the calculating of semantic loss based on the first category word-grams and the second category word-grams includes: Respectively extracting a first sentence vector from the first category of words and a second sentence vector from the second category of words; The semantic loss is calculated based on the first sentence vector and the second sentence vector.

[0015] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the watermark adding method, watermark detection method, or watermark adding model training method as described above is implemented.

[0016] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the watermark adding method, watermark detection method, or watermark adding model training method as described above is implemented.

[0017] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the watermark adding methods, watermark detection methods, or watermark adding model training methods described above.

[0018] The watermark adding method, watermark detection method and watermark adding model training method provided by the present invention introduce a splitting model in the method, which is used to determine the word unit classification parameters according to the historical word units, so that the proportions of the same category of words in the word unit dictionary corresponding to different outputs of the text generation model are different, which can be applied to the generation of texts in various situations, avoiding the destruction of the accuracy and availability of the generated content of the text generation model by rigidly setting a fixed proportion, making the watermark adding effect stable, and thus ensuring the subsequent watermark detection effect. Moreover, the method also introduces a deviation model to determine the probability deviation parameters of different word unit categories in the word unit dictionary according to the historical word units, so that the watermark adding module combines the probability deviation parameters and the word unit classification parameters to update the probability distribution of the first word unit, which can change the probability values ​​of words of different word unit categories in the word unit dictionary being selected, and further improve the subsequent watermark detection effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 This is one of the flow charts of the watermark adding method provided by the present invention.

[0021] Figure 2 This is the second flow chart of the watermark adding method provided by the present invention.

[0022] Figure 3 This is one of the flow charts of the watermark detection method provided by the present invention.

[0023] Figure 4 This is the second flow chart of the watermark detection method provided by the present invention.

[0024] Figure 5 This is one of the flow charts of the watermark adding model training method provided by the present invention.

[0025] Figure 6 This is the second flow chart of the watermark adding model training method provided by the present invention.

[0026] Figure 7 It is a structural schematic diagram of the watermark adding device provided by the present invention.

[0027] Figure 8 It is a structural schematic diagram of the watermark detection device provided by the present invention.

[0028] Fig. 9 It is a structural schematic diagram of the watermark adding model training device provided by the present invention.

[0029] Fig.10 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0031] Due to the existing watermark adding method, a fixed proportion is used to divide the green list and red list in the model word list when the model is driven. The division is rigid and fails to set exclusive red and green partitions for the current output based on the output historical word, which is not suitable for texts in various situations. Moreover, the model drive only rewards the green zone, but not the red zone, resulting in the original high-probability red zone elements still having high probability, resulting in unstable watermark adding effect and affecting the subsequent watermark detection effect.

[0032] Based on this, a watermark adding method is provided in an embodiment of the present invention.

[0033] Figure 1 FIG. 1 is a flow chart of a watermark adding method provided in an embodiment of the present invention. Figure 1 As shown, the method includes: S11, obtaining the historical word-units output by the text generation model and the probability distribution of the first word-unit currently output by the text generation unit in the text generation model; S12, inputting the historical word-grams into the splitting model and the deviation model of the watermark adding model in the text generation model, obtaining word-gram classification parameters output by the splitting model and probability deviation parameters of different word-gram categories in the word-gram dictionary output by the deviation model; the word-gram classification parameters are used to classify each word-gram in the word-gram dictionary; S13, inputting the first word-unit probability distribution, the word-unit classification parameter and the probability deviation parameter into the watermark adding module of the watermark adding model, and obtaining the watermarked second word-unit probability distribution output by the watermark adding module; S14: Determine a current word-unit currently output by the text generation model based on the second word-unit probability distribution.

[0034] Specifically, the watermark adding method provided in the embodiment of the present invention is executed by a watermark adding device, which can be configured in a computer, which can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., which is not specifically limited here.

[0035] First, step S11 is performed to obtain the historical word-grams output by the text generation model and the probability distribution of the first word-gram currently output by the text generation unit in the text generation model. The adopted text generation model may include a text generation unit and a watermark adding model connected in sequence, and the text generation unit may be an LLM or other models or tools with text generation function, which are not specifically limited here.

[0036] It can be understood that the text generation unit takes the user interaction content and the output historical word-grams as input, generates text content by autoregressive decoding, and in each output, contains the probability distribution of each word-gram in the word-gram dictionary of that sub-time. Therefore, the probability distribution of the first word-gram output at the current time refers to the probability distribution of each word-gram in the word-gram dictionary of the current sub-time. The word-gram dictionary can be the word-gram set of the text generation unit, and is a reference database used by the text generation unit to compare and output the word-gram with the optimal probability.

[0037] It can be understood that after obtaining the probability distribution of the first word unit currently output by the text generation unit, the word unit corresponding to the optimal probability in the first word unit probability distribution is not directly used as the current word unit currently output, but it is necessary to use the watermark adding model to obtain the current word unit currently output by the text generation model with the help of the operations of steps S12-S14.

[0038] The historical word-grams output by the text generation model may include one or more, each of which is a word-gram probability distribution of the word-gram output by the text generation unit at the time, combined with the word-gram output before the current time, and the word-gram obtained by performing the operations of steps S12-S14. In particular, the number of historical word-grams may also be 0, in which case the current output is the first output of the text generation model. Since there are no historical word-grams in the first output, steps S12-S14 cannot be continued. In this case, the optimal probability can be directly selected from the word-gram probability distribution of the first output, and the word-gram corresponding to the optimal probability is used as the word-gram output for the first time.

[0039] Then, step S12 is executed, where the watermark adding model in the text generation model includes a split model (SplitModel), a bias model (BiasModel) and a watermark adding module.

[0040] The split model is used to parse the historical word-grams and obtain word-gram classification parameters. The word-gram classification parameters are used to classify each word-gram in the word-gram dictionary. Each word-gram in the word-gram dictionary can be divided into two categories: one is a watermark word-gram used as a watermark, and all watermark words in the word-gram dictionary can constitute a watermark word-gram list, which has the same function as the existing green list, but contains different contents; the other is a non-watermark word-gram used as a non-watermark, and all non-watermark words in the word-gram dictionary can constitute a non-watermark word-gram list, which has the same function as the existing red list, but contains different contents.

[0041] Each time the text generation unit outputs a word-gram probability distribution, it can determine the word-gram output each time by combining the historical word-grams output previously. This word-gram will continue to be used as a historical word-gram to determine the word-gram output by the text generation model next time, causing the input of the splitting model to change each time. Therefore, the word-gram classification parameters output by the splitting model each time will also change, and then the categories of each word-gram in the word-gram dictionary will also change, that is, the categories of each word-gram in the word-gram dictionary corresponding to different outputs of the text generation model are not the same, and the proportion of watermark word-grams and non-watermark word-grams are also different.

[0042] The deviation model is used to parse historical word units to obtain probability deviation parameters of different word unit categories. The probability deviation parameter is used to compensate for the probability values ​​of different word unit categories in the probability distribution of the first word unit currently output. Since the categories of each word unit in the word unit dictionary may be the same or different, the probability deviation parameters corresponding to the words of the same category in the word unit dictionary are the same, and the probability deviation parameters corresponding to the words of different categories are different, so as to increase the probability difference corresponding to the words of different categories in the word unit dictionary. For example, the probability deviation parameter may include a first type of deviation corresponding to a watermark word unit and a second type of deviation corresponding to a non-watermark word unit, and the first type of deviation and the second type of deviation are different.

[0043] Thereafter, step S13 is executed to input the first word-gram probability distribution, word-gram classification parameter, and probability deviation parameter into the watermark adding module of the watermark adding model, and the watermark adding module is used to determine and output the second word-gram probability distribution of the watermarked word in combination with the word-gram classification parameter and the probability deviation parameter. Here, the watermark adding module can use the word-gram classification parameter to determine the category of each word-gram in the word-gram dictionary corresponding to the current output, and then determine the second word-gram probability distribution of the watermarked word in combination with the deviation corresponding to the word-grams of different categories in the probability deviation parameter. Among them, the second word-gram probability distribution is the probability distribution of each word-gram in the word-gram dictionary after the watermark is added.

[0044] It should be noted that the watermark added here is not an explicit, specific content (such as a specific string or identifier), but an implicit statistical pattern. This pattern is reflected by the preference for word units. Specifically, the watermark is encoded through the statistical characteristics of the watermark word unit list in the word unit dictionary.

[0045] Finally, step S14 is performed to determine the current word unit output by the text generation model using the second word unit probability distribution. For example, the word unit corresponding to the optimal probability can be selected from the second word unit probability distribution as the current word unit output.

[0046] In the embodiment of the present invention, the process of steps S11-S14 is a process of continuous iterative execution, and each word element output by the text generation model serves as the historical word element for the next output, and also serves as the next input, until the current word element currently output by the text generation model is the terminator, completing the text generation action of the text generation model, and all the generated word elements constitute the generated text with the watermark added.

[0047] The watermark adding method provided in the embodiment of the present invention first obtains the historical word-grams generated by the text generation model and the probability distribution of the first word-gram currently output by the text generation unit in the text generation model, inputs the historical word-grams into the split model and the deviation model of the watermark adding model in the text generation model, obtains the word-gram classification parameters through the split model, obtains the probability deviation parameters through the deviation model, and then inputs the first word-gram probability distribution, the word-gram classification parameters and the probability deviation parameters into the watermark adding module of the watermark adding model to determine the probability distribution of the second word-gram to be watermarked, and finally determines the current word-gram currently output by the text generation model using the second word-gram probability distribution. The method introduces a split model for determining the word-gram classification parameters based on the historical word-grams, so that the proportions of the same category of words in the word-gram dictionary corresponding to different outputs of the text generation model are different, and can be applied to the generation of text in a variety of situations, avoiding the rigid setting of fixed proportions that undermines the accuracy and availability of the generated content of the text generation model, making the watermark adding effect stable, thereby ensuring the subsequent watermark detection effect. Moreover, the method also introduces a deviation model to determine the probability deviation parameters of different word-unit categories in the word-unit dictionary based on historical word-units, so that the watermark adding module updates the probability distribution of the first word-unit in combination with the probability deviation parameters and the word-unit classification parameters, which can change the probability values ​​of words of different word-unit categories in the word-unit dictionary being selected, and further improve the subsequent watermark detection effect.

[0048] Based on the above embodiment, the watermark adding module is specifically used for: Based on the word-gram classification parameters, determining the current category of each word-gram in the word-gram dictionary; Based on the probability deviation parameter and the current category of each word unit, the first word unit probability distribution is updated to obtain the second word unit probability distribution.

[0049] Specifically, when the watermark adding module determines the probability distribution of the second word-gram, the word-gram classification parameters can be used to first determine the current category of each word-gram in the word-gram dictionary. Here, the word-gram classification parameters may include random variables that are the same number as the number of words in the word-gram dictionary and correspond one-to-one to each word-gram in the word-gram dictionary. The random variables corresponding to each word-gram in the word-gram dictionary can be greater than 0 and less than or equal to 1. By judging the size of the random variables corresponding to each word-gram in the word-gram dictionary and the specified threshold, each word-gram in the word-gram dictionary is classified. For example, if the random variable corresponding to a certain word-gram is greater than the specified threshold, the current category of the word-gram can be determined as a watermark word-gram, otherwise the current category of the word-gram is determined to be a non-watermark word-gram. The specified threshold can be set as needed and is not specifically limited here. By outputting different combinations of random variables each time, the current category of each word-gram in the word-gram dictionary can be made random.

[0050] Afterwards, the sum of each watermark word and the corresponding first-type deviation can be used as the current probability value of each watermark word, and the sum of each non-watermark word and the corresponding second-type deviation can be used as the current probability value of each non-watermark word. The current probability values ​​of each watermark word and the current probability values ​​of each non-watermark word are arranged according to the order of each word in the first word probability distribution to obtain the second word probability distribution.

[0051] In the embodiment of the present invention, the watermark adding module updates the first word unit probability distribution in combination with the probability deviation parameter and the current category of each word unit in the word unit dictionary, so that each probability value in the second word unit probability distribution is related to the current category of each word unit, further improving the watermark adding effect.

[0052] Based on the above embodiment, the word-unit classification parameters include a random seed and a distribution parameter of a preset distribution; Accordingly, the watermark adding module is specifically used for: Based on the random seed, the positions of the word-grams in the word-gram dictionary are rearranged, and probability values ​​equal to the number of word-grams in the word-gram dictionary are obtained by sampling from the preset distribution; The rearranged position of each word in the word-gram dictionary is matched to the probability value one by one, and the current category of each word in the word-gram dictionary is determined based on the probability value corresponding to the rearranged position of each word in the word-gram dictionary.

[0053] Specifically, the word-unit classification parameters in the embodiment of the present invention may include a random seed and a distribution parameter of a preset distribution. The preset distribution may be set as needed, for example, a normal distribution, a binomial distribution, a Poisson distribution, an exponential distribution, a t distribution, an F distribution, a chi-square distribution, a hypergeometric distribution, and a geometric distribution. Taking the normal distribution as an example, its distribution parameters may include a mean, a and standard deviation .

[0054] Furthermore, when determining the current category of each word in the word-gram dictionary, the position of each word in the word-gram dictionary can be rearranged by using a random seed, that is, the position of each word in the word-gram dictionary can be shuffled. Simultaneously, probability values ​​equal to the number of word-grams in the word-gram dictionary can be sampled from a preset distribution. For example, if the word-gram dictionary contains N word-grams, N probability values ​​can be sampled from the preset distribution.

[0055] Thereafter, the rearranged position of each word in the word-gram dictionary is matched to the sampled probability value one by one, that is, a sampled probability value is reconfigured for the rearranged position of each word in the word-gram dictionary. Furthermore, the probability value corresponding to the rearranged position of each word in the word-gram dictionary can be used to determine the current category of each word in the word-gram dictionary. For example, for any word in the word-gram dictionary, if the probability value corresponding to the rearranged position of any word is greater than the probability threshold, then any word can be determined as a watermark word, otherwise any word is determined as a non-watermark word.

[0056] In the embodiment of the present invention, by introducing a random seed, the randomness of the probability value corresponding to each word in the word-gram dictionary is improved, and by introducing a distribution parameter of a preset distribution, a probability value is configured for each word in the word-gram dictionary, providing a theoretical basis for determining the current category of each word in the word-gram dictionary.

[0057] Based on the above embodiment, each word in the word-word dictionary includes a watermark word and a non-watermark word; accordingly, the probability deviation parameter includes a gain parameter corresponding to the watermark word and a loss parameter of the non-watermark word.

[0058] Specifically, in the embodiment of the present invention, in order to further improve the accuracy of classification of each word in the word word dictionary, the probability deviation parameter may include a gain parameter corresponding to the watermark word and a loss parameter of the non-watermark word, that is, the first type of deviation corresponding to the watermark word is a gain parameter, which is a positive value, and the second type of deviation corresponding to the non-watermark word is a loss parameter, which is a negative value. In this way, the probability value of the watermark word can be increased, and the probability value of the non-watermark word can be reduced, so as to achieve dynamic adjustment, avoid the situation where the non-watermark word has a high probability and is always high, and improve the subsequent detection effect.

[0059] For example, for any word in the word-gram dictionary, the probability value of the word in the second word-gram probability distribution can be expressed as: ; in, is the probability value of the i-th word in the word dictionary in the first word probability distribution, is the probability value of the i-th word in the word-gram dictionary in the probability distribution of the second word-gram, is the probability deviation parameter. If the i-th word in the word dictionary is a watermark word, then , is the gain parameter; if the i-th word in the word dictionary is a non-watermark word, then , is the impairment parameter.

[0060] On the basis of the above-mentioned embodiments, the watermark adding method provided in the embodiments of the present invention is applied to the text generation stage of the text generation model, and is mainly in the application period of the text generation model. The core process of the watermark adding method is that the probability distribution of the first word element of each word element in the word element dictionary of the current output will be selected through the watermark adding model, and the word element with the best probability after the watermark rule will be selected as the current word element of the current output. After obtaining the current word element of the current output, the historical word element will be updated, and the next round of autoregression will be performed. The watermark adding model is also used to obtain the current word element output according to the watermark rule until the autoregression ends.

[0061] like Figure 2 As shown, the complete process of the watermark adding method includes: Get user interaction content; The user interaction content and the output historical word-units are input into the text generation unit in the text generation model, and the text generation unit is used to autoregressively obtain the probability distribution of the first word-unit of the current output; The historical words are input into the split model and the deviation model of the watermark adding model in the text generation model respectively, the word classification parameters are obtained through the split model, and the gain parameters of the watermark words and the loss parameters of the non-watermark words in the word dictionary are obtained through the deviation model; Inputting the first word unit probability distribution, word unit classification parameters, watermark word unit gain parameters and non-watermark word unit loss parameters into the watermark adding module of the watermark adding model, and outputting the watermarked second word unit probability distribution through the watermark adding module; Select the word corresponding to the optimal probability from the second word probability distribution as the current word of the current output, and use the current word as the historical word to perform the next autoregression; Finally, when the current word is the end symbol, the autoregression ends, and the generated text with a hidden watermark is obtained as the output of the text generation model.

[0062] Based on the above embodiments, Figure 3 As shown, an embodiment of the present invention further provides a watermark detection method, including: S21, determining a window to be detected of the text to be detected and a full text of the text to be detected that includes the window to be detected and its preceding text, and obtaining word-element fragments of gradually increasing length in the full text; S22, for any word-meta segment, based on the text generation unit in the text generation model, determine the third word-meta probability distribution of the last word-meta in the any word-meta segment, and apply the watermark adding method provided in the above embodiments to determine the fourth word-meta probability distribution of the watermarked word-meta and each word-meta category in the word-meta dictionary; S23, calculating the perplexity information entropy of the last word based on the fourth word probability distribution, and calculating the weight of the last word in the hypothesis test based on the perplexity information entropy; S24, based on the number of watermark word units in the window to be detected, the word unit categories in the word unit dictionary corresponding to each word unit in the window to be detected and the proportion weight, taking the score of the hypothesis test as the detection index, determining whether the window to be detected contains a watermark.

[0063] Specifically, the watermark detection method provided in the embodiment of the present invention is executed by a watermark detection device, which can be configured in a computer. The computer can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., which is not specifically limited here.

[0064] First, step S21 is performed to determine the window to be detected of the text to be detected and the full text of the text to be detected that contains the window to be detected and its preceding text, and obtain each word-element fragment with gradually increasing length in the full text. The text to be detected is a text that needs to be judged whether it is generated by a text generation model. The window to be detected can be a text fragment of any length in the text to be detected that needs to be judged whether it is generated by a text generation model. For example, the detection window can be the entire text to be detected, or it can be a paragraph in the text to be detected or contain multiple paragraphs, which is not specifically limited here.

[0065] After obtaining the text to be detected, the text to be detected can be preprocessed, such as performing format conversion and word unit matching, so that the text to be detected after format conversion can meet the input format of the text generation model, and the word units contained therein are all within the word unit dictionary used by the text generation model, and the word units that are not within the word unit dictionary are deleted.

[0066] The full text in the to-be-detected text refers to all the text that appears in the to-be-detected text up to the to-be-detected window, including all the contents of the to-be-detected window and its preceding text.

[0067] Afterwards, the full text can be sliced ​​in a manner of gradually increasing word-meta fragment lengths according to a fixed step size, so that word-meta fragments with gradually increasing lengths in the full text can be obtained. Each word-meta fragment obtained by slicing is based on the previous word-meta fragment and adds a fixed step size of words, so that each word-meta fragment contains contextual information.

[0068] Then, step S22 is executed, and the same operation is performed on each word-meta segment obtained by the division, that is, for any word-meta segment, the any word-meta segment is input into the text generation model, and the text generation model uses other word-meta before the last word-meta in any word-meta segment to generate the third word-meta probability distribution of the last word-meta in any word-meta segment, and applies the watermark adding method provided in the above embodiments to determine the fourth word-meta probability distribution to add the watermark and the word-meta categories in the word-meta dictionary.

[0069] Since the length of each word-gram fragment increases word-by-word, the third-word probability distribution of each word-gram in the full text can be obtained through the text generation model, and then the fourth-word probability distribution corresponding to each word-gram in the full text and the word-gram categories in the word-gram dictionary can be determined through the watermark adding method provided in the above embodiments.

[0070] Then, step S23 is executed to calculate the perplexity information entropy of the last word using the probability distribution of the fourth word. The perplexity information entropy of the i-th word in the full text is It can be calculated by the following formula: ; Where N is the number of words in the word dictionary, is the probability value of the j-th word in the word-word dictionary in the fourth word-word probability distribution.

[0071] After that, the perplexity information entropy can be used to calculate the weight of the last word in the hypothesis test. The type of the hypothesis test can be set as needed, for example, it can be a z test, a t test, a chi-square test, etc. The weight can be the ratio of the difference between the perplexity information entropy of the last word and the maximum perplexity information entropy to the maximum perplexity information entropy. The weight w(i) of the i-th word in the hypothesis test in the full text can be calculated by the following formula: ; in, is the maximum value of perplexity information entropy.

[0072] Finally, step S24 is executed. Since step S22 can determine the word-unit categories in the word-unit dictionary corresponding to each word-unit in the full text, the category of each word-unit in the window to be detected can be directly determined, and the number of watermark word-units can be obtained from them.

[0073] After that, the number of watermark word units in the window to be detected, the word unit categories and proportion weights in the word unit dictionary corresponding to each word unit in the window to be detected are used to calculate the score of the hypothesis test, and the score of the hypothesis test is used as the detection index to determine whether the window to be detected contains a watermark. If the score of the hypothesis test is greater than the preset threshold, it is determined that the window to be detected contains a watermark. Otherwise, it is determined that the window to be detected does not contain a watermark. Among them, the preset threshold can be set as needed and is not specifically limited here.

[0074] The watermark detection method provided in the embodiment of the present invention makes watermark detection focus on texts with non-fixed expressions by calculating perplexity information entropy, thereby improving the detection effect.

[0075] On the basis of the above embodiment, the method of judging whether the window to be detected contains a watermark based on the number of watermark word units in the window to be detected, the word unit categories in the word unit dictionary corresponding to each word unit in the window to be detected, and the proportion weight, and taking the score of the hypothesis test as the detection index, comprises: Based on each word-gram category in the word-gram dictionary corresponding to each word-gram in the window to be detected, calculating the proportion of watermark word-grams in the word-gram dictionary corresponding to each word-gram in the window to be detected; Calculating the score of the hypothesis test based on the number of watermark word units in the window to be detected, the proportion of the watermark word units corresponding to each word unit in the window to be detected, and the proportion weight; If the score of the hypothesis test is greater than a preset threshold, it is determined that the window to be detected contains a watermark.

[0076] Specifically, when determining whether a watermark is contained in the window to be detected, the word-gram categories in the word-gram dictionary corresponding to each word-gram in the window to be detected can be used to calculate the proportion of watermark word-grams in the word-gram dictionary corresponding to each word-gram in the window to be detected, that is, the ratio of the number of watermark word-grams in the word-gram dictionary to the number of word-grams in the word-gram dictionary.

[0077] After that, the score of the hypothesis test is calculated using the number of watermark words in the window to be detected, the proportion of watermark words corresponding to each word in the window to be detected, and the proportion weight. Taking the hypothesis test as an example, the score of the hypothesis test is It can be calculated by the following formula: ; in, is the number of watermark words in the window to be detected, is the proportion of watermark word corresponding to the mth word in the window to be detected, is the weight corresponding to the mth word in the window to be detected. m1 is the sequence number of the starting word in the window to be detected, and m2 is the sequence number of the ending word in the window to be detected.

[0078] Afterwards, judge The relationship between the size of and the preset threshold value is If the value is greater than a preset threshold, it can be determined that the window to be detected contains a watermark. Otherwise, it is determined that the window to be detected does not contain a watermark.

[0079] In the embodiment of the present invention, the score of the hypothesis test is calculated by the watermark word unit ratio and the ratio weight, so as to determine whether the window to be detected contains a watermark.

[0080] On the basis of the above-mentioned embodiments, the watermark detection method provided in the embodiments of the present invention is applied to the stage of detecting whether a section of text to be detected contains a hidden watermark, mainly in the stage of officially detecting whether the text content belongs to the generation period of the text generation model.

[0081] like Figure 4 As shown, the core process of the watermark detection method includes: By sequentially inputting each word-meta fragment with gradually increasing length in the full text of the text to be detected into the text generation unit in the text generation model, the third word-meta probability distribution of each word-meta in the full text output by the text generation unit can be obtained, and then the watermark adding method provided in the above embodiments is applied, the word-meta classification parameters are obtained through the splitting model of the watermark adding model in the text generation model, the probability deviation parameters of different word-meta categories in the word-meta dictionary are obtained through the deviation model of the watermark adding model, and the fourth word-meta probability distribution of the watermark is obtained through the watermark adding module of the watermark adding model; After that, the perplexity information entropy of each word in the full text is calculated through the probability distribution of the fourth word, and the perplexity information entropy is used to calculate the weight of each word in the full text in the hypothesis test; The score of the hypothesis test is calculated using the number of watermark word units in the window to be detected, the word unit categories and proportion weights in the word unit dictionary corresponding to each word unit in the window to be detected; The score of the hypothesis test is used to determine whether the window to be detected contains a watermark.

[0082] like Figure 5 and Figure 6 As shown, based on the above embodiment, an embodiment of the present invention further provides a watermark adding model training method, including: S31, obtaining a text sample, inputting the text sample into an initial text generation model, obtaining a sample probability distribution output by a text generation unit in the initial text generation model and a first type of word unit output by the initial text generation model, and determining a second type of word unit based on the sample probability distribution; S32, based on the watermark detection method provided in the above embodiments, respectively calculating a first score of the hypothesis test corresponding to the first category of words and a second score of the hypothesis test corresponding to the second category of words; S33, calculating a semantic loss based on the first category word-grams and the second category word-grams, and calculating a detection loss based on the first score and the second score; S34, based on the semantic loss and the detection loss, the initial watermark adding model in the initial text generation model is trained to obtain a trained watermark adding model.

[0083] Specifically, the watermark adding model training method provided in the embodiment of the present invention is executed by a watermark adding model training device, which can be configured in a computer, which can be a local computer or a cloud computer. The local computer can be a computer, a tablet, etc., which is not specifically limited here.

[0084] First, step S31 is executed to obtain a text sample. To ensure the watermark adding effect and watermark detection effect of the trained watermark adding model, the text sample may be a manually written text content.

[0085] The text sample is input into the initial text generation model, which may include a general text generation unit and an initial watermark adding model to be trained. The sample probability distribution may be output through the text generation unit, which refers to the probability distribution of each word in the word word dictionary. From the sample probability distribution, the word corresponding to the optimal probability may be selected as the first type of word.

[0086] The sample probability distribution is input into the initial watermark adding model, and can be processed by the splitting model, the deviation model and the watermark adding module in the initial watermark adding model to obtain the second type of word element output by the initial watermark adding model.

[0087] Then, step S32 is performed to calculate the first score of the hypothesis test corresponding to the first type of word by using the watermark detection method provided in the above embodiments. The second score of the hypothesis test corresponding to the second type of word .

[0088] Thereafter, step S33 is executed to calculate the semantic loss using the first type of word-gram and the second type of word-gram. Here, the first sentence vector in the first type of word-gram and the second sentence vector in the second type of word-gram can be extracted respectively. For example, a pre-trained language model such as Bert can be introduced, and the first sentence vector output by the pre-trained language model can be obtained by inputting the first type of word-gram into the pre-trained language model, and the second sentence vector output by the pre-trained language model can be obtained by inputting the second type of word-gram into the pre-trained language model.

[0089] Then, the semantic loss is calculated using the first sentence vector and the second sentence vector. The semantic loss can be determined by the similarity between the first sentence vector and the second sentence vector, and the similarity can be cosine similarity or other similarity. For example, the semantic loss can be calculated by the following formula: ; in, is the semantic loss, is the first type of word, For the second type of word, To pre-train the language model, is the first sentence vector, is the second sentence vector, is the similarity between the first sentence vector and the second sentence vector.

[0090] This semantic loss can ensure that the original semantics will not be destroyed after adding the watermark, so as to prioritize the practicability of the text generation model.

[0091] At the same time, the first score and the second score can also be used to calculate the detection loss. The detection loss can be calculated by the following formula: ; in, To detect the loss, For the first score, For the second score, is the first weight value.

[0092] Finally, using the semantic loss and detection loss, we can perform a weighted sum to get the total loss, that is: ;in, is the total loss, is the second weight value, is the third weight value.

[0093] By using the total loss, the initial watermark adding model in the initial text generation model can be iteratively trained until the iteration reaches a specified number of times or the total loss converges, so as to obtain the trained watermark adding model, and then obtain the text generation model, and then generate the text content with watermark added through the watermark adding method provided in the above embodiments.

[0094] The watermark adding model training method provided in the embodiment of the present invention, with the help of the watermark detection method provided in the above embodiments, realizes the adversarial training of watermark adding and watermark detection, taking into account the semantic similarity of watermark adding and the accuracy of watermark detection.

[0095] On the basis of the above embodiments, the watermark adding model training method provided in the embodiments of the present invention is applied to the training confrontation stage of the watermark adding process and the watermark detecting process, and is mainly in the research and development period before commercial use.

[0096] like Figure 7 As shown, based on the above embodiment, an embodiment of the present invention provides a watermark adding device, including: A first acquisition unit 61 is used to acquire the historical word-grams output by the text generation model and the probability distribution of the first word-gram currently output by the text generation unit in the text generation model; The watermark adding unit 62 is used to input the historical word-grams into the splitting model and the deviation model of the watermark adding model in the text generation model, and obtain the word-gram classification parameters output by the splitting model and the probability deviation parameters of different word-gram categories in the word-gram dictionary output by the deviation model; the word-gram classification parameters are used to classify each word-gram in the word-gram dictionary; The watermark adding unit 62 is further used to input the first word-unit probability distribution, the word-unit classification parameter and the probability deviation parameter into the watermark adding module of the watermark adding model, and obtain the watermarked second word-unit probability distribution output by the watermark adding module; The output word unit determining unit 63 is used to determine the current word unit currently output by the text generation model based on the second word unit probability distribution.

[0097] On the basis of the above embodiments, in the watermark adding device provided in the embodiments of the present invention, the watermark adding module is specifically used for: Based on the word-gram classification parameters, determining the current category of each word-gram in the word-gram dictionary; Based on the probability deviation parameter and the current category of each word unit, the first word unit probability distribution is updated to obtain the second word unit probability distribution.

[0098] On the basis of the above-mentioned embodiment, in the watermark adding device provided in the embodiment of the present invention, the word-unit classification parameter includes a random seed and a distribution parameter of a preset distribution; Accordingly, the watermark adding module is specifically used for: Based on the random seed, the positions of the word-grams in the word-gram dictionary are rearranged, and probability values ​​equal to the number of word-grams in the word-gram dictionary are obtained by sampling from the preset distribution; The rearranged position of each word in the word-gram dictionary is matched to the probability value one by one, and the current category of each word in the word-gram dictionary is determined based on the probability value corresponding to the rearranged position of each word in the word-gram dictionary.

[0099] On the basis of the above-mentioned embodiment, in the watermark adding device provided in the embodiment of the present invention, each word in the word-word dictionary includes a watermark word and a non-watermark word; Correspondingly, the probability deviation parameter includes a gain parameter corresponding to the watermark word element and a loss parameter of the non-watermark word element.

[0100] Specifically, the functions of each module in the watermark adding device provided in the embodiment of the present invention correspond one-to-one to the operation flow of each step in the above method embodiment, and the effects achieved are also consistent. Please refer to the above embodiment for details, which will not be repeated in the embodiment of the present invention.

[0101] like Figure 8 As shown, based on the above embodiment, an embodiment of the present invention provides a watermark detection device, including: The second acquisition unit 71 is used to determine a window to be detected of the text to be detected and a full text of the text to be detected that includes the window to be detected and its preceding text, and to acquire word-element fragments of gradually increasing length in the full text; The watermark adding unit 72 is used to determine, for any word-meta segment, the third word-meta probability distribution of the last word-meta in the any word-meta segment based on the text generation unit in the text generation model, and apply the watermark adding method provided by the above embodiments to determine the fourth word-meta probability distribution of the watermark and each word-meta category in the word-meta dictionary; A weight proportion calculation unit 73, configured to calculate the perplexity information entropy of the last word based on the fourth word probability distribution, and calculate the weight proportion of the last word in the hypothesis test based on the perplexity information entropy; The judging unit 74 is used to judge whether the window to be detected contains a watermark based on the number of watermark word units in the window to be detected, the word unit categories in the word unit dictionary corresponding to each word unit in the window to be detected and the proportion weight, and the score of the hypothesis test as a detection index.

[0102] On the basis of the above embodiments, in the watermark detection device provided in the embodiments of the present invention, the judgment unit is specifically used for: Based on each word-gram category in the word-gram dictionary corresponding to each word-gram in the window to be detected, calculating the proportion of watermark word-grams in the word-gram dictionary corresponding to each word-gram in the window to be detected; Calculating the score of the hypothesis test based on the number of watermark word units in the window to be detected, the proportion of the watermark word units corresponding to each word unit in the window to be detected, and the proportion weight; If the score of the hypothesis test is greater than a preset threshold, it is determined that the window to be detected contains a watermark.

[0103] Specifically, the functions of each module in the watermark detection device provided in the embodiment of the present invention correspond one-to-one to the operation flow of each step in the above method embodiment, and the effects achieved are also consistent. Please refer to the above embodiment for details, which will not be repeated in the embodiment of the present invention.

[0104] like Fig. 9 As shown, based on the above embodiment, an embodiment of the present invention provides a watermark adding model training device, including: A third acquisition unit 81 is used to acquire a text sample, input the text sample into the initial text generation model, obtain the sample probability distribution output by the text generation unit in the initial text generation model and the first type of word unit output by the initial text generation model, and determine the second type of word unit based on the sample probability distribution; A score calculation unit 82, configured to calculate a first score of the hypothesis test corresponding to the first category of words and a second score of the hypothesis test corresponding to the second category of words, respectively, based on the watermark detection method provided in the above embodiments; a loss calculation unit 83, configured to calculate a semantic loss based on the first category word-grams and the second category word-grams, and to calculate a detection loss based on the first score and the second score; The training unit 84 is used to train the initial watermark adding model in the initial text generation model based on the semantic loss and the detection loss to obtain a trained watermark adding model.

[0105] On the basis of the above-mentioned embodiment, in the watermark adding model training device provided in the embodiment of the present invention, the loss calculation unit is specifically used for: Respectively extracting a first sentence vector from the first category of words and a second sentence vector from the second category of words; The semantic loss is calculated based on the first sentence vector and the second sentence vector.

[0106] Specifically, the functions of each module in the watermark adding model training device provided in the embodiment of the present invention correspond one-to-one to the operation flow of each step in the above method embodiment, and the effects achieved are also consistent. Please refer to the above embodiment for details, which will not be repeated in the embodiment of the present invention.

[0107] Fig.10 An example of a physical structure diagram of an electronic device is shown in FIG. Fig.10As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830 and a communication bus 840, wherein the processor 810, the communication interface 820 and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the watermark adding method, the watermark detection method or the watermark adding model training method provided in the above embodiments.

[0108] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the relevant technology or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0109] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the watermark adding method, watermark detection method, or watermark adding model training method provided in the above embodiments.

[0110] In another aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to execute the watermark adding method, watermark detection method, or watermark adding model training method provided in the above embodiments. The computer-readable storage medium can be either a non-transitory computer-readable storage medium or a transient computer-readable storage medium, which is not specifically limited here.

[0111] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0112] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.

[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A watermark adding method, characterized in that: include: Obtaining historical word units output by a text generation model and a probability distribution of a first word unit currently output by a text generation unit in the text generation model; Input the historical word-grams into the split model and the deviation model of the watermark adding model in the text generation model, and obtain the word-gram classification parameters output by the split model and the probability deviation parameters of different word-gram categories in the word-gram dictionary output by the deviation model; the word-gram classification parameters are used to classify each word-gram in the word-gram dictionary; Inputting the first word-unit probability distribution, the word-unit classification parameter and the probability deviation parameter into the watermark adding module of the watermark adding model, and obtaining the watermarked second word-unit probability distribution output by the watermark adding module; Based on the second word-unit probability distribution, a current word-unit currently output by the text generation model is determined.

2. The watermark adding method according to claim 1, characterized in that: The watermark adding module is specifically used for: Determining the current category of each word in the word-gram dictionary based on the word-gram classification parameter; Based on the probability deviation parameter and the current category of each word unit, the first word unit probability distribution is updated to obtain the second word unit probability distribution.

3. The watermark adding method according to claim 2, characterized in that: The word-unit classification parameters include a random seed and a distribution parameter of a preset distribution; Accordingly, the watermark adding module is specifically used for: Based on the random seed, the positions of the word-grams in the word-gram dictionary are rearranged, and probability values ​​equal to the number of word-grams in the word-gram dictionary are obtained by sampling from the preset distribution; The rearranged position of each word in the word-gram dictionary is matched to the probability value one by one, and the current category of each word in the word-gram dictionary is determined based on the probability value corresponding to the rearranged position of each word in the word-gram dictionary.

4. The watermark adding method according to any one of claims 1 to 3, characterized in that: Each word in the word-word dictionary includes watermark words and non-watermark words; Correspondingly, the probability deviation parameter includes a gain parameter corresponding to the watermark word element and a loss parameter of the non-watermark word element.

5. A watermark detection method, characterized in that: include: Determine a window to be detected of the text to be detected and a full text of the text to be detected that includes the window to be detected and its preceding text, and obtain word-element fragments of gradually increasing length in the full text; For any word-meta segment, based on the text generation unit in the text generation model, determine the third word-meta probability distribution of the last word-meta in the any word-meta segment, and apply the watermark adding method according to any one of claims 1 to 4 to determine the fourth word-meta probability distribution of the watermarked word-meta and the word-meta categories in the word-meta dictionary; Based on the fourth word-gram probability distribution, calculating the perplexity information entropy of the last word-gram, and based on the perplexity information entropy, calculating the weight of the last word-gram in the hypothesis test; Based on the number of watermark word units in the window to be detected, the word unit categories in the word unit dictionary corresponding to each word unit in the window to be detected and the proportion weight, the score of the hypothesis test is used as the detection index to determine whether the window to be detected contains a watermark.

6. The watermark detection method according to claim 5, characterized in that: The step of judging whether the window to be detected contains a watermark based on the number of watermark word units in the window to be detected, the word unit categories in the word unit dictionary corresponding to each word unit in the window to be detected, and the proportion weight, and taking the score of the hypothesis test as a detection index, comprises: Based on each word-gram category in the word-gram dictionary corresponding to each word-gram in the window to be detected, calculating the proportion of watermark word-grams in the word-gram dictionary corresponding to each word-gram in the window to be detected; Calculating the score of the hypothesis test based on the number of watermark word units in the window to be detected, the proportion of the watermark word units corresponding to each word unit in the window to be detected, and the proportion weight; If the score of the hypothesis test is greater than a preset threshold, it is determined that the window to be detected contains a watermark.

7. A watermark adding model training method, characterized in that: include: Acquire a text sample, input the text sample into an initial text generation model, obtain a sample probability distribution output by a text generation unit in the initial text generation model and a first type of word unit output by the initial text generation model, and determine a second type of word unit based on the sample probability distribution; Based on the watermark detection method according to any one of claims 5 to 6, respectively calculating a first score of the hypothesis test corresponding to the first category of words and a second score of the hypothesis test corresponding to the second category of words; Calculate semantic loss based on the first category word-gram and the second category word-gram, and calculate detection loss based on the first score and the second score; Based on the semantic loss and the detection loss, the initial watermark adding model in the initial text generation model is trained to obtain a trained watermark adding model.

8. The watermark adding model training method according to claim 7, characterized in that: The calculating the semantic loss based on the first category of words and the second category of words includes: Extracting a first sentence vector from the first category of words and a second sentence vector from the second category of words respectively; The semantic loss is calculated based on the first sentence vector and the second sentence vector.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the watermark adding method according to any one of claims 1 to 4, or the watermark detection method according to any one of claims 5 to 6, or the watermark adding model training method according to any one of claims 7 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the watermark adding method according to any one of claims 1 to 4, or the watermark detection method according to any one of claims 5 to 6, or the watermark adding model training method according to any one of claims 7 to 8 is implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the watermark adding method according to any one of claims 1 to 4, or the watermark detection method according to any one of claims 5 to 6, or the watermark adding model training method according to any one of claims 7 to 8 is implemented.

Citation Information

Patent Citations

  • Text watermark mark generation and detection method based on biased output large language model

    CN117494081A

  • Watermark embedding and detecting method and device for large model generation text

    CN118608367A

  • Text watermark embedding method in Logits generation period based on large language model

    CN119337344A

  • Big language model watermark detection method and system capable of being publicly verified

    CN119577708A

  • Text watermark adding and detecting method and device, equipment and storage medium

    CN119598427A