Watermark generation method and device based on large model, equipment and medium
By embedding watermarks in text generated by large models, the problem of how to protect intellectual property rights without affecting the quality of the text is solved, and efficient watermark embedding and detection is achieved.
Patent Information
- Application Number
- CN202510135226.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-07
AI Technical Summary
How to embed watermarks to protect intellectual property and prevent intellectual property abuse without affecting the quality of text generated by the big model.
By obtaining the target prompt word sequence, a sequence of words to be generated is generated, and some words are randomly selected as target words in the word sequence, and watermark information is embedded into these words according to the preset watermark intensity to form watermark mark words, and splicing them in order to generate watermark generated text.
It realizes embedding of watermarks without significantly affecting text quality, reducing the impact of watermarks on the overall generated text quality, while ensuring the existence and feasibility of detection of watermarks.
Smart Images

Figure CN120068027A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of watermark technology, and in particular, to a watermark generation method, device, equipment and medium based on a large model. Background Art
[0002] Large models can generate highly realistic and logical texts, but their powerful generation capabilities also bring potential risks of intellectual property abuse. Watermark technology, as a means of information hiding, has been introduced into the field of intellectual property protection in the era of large models. Large model watermarks have attracted much attention due to their unique embedding and detection methods. How to embed watermarks without affecting the quality of model-generated text has become an urgent problem to be solved. Summary of the invention
[0003] The present application provides a large model-based watermark generation method, device, equipment and medium to solve one or more technical problems existing in the prior art and at least provide a beneficial choice or create conditions.
[0004] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.
[0005] According to one aspect of an embodiment of the present application, a method for generating a watermark based on a large model is provided, the method comprising:
[0006] Acquire a target prompt word sequence, and acquire a word sequence to be generated according to the target prompt word sequence, wherein the word sequence includes a plurality of words in order;
[0007] Selecting one or more of the words in the word sequence as target words according to a preset randomly selected seed, and embedding watermark information into each of the target words according to a preset watermark strength to obtain each watermarked word;
[0008] According to the order of each of the words, each of the watermarked words and each of the words other than the watermarked words are concatenated to obtain a watermark generated text;
[0009] A target watermark text is generated according to the target prompt word sequence and the watermark generation text.
[0010] In one embodiment of the present application, based on the above scheme, the target prompt word sequence is obtained by the following steps:
[0011] Get the target prompt word text;
[0012] Splitting the target prompt word text according to a preset large language model to obtain multiple prompt words;
[0013] Generate the target prompt word sequence according to each of the foregoing prompt words.
[0014] In one embodiment of the present application, based on the foregoing solution, the obtaining the word sequence to be generated according to the target prompt word sequence includes:
[0015] Obtain each of the words in turn according to the target prompt word sequence and the vocabulary set of the large language model;
[0016] Generate the word sequence according to each of the words.
[0017] In one embodiment of the present application, based on the foregoing solution, the selecting one or more of the words as target words in the word sequence according to a preset random selection seed includes:
[0018] Group each of the words according to the number of words in the word sequence to obtain multiple groups of word subsequences, and each group of word subsequences includes at least two of the words;
[0019] For each group of word subsequences, select any one of the words in the word subsequence as the target word according to the random selection seed.
[0020] In one embodiment of the present application, based on the foregoing solution, the embedding watermark information into each of the target words according to a preset watermark strength to obtain each watermark-labeled word includes:
[0021] Generate a watermark attack matrix according to the vocabulary set;
[0022] Attack each of the target words according to the watermark attack matrix and the watermark strength to embed watermark information, and obtain each of the watermark-labeled words.
[0023] In one embodiment of the present application, based on the foregoing solution, the generating a target watermark text according to the target prompt word sequence and the watermark generation text includes:
[0024] Perform text splicing on the target prompt word text corresponding to the target prompt word sequence and the watermark generation text to obtain the target watermark text.
[0025] In one embodiment of the present application, based on the foregoing solution, after generating the target watermark text according to the target prompt word sequence and the watermark generation text, the method further includes:
[0026] Obtain the first total distribution probability of each word in the watermark generation text and the second total distribution probability of each word in the word sequence;
[0027] If the first total distribution probability is greater than the second total distribution probability, it is determined that the generated target watermark text does not contain watermark information;
[0028] If the first total distribution probability is greater than the second total distribution probability, it is determined that the generated target watermark text does not contain a watermark;
[0029] If the first total distribution probability is less than the second total distribution probability, it is determined that the generated target watermark text contains a watermark.
[0030] According to one aspect of the embodiments of the present application, a watermark generation device based on a large model is provided. The device includes:
[0031] An acquisition unit, configured to acquire a target prompt word sequence and acquire a sequence of words to be generated according to the target prompt word sequence. The sequence of words includes a plurality of words in order;
[0032] A watermark embedding unit, configured to select one or more of the words in the sequence of words as target words according to a preset randomly selected seed, and embed watermark information into each of the target words according to a preset watermark strength to obtain each watermark marked word;
[0033] A splicing unit, configured to splice each of the watermark marked words and each word other than the watermark marked words in the order of each word to obtain a watermark generation text;
[0034] A watermark text generation unit, configured to generate a target watermark text according to the target prompt word sequence and the watermark generation text.
[0035] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. The computer program includes executable instructions, and when the executable instructions are executed by a processor, the method described in the above embodiments is implemented.
[0036] According to one aspect of the embodiments of the present application, an electronic device is provided, including: one or more processors; a memory for storing executable instructions of the processor, and when the executable instructions are executed by the one or more processors, the one or more processors implement the method described in the above embodiments.
[0037] Advantageous effects of the present application: Through the preset target prompt word sequence given in advance, a corresponding sequence of words can be generated. By selecting some of the words in the sequence of words as target words for watermark embedding to obtain watermark marked words, and then splicing the words in order according to the order of the watermark marked words in the sequence of words and the words without watermark embedding, a watermark generation text with a watermark is obtained.
[0038] Further, by concatenating the target prompt word sequence with the watermark generation text, a complete target watermark text is obtained. It can be seen that only some words are embedded with watermarks in this application, rather than all words, which can reduce the impact of the watermark on the overall generated text quality. Therefore, both the concatenated watermark generation text and the target watermark text have watermarks, and the text quality is guaranteed, and there will be no large deviation in the understanding of the prompt word sequence.
[0039] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. Description of the Drawings
[0040] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0041] Figure 1 is a conventional token generation method shown according to an embodiment of this application;
[0042] Figure 2 is a flowchart of a watermark generation method based on a large model shown according to an embodiment of this application;
[0043] Figure 3 is a comparison diagram of grouped watermark embedding and global watermark embedding for grouping each word in a word sequence and embedding watermarks according to randomly selected seeds shown according to an embodiment of this application;
[0044] Figure 4 is an architecture diagram of the TWT-LLM model network architecture shown according to an embodiment of this application;
[0045] Figure 5 is a text quality comparison diagram under different watermark strengths shown according to an embodiment of this application;
[0046] Figure 6 is a block diagram of a watermark generation device based on a large model shown according to an embodiment of this application;
[0047] Figure 7 is a structural diagram of an electronic device shown according to this application. Detailed Embodiments
[0048] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0049] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this application. However, those skilled in the art will realize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other instances, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of this application.
[0050] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontrol node devices.
[0051] The flowcharts shown in the drawings are merely illustrative and do not necessarily include all the content and operations / steps, nor do they necessarily have to be executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0052] It should be noted that: "a plurality of" as mentioned herein means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0053] The background art of the embodiments of this application will be introduced in detail below:
[0054] With the rapid development of artificial intelligence technology, especially the wide application of large language models such as GPT, Claude, and LLaMA, text similar to that of humans can be generated quickly. The text generated by large language models looks very realistic and is therefore easily misused to create and spread misleading information such as fake news and speech. The content generated by large models may contain unauthorized personal information, infringe on personal privacy, and even harass or threaten specific individuals. As these large models become more common, the risk of their being used for malicious purposes is increasing. There is a risk of using large models to infringe on the copyrights of the creative works of many authors on the Internet and to generate a large amount of false information on major social platforms. At the same time, on platforms such as social media and news websites, malicious users or organizations may use large models to generate false information to mislead public opinion, undermine social stability, or seek improper benefits. Such false information often spreads rapidly and widely, causing confusion and panic among the public and even affecting government decision-making and judicial fairness. How to effectively protect the copyrights of these complex models and human intellectual property has become an urgent problem to be solved. Therefore, the ability to detect and audit the use of machine-generated text and to detect and mark text generated by LLMs has become a key method for large language models to reduce harm.
[0055] An effective method to accurately detect text generated by large models is through watermarking technology. The watermarking technology used in large models is a method of embedding specific information or identifiers in deep learning models. The generated text will be marked imperceptibly, and finally its source can be determined based on the mark. Its purpose is to achieve copyright protection, traceability, or integrity verification of the model without significantly affecting the performance of the model. There are inherent dimensional feature differences between watermarked text with characteristic information and normally generated text. Figure 1It shows that the large language model GPT2-XL has different intrinsic dimensions in the normal text and watermarked text generated by inputting a prompt text sequence. In addition, we can also find that compared with the watermarked text, the text generated without adding watermark is relatively regular in the process of generating words, while the latter is interfered in the generation process due to the addition of watermark information, so it appears relatively chaotic. Therefore, this also poses a challenge to our research work. In the process of adding watermark, we need to find ways to minimize the impact on the semantics of the generated text. In the research of this application, the watermark generation method based on the large model proposed in this application can run on its own through social media platforms, or can be kept private and run in the background after being encapsulated into an API. This application can perform algorithm detection without knowing the model parameters or accessing the language model API, and allows the detection algorithm to be open source, which can also make the detection cheap and fast because the LLM does not need to be loaded or run. Moreover, the watermarked text can be generated using a standard language model without retraining. At the same time, the target watermarked text detected in this application is continuous, so that when the target watermarked text is put into a large document, the existence of the watermark can still be detected.
[0056] The implementation details of the technical solutions of the embodiments of this application are elaborated in detail below:
[0057] According to one aspect of this application, a watermark generation method based on a large model is provided. Figure 2 As shown in the flowchart of the watermark generation method based on the large model according to the embodiments of this application, the watermark generation method based on the large model at least includes steps S1 to S4, which are introduced in detail as follows:
[0058] In step S1, obtain a target prompt sequence, and obtain a sequence of words to be generated according to the target prompt sequence, where the sequence of words includes a plurality of words in order.
[0059] Specifically, the target prompt sequence is obtained through the following steps:
[0060] Obtain a target prompt text;
[0061] According to a preset large language model, split the target prompt text into multiple prompts;
[0062] Generate the target prompt sequence according to each of the prompts.
[0063] The target prompt text can be as Figure 1The prompt text shown in 1 , namely "I'm glad we have a vacation now! The weather is nice today," 2 , can be split into prompt words through a large language model, which is represented by the letter L in this application. That is, it is split into 12 prompt words: "very", "glad", "we", "already", "have", "vacation", "now", "!", "today", "weather", "nice", ",". Then, after inputting these 12 prompt words into the large language model L, they will be converted into a corresponding sequence of prompt words, that is, the target sequence of prompt words described in this application. The target sequence of prompt words is represented by T=(t n ).
[0064] The watermark generation method based on a large model proposed in this application is called TWT-LLM. The TWT method will embed a certain amount of watermark information during the sampling process in the process of generating the subsequent word sequence according to the target sequence of prompt words.
[0065] In an embodiment of this application, obtaining the word sequence to be generated according to the target sequence of prompt words includes:
[0066] Obtaining each of the words in turn according to the target sequence of prompt words and the vocabulary set of the large language model;
[0067] Generating the word sequence according to each of the words.
[0068] Continuing from the above, before inputting the 12 prompt words into the large language model L, the obtained sequence can be represented by χ=(x 1 , x 2 , …, x n ), and a probability distribution P n can be generated on the vocabulary set ν of the large language model. The size of this vocabulary set is |ν|. Therefore, the probability distribution of the next token can be obtained through P n = L(χ). Here, a token is the word described in this application, and the word sequence described in this application is composed of multiple tokens in order.
[0069] The watermark generation function can be represented by the mathematical symbol M. In fact, it is also a slight change to the model L. If the change is too large, it will cause the generated distribution probability to be chaotic, thus affecting the text quality. By inputting the sequence χ=(x 1 , x 2 , …, x n ), the probability distribution of the next token with a watermark can be obtained using P′ n = M(χ).
[0070] In step S2, one or more of the words are selected as target words from the word sequence according to a preset random selection seed, and watermark information is embedded into each of the target words according to a preset watermark strength to obtain each watermarked word.
[0071] In one embodiment of the present application, the selecting one or more of the words as target words from the word sequence according to a preset random selection seed includes:
[0072] Grouping each of the words according to the number of words in the word sequence to obtain multiple groups of word subsequences, where each group of word subsequences includes at least two of the words;
[0073] For each group of word subsequences, any one of the words in the word subsequence is selected as the target word according to the random selection seed.
[0074] Specifically, since global watermark embedding is performed on the entire process of generating words by the large language model L, it may not only affect the quality of the generated text, but also make the model TWT-LLM lose its security. Therefore, in the embodiments of the present application, an optimization strategy is proposed to improve the quality of the generated text after embedding, while it is difficult for the human eye to see that watermark information is hidden. As Figure 3 shown, texts are generated using two different watermark embedding strategies. The upper part of the figure shows the strategy of global watermark embedding, which embeds watermarks into all tokens during the generation process and has some impact on the quality of text generation. The lower part of the figure shows that based on the group size n group the tokens to be generated are divided into k groups. It should be noted that the TWT-LLM model is also the model network architecture shown in the appendix of the present application Figure 4 shown.
[0075] X = G 1 ∪G 2 ∪…∪G k
[0076] For each group G i apply the watermark embedding function M to obtain the result M(G i ), where for each group G i randomly select N token_mk tokens for watermark embedding, and the random seeds for selecting tokens are all set to S win ,
[0077]
[0078] Since the global embedding strategy embeds watermarks in all tokens, error accumulation will occur during the generation process, which may lead to the phenomenon of hallucinations in the model TWT-LLM. However, for the local watermark embedding based on grouping, the interference generated is very subtle. And since the large language model generates tokens based on the previous sequence of prompt words and tends to generate subsequent tokens by understanding the correct meaning, there will be no error accumulation. Therefore, it can be clearly analyzed that compared with the global embedding strategy, the local embedding strategy based on grouping has less impact on text generation.
[0079] In an embodiment of the present application, the embedding of watermark information into each of the target words according to a preset watermark strength to obtain each watermarked word includes:
[0080] Generating a watermark attack matrix according to the vocabulary set;
[0081] Attacking each of the target words according to the watermark attack matrix and the watermark strength to embed watermark information, so as to obtain each of the watermarked words.
[0082] Specifically, reference can be made to Figure 4 as shown Figure 4 which is the logical diagram of the entire watermark embedding of the present application. First, the target prompt word text is input into the large language model, that is, "I'm glad we have a holiday now! The weather is nice today," and then a target prompt word sequence that can be recognized and used by the watermark generator is obtained. Further, by adding the interference of the watermark strength to the generation of each token, that is, watermarked generated words (i.e., the watermarked words described in the present application) can be obtained.
[0083] In an embodiment of the present application, the embedding of watermark information into each of the target words according to a preset watermark strength to obtain each watermarked word includes:
[0084] Generating a watermark attack matrix according to the vocabulary set;
[0085] Attacking each of the target words according to the watermark attack matrix and the watermark strength to embed watermark information, so as to obtain each of the watermarked words.
[0086] Specifically, the steps of watermark generation are as follows:
[0087] 1: Input: Watermark generation model: M. Watermark strength: α. Prompt word sequence: T = (t 1 , t 2 , …, t n ). Generation length: δ
[0088] 2: Calculate the given word sequence χ(x 1, x 2 , …, x n ) The output logit P x0:n , obtaining the probability distribution of the next word sequence
[0089] Generate the watermark attack matrix WK, and the size of this matrix is determined by the vocabulary ν of the large model.
[0090] 4: Use the watermark strength α for attack and calculation α ∈ [0, 1]. When the value of the watermark strength α is 0, it is equivalent to not using the watermark algorithm. When the value of the watermark strength α is 1, the quality of the generated text will be very poor.
[0091] 5: Concatenate the prompt word sequence and the generated watermark sequence
[0092] 6: Obtain a reconstructed large model watermark model M.
[0093] 7: Output: The fine-tuned large model watermark model M.
[0094] In step S3, according to the order of each of the said words, splice each of the said watermark marked words and each word other than the said watermark marked words to obtain the watermark generated text.
[0095] Continue to refer to Figure 4 As shown, by splicing each of the said watermark marked words and each word other than the said watermark marked words in order, the target watermark text "I'm glad we have a vacation now! The weather is nice today, but we don't want to go shopping together anymore" with watermark can be obtained. Then, "but we don't want to go shopping together anymore" is the watermark generated text described in this application.
[0096] In step S4, generate the target watermark text according to the said target prompt word sequence and the said watermark generated text
[0097] The generating the target watermark text according to the said target prompt word sequence and the said watermark generated text includes:
[0098] Splice the target prompt word text corresponding to the said target prompt word sequence and the said watermark generated text to obtain the said target watermark text.
[0099] Through Figure 4 It can be seen that "I'm glad we have a vacation now! The weather is nice today, but we don't want to go shopping together anymore" is the complete target watermark text. According to Figure 4As shown, first, the target prompt word sequence generates a token sequence through the large language model. The token sequence is the expression that the large language model can understand. Then, through the watermark generator, watermark information can be embedded while minimizing the impact on the text generation quality, thereby affecting the distribution of the next word (token) to be generated. After that, the target prompt word sequence and the watermark generated text are concatenated and transmitted to the watermark detection module. Finally, the watermark detection can detect the complete concatenated text according to the output rules, and detect and determine the existence of the watermark.
[0100] In an embodiment of the present application, after generating the target watermark text according to the target prompt word sequence and the watermark generated text, the method further includes:
[0101] Obtaining the first total distribution probability of each word in the watermark generated text and the second total distribution probability of each word in the word sequence;
[0102] If the first total distribution probability is greater than the second total distribution probability, it is determined that the generated target watermark text does not contain watermark information;
[0103] If the first total distribution probability is greater than the second total distribution probability, it is determined that the generated target watermark text does not have a watermark;
[0104] If the first total distribution probability is less than the second total distribution probability, it is determined that the generated target watermark text has a watermark.
[0105] Specifically, the function of the watermark detection method can be represented by the mathematical symbol D, which can quickly and effectively detect the watermark generated text. By inputting the prompt text T=(t 1 ,t 2 ,…,t n ), calculate Y 1 ,Y 2 =D(T).
[0106] Assume that Y 1 is calculated under the condition that it uses the watermark, that is, the first total distribution probability, and Y 2 is calculated under the assumption that it has no watermark, that is, the second total distribution probability. If Y 1 -Y 2 satisfies the following formula, it means there is a watermark, otherwise there is no watermark.
[0107]
[0108] Furthermore, the present application uses a custom detection method to determine the existence of the watermark. Given a fixed large language model L, watermark strength α, and the number of generated words δ. Denote the probability of each token as Calculate the total distribution probability of the word sequence with and without watermark respectively, denoted as Y 1 ,Y 2 , the formula is as follows
[0109]
[0110] When Y 1 ,Y 2 When there is a certain relationship as follows, the existence of the watermark can be judged.
[0111]
[0112] Considering that as the scale of the large language model increases, the vocabulary size of the large model also increases. For example, for the vocabulary size |ν| of GPT2, the probability of a certain word may be very small. If the above detection method is directly used, the two evaluation values Y 1 ,Y 2 are very difficult to compare. In order to amplify the internal feature gap between the text without adding watermark and the text with added watermark, this application improves the evaluation and detection method.
[0113] Similarly, denote the probability of each token as Calculate the total distribution probability of the word sequence with and without watermark respectively, denoted as Y 1 ,Y 2 , the formula is as follows
[0114]
[0115] When the Y obtained by using the improved detection strategy 1 ,Y 2 When there is a certain relationship as follows, the existence of the watermark can be judged.
[0116]
[0117] Furthermore, the preset watermark strength α is 0.15. The size of the watermark strength α will have a direct and significant impact on Figure 4 the effect of the model network architecture in. For example Figure 5 , it can be shown that when the watermark strength α is too large, it will affect the text generation quality, and when the watermark strength α is too small, it will be very difficult to detect the existence of the text watermark. Through experiments, it is found that the effect is the best when α = 0.15.
[0118]
[0119] Figure 5It shows the generation situation of watermarked text when GPT is selected as the original model, by setting different values of the watermark strength α and generating text with watermarks through the given prompt text. It can be seen that when α = 0.1, the generated text content is basically not affected. However, when α = 1.0, the model is like a drunkard talking nonsense.
[0120] In summary, the method proposed in this application can embed watermarks on the premise of minimizing the impact on the quality of the text generated by the model, and can quickly detect the watermarked generated text.
[0121] According to one aspect of the embodiments of the present application, a watermark generation device 300 based on a large model is proposed. Figure 6 It is a schematic diagram of the watermark generation device 300 based on the large model proposed in the embodiments of the present application. The device 300 includes: an acquisition unit 301, a watermark embedding unit 302, a splicing unit 303, and a watermark text generation unit 304.
[0122] The acquisition unit 301 is configured to acquire a target prompt word sequence, and acquire a sequence of words to be generated according to the target prompt word sequence. The sequence of words includes a plurality of words in order.
[0123] The watermark embedding unit 302 is configured to select one or more of the words in the sequence of words as target words according to a preset randomly selected seed, and embed watermark information into each of the target words according to a preset watermark strength to obtain each watermark-tagged word.
[0124] The splicing unit 303 is configured to splice each of the watermark-tagged words and each of the words other than the watermark-tagged words according to the order of each word to obtain a watermark-generated text.
[0125] The watermark text generation unit 304 is configured to generate a target watermark text according to the target prompt word sequence and the watermark-generated text.
[0126] As another aspect, the present application also provides a computer-readable storage medium, on which a program product capable of implementing the method provided in the above description of this specification is stored. In some possible implementation manners, each aspect of the present application can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present application described in the "Embodiment Method" section of the above description of this specification.
[0127] A program product for implementing the above method according to an embodiment of the present application may be a portable compact disc read-only memory (CD-ROM), include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0128] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0129] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0130] The program code contained on the readable medium may be transmitted by any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0131] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0132] Reference will now be made to Figure 7 describe the electronic device 400 according to this embodiment of the present application. Figure 7 The electronic device 400 shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.
[0133] As Figure 7 shown, the electronic device 400 is presented in the form of a general-purpose computing device. The components of the electronic device 400 may include, but are not limited to: at least one of the above-mentioned processing units 410, at least one of the above-mentioned storage units 420, and a bus 430 connecting different system components (including the storage unit 420 and the processing unit 410).
[0134] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 410, so that the processing unit 410 executes the steps according to various exemplary embodiments of the present application described in the "Embodiment Method" section of this specification.
[0135] The storage unit 420 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 421 and / or a cache storage unit 422, and may further include a read-only storage unit (ROM) 423.
[0136] The storage unit 420 may also include a program / utility 424 having a set (at least one) of program modules 425. Such program modules 425 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0137] The bus 430 may represent one or more of several types of bus structures, including a memory bus or a memory control node, a peripheral bus, an Accelerated Graphics Port, a processor unit, or a local bus using any of the various bus structures.
[0138] The electronic device 400 may also communicate with one or more external devices 1200 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 400, and / or may communicate with any device that enables the electronic device 400 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through the input / output (I / O) interface 450. Moreover, the electronic device 400 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 460. As shown in the figure, the network adapter 460 communicates with other modules of the electronic device 400 through the bus 430. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 400, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0139] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by the way of software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which may be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0140] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present application, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes may be executed synchronously or asynchronously in, for example, multiple modules.
[0141] It should be understood that the present application is not limited to the exact structures that have been described and shown in the drawings above, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A watermark generation method based on a large model, characterized in that: The method comprises: Acquire a target prompt word sequence, and acquire a word sequence to be generated according to the target prompt word sequence, wherein the word sequence includes a plurality of words in order; Selecting one or more of the words in the word sequence as target words according to a preset randomly selected seed, and embedding watermark information into each of the target words according to a preset watermark strength to obtain each watermarked word; According to the order of each of the words, each of the watermarked words and each of the words other than the watermarked words are concatenated to obtain a watermark generated text; A target watermark text is generated according to the target prompt word sequence and the watermark generation text.
2. The watermark generation method based on a large model according to claim 1 is characterized in that: The target prompt word sequence is obtained by the following steps: Get the target prompt word text; Splitting the target prompt word text according to a preset large language model to obtain multiple prompt words; The target prompt word sequence is generated according to each of the prompt words.
3. The watermark generation method based on a large model according to claim 2 is characterized in that: The step of obtaining a word sequence to be generated according to the target prompt word sequence includes: Acquire each of the words in sequence according to the target prompt word sequence and the vocabulary set of the large language model; The word sequence is generated according to each of the words.
4. The watermark generation method based on a large model according to claim 3 is characterized in that: The step of selecting one or more words as target words in the word sequence according to a preset random selection seed includes: According to the number of the words in the word sequence, each of the words is grouped to obtain a plurality of groups of word subsequences, each group of the word subsequences including at least two of the words; For each group of the word subsequences, any one of the words in the word subsequences is selected as the target word according to the randomly selected seed.
5. The watermark generation method based on a large model according to claim 4 is characterized in that: The step of embedding watermark information into each target word according to a preset watermark strength to obtain each watermarked word includes: generating a watermark attack matrix according to the vocabulary set; Attack each of the target words according to the watermark attack matrix and the watermark strength to embed watermark information, thereby obtaining each of the watermarked words.
6. The watermark generation method based on a large model according to claim 5 is characterized in that: The step of generating a target watermark text according to the target prompt word sequence and the watermark generation text comprises: The target prompt word text corresponding to the target prompt word sequence and the watermark generation text are concatenated to obtain the target watermark text.
7. The watermark generation method based on a large model according to claim 6 is characterized in that: After generating the target watermark text according to the target prompt word sequence and the watermark generation text, the method further includes: Obtaining a first total distribution probability of each word in the watermark generation text and a second total distribution probability of each word in the word sequence; If the first total distribution probability is greater than the second total distribution probability, it is determined that the generated target watermark text does not contain watermark information; If the first total distribution probability is greater than the second total distribution probability, it is determined that the generated target watermark text does not have a watermark; If the first total distribution probability is less than the second total distribution probability, it is determined that the generated target watermark text has a watermark.
8. A watermark generation device based on a large model, characterized in that: The device comprises: An acquisition unit, used to acquire a target prompt word sequence, and acquire a word sequence to be generated according to the target prompt word sequence, wherein the word sequence includes a plurality of words in order; A watermark embedding unit, configured to select one or more words in the word sequence as target words according to a preset randomly selected seed, and embed watermark information into each of the target words according to a preset watermark strength to obtain each watermarked word; A concatenation unit, configured to concatenate the watermarked words and the words other than the watermarked words in the order of the words to obtain a watermark generated text; The watermark text generating unit is used to generate a target watermark text according to the target prompt word sequence and the watermark generation text.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the large model-based watermark generation method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the large model-based watermark generation method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Digital watermark adding method and device, electronic equipment and storage medium
CN112948776A
Response text generation method based on large language model, electronic equipment and medium
CN117932010A
Model tracing method, device and equipment and readable storage medium
CN117932572A
Text watermark detection and watermark adding method, program product, equipment and medium
CN118656810A
Sentence semantics-based watermarking method for large language model transfer attack
CN118821086A
Cited By
Black box large language model watermark based on sampling and preferential output and detection method
CN122433060A