Controllable text watermark embedding method based on reinforcement learning strategy model
By combining a reinforcement learning strategy model with a multi-objective reward function and a watermark detector, high-quality and highly detectable text watermark embedding is achieved, solving the problem of fixed embedding positions in existing technologies and improving the naturalness and detectability of text watermarks.
Patent Information
- Application Number
- CN202511509238.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-10-22
AI Technical Summary
Existing text watermarking technologies rely on fixed strategies for embedding positions, lacking flexibility, potentially disrupting language fluency and contextual consistency, failing to dynamically adapt to the text generation environment or user needs, and struggling to achieve a balance between high quality and high detectability.
We adopt a controllable text watermarking embedding method based on reinforcement learning policy model. By constructing a multi-objective fusion reward function and policy network, and combining word substitution, grammatical perturbation and structural addition, we achieve token-level dynamic control. We also introduce a watermark detector to form a closed-loop optimization system and integrate it into a large language model.
Without interfering with the main LLM model structure, it significantly improves the naturalness, robustness, and detectability of text watermarks, and has good deployment flexibility and traceability capabilities.
Smart Images

Figure CN120974466A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of natural language processing and artificial intelligence generated content (AIGC) security, in particular to a controllable text watermark embedding method based on a reinforcement learning strategy model. BACKGROUND
[0002] In the prior art, with the wide application of large language models (LLMs) such as GPT, LLaMA, PaLM, etc., they have shown strong generation capabilities in natural language generation, dialogue systems, code completion, content creation, etc.
[0003] However, the "uncontrollability" of text generated by such models also brings potential abuse risks, including but not limited to the spread of false information, academic plagiarism, and automated water army content creation. In order to alleviate such risks, researchers have proposed a technology called "text watermarking" that aims to maintain naturalness while embedding verifiable and extractable hidden information into generated text.
[0004] Traditional text watermarking techniques include methods based on word replacement, syntactic structure reconstruction, and sentence tail appending, such as the "text watermark embedding and extraction method" disclosed in the authorized announcement number: CN 110414194 B. Although some progress has been made in embedding information, there are obvious shortcomings: (1) the embedding position depends on fixed strategies, lacking flexibility; (2) it may damage language fluency and context consistency; (3) it cannot dynamically adapt to the text generation environment or user needs. In addition, most existing methods cannot form an end-to-end closed-loop optimization system, making it difficult to balance the generation quality and watermark detectability.
[0005] In recent years, reinforcement learning (RL) as a goal-oriented learning mechanism has been widely used in policy control scenarios in generation tasks such as text summarization optimization and dialogue strategy generation. RL has the natural "goal-driven" characteristic, which is suitable for handling optimization problems with multiple conflicting goals. Therefore, introducing control signals and feedback mechanisms based on reinforcement learning in text watermark embedding has good potential.
[0006] However, directly applying reinforcement learning to text watermark embedding still has technical challenges, such as how to design appropriate reward functions to balance language quality and watermark detectability, and how to seamlessly integrate RL strategies into the LLM generation process to achieve token-level selection control.
[0007] Therefore, there is an urgent need to propose a new and feasible technical solution that can embed high-quality and highly detectable watermarks without interfering with the structure of the LLM master model and has good deployability and scalability. SUMMARY
[0008] The purpose of the present application is to provide a black-box prompt attack method based on a universal trigger capable of efficiently performing cross-task backdoor attacks; the technical solution is as follows: a controllable text watermark embedding method based on a reinforcement learning strategy model, comprising the following steps:
[0009] Step 1, construct a reinforcement learning strategy model to control the token output of the large language model through the reinforcement learning strategy model;
[0010] Step 2, under the strategy control of the reinforcement learning strategy model, embed watermarks in the output text based on the watermark embedding strategy to form a candidate token sequence with hidden information;
[0011] Step 3, establish a watermark detector to determine whether the watermark information has been successfully embedded in the text, and use the detection result as a signal to feed back to the reinforcement learning strategy network of the reinforcement learning strategy model to form a closed-loop training system;
[0012] Step 4, implement hierarchical encoding and mapping of identity information, encode specific identifiable information into watermark payload, and embed it into the text to realize the functions of "traceability, verifiability, and attribution";
[0013] Step 5, deploy the watermark embedder and complete model integration to integrate the watermark embedding mechanism into the existing large language model framework in a way that minimizes changes to support engineering deployment and large-scale applications.
[0014] Further, the reinforcement learning strategy model constructed in step 1 is as follows:
[0015] Step 1.1: a multi-objective fusion reward function design method is proposed to consider both text generation quality and watermark embedding feasibility, specifically, the reward function is designed as follows:
[0016] wherein, (1)
[0017] wherein, represents the confidence score of the embedded text being identified as "watermarked" by the watermark detector, taking a probability value of [0, 1], which can be obtained by the watermark detector in step 3 after inputting the text; denotes the text naturalness score, which measures the language style and semantic consistency between the generated text and the original language model output, defined as the KL divergence between the base language model distribution:
[0018] (2)
[0019] wherein, denotes the token probability distribution generated by the current reinforcement learning policy model, is the output distribution of the original language model; the smaller the KL divergence, the more similar the two are, i.e., the better the naturalness; therefore, take its opposite as the naturalness score item; at the same time, set the corresponding weight parameter and , adjust the balance between naturalness and embedding ability according to the actual application scene;
[0020] Step 1.2: By constructing and optimizing a reinforcement learning policy network to learn the most appropriate token- selection strategy under the condition of the given current generation context , so as to maximize the pre-defined reward function ; the training method adopts the standard policy gradient optimization algorithm, and its gradient expression is as follows:
[0021] (3)
[0022] wherein, denotes the token selected at time step , denotes the current context state, such as the generated text sequence; is the trainable parameter of the policy network; denotes the gradient operator of the parameter , i.e., the partial derivative vector; denotes the objective function of policy optimization , the gradient of the parameter θ;
[0023] Step 1.3: Deploy the reinforcement learning policy network trained in step 1.2 as a token reranker, which is used to adjust the token selection logic in real time during the actual text generation process, so as to realize the accurate control of embedding behavior at the token level, and select the optimal token according to the following formula:
[0024] (4)
[0025] wherein, wherein, y denotes any one token in the candidate set, Vt represents the candidate token set (such as the top-k candidate set) output by the large language model at time step t, which contains all possible tokens that can be selected, represents the token finally selected, is the probability of the candidate token scored by the policy network at the current state.
[0026] Further, the specific method of embedding the watermark in the output text based on the watermark embedding strategy in step 2 is as follows:
[0027] Step 2.1: Obtain the candidate token set of the language model At each generation step , extract the top-k tokens with the highest probability from the distribution output by the language model to form the candidate set , that is:
[0028] (5)
[0029] This set is the input basis for subsequent watermark bit and token mapping; where LLM is a Large Language Model, i.e., a large language model;
[0030] Step 2.2: Assume that the watermark information to be embedded is represented as a bit sequence where is the i-th bit; the system selects a suitable token from the candidate set according to the current bit at each step t:
[0031] (6)
[0032] This mapping function can be implemented through position bias, hash mapping or lookup table, and is suitable for different language model output distributions.
[0033] Step 2.3: Use a semantic-preserving syntactic perturbation technique to perturb the sentence structure, including subject-predicate-object order replacement, phrase insertion, and use of attributive clauses, so that the original sentence and the perturbed sentence remain close in semantic space:
[0034] (7)
[0035] where represents the semantic embedding representation, a sentence vector extracted by a model such as BERT, to ensure that the sentence content remains consistent after perturbation; represents the original sentence, i.e., the large language model generated text before any syntactic perturbation. denotes the perturbed sentence, i.e. the sentence generated by syntactic transformation while keeping the semantic unchanged;
[0036] Step 2.4: Additional structural token sequences are added as watermark carrying areas, and the structured token sequence is added at the end of the text when the context length allows to encode the remaining watermark bits:
[0037] (8)
[0038] wherein, is a function for mapping bit string to token sequence, which can be implemented based on the template library; denotes the original generated text sequence, denotes the complete text after adding the structured watermark token sequence;
[0039] Step 2.5: In order to adapt to the embedding needs of different contexts, a strategy controller is used to determine the optimal embedding strategy according to the current context state and the embedding state Output strategy score:
[0040] (9)
[0041] wherein, is a set of candidate strategies, is a strategy score function, which comprehensively evaluates the embeddability, semantic preservation and context adaptability, and finally realizes dynamic embedding path planning.
[0042] Further, the specific method for establishing a watermark detector and participating in joint training feedback in step 3 is:
[0043] Step 3.1: Construct a watermark detector based on the Transformer architecture as a discrimination module to join the entire watermark training system. The watermark detector judges any input text x and outputs the probability score of containing watermark. The sigmoid function is used for normalization to get the confidence score:
[0044] (10)
[0045] wherein is the logit value output by the Transformer encoder, denotes the sigmoid activation function, and the output value is used to measure the strength of the current text watermark signal;
[0046] Step 3.2: Construct the training set using the corpus with and without watermark, and perform supervised training, with the loss function being cross-entropy:
[0047] (11)
[0048] wherein, represents the label of whether the sample contains a watermark, is the confidence of the model prediction; this training process enables the watermark detector to distinguish between watermarked and non-watermarked text;
[0049] Step 3.3: To form a joint optimization mechanism for watermark embedding and detection, the watermark confidence output by the watermark detector is introduced into the reward function of the reinforcement learning policy model as an evaluation indicator of watermark embedding effectiveness:
[0050] (12)
[0051] Further, the specific method for implementing hierarchical bit encoding of identity information in step 4 is as follows:
[0052] Step 4.1: First, extract the core metadata from the generation task, including the user identity UserID, generation timestamp Timestamp, and task number TaskID, and encode them using the encoding function to uniformly encode the user identity UserID, generation timestamp Timestamp, and task number TaskID into a bit stream .
[0053] Step 4.2: Based on the candidate set at each token generation moment , the bit stream is mapped bit by bit to the corresponding position token; the mapping function determines the corresponding candidate token according to the bit value, mapping bit 0 to the natural token with a distribution closer to the front, and mapping bit 1 to the token subset that has been disturbed or offset, thereby completing information embedding without significantly compromising language quality.
[0054] Step 4.3: A hierarchical embedding mechanism is introduced, with the watermark bit stream being divided into a macro layer and a micro layer: the macro layer is used to embed global information, and the micro layer is used to embed task-level information.
[0055] Further, the specific method of step 5 is as follows:
[0056] Step 5.1: The watermark embedder is encapsulated as an independent "post-processing module" to perform strategic reordering and selection of the output sequence in an intercepting manner after the large language model LLM completes the token probability distribution output; the "post-processing module" does not modify any network parameters or architecture of the main model, but is attached to the decoding process as a token selection reorderer to realize fine-tuning embedding of the output token;
[0057] Step 5.2: The watermark embedder is integrated through a standardized API interface call, and the user loads the watermark control strategy through HTTP request or local call, including generated text content, user identity information, and task configuration parameters, and automatically completes watermark encoding, embedding, and control output.
[0058] Further, the encoding method in step 4.1 is set to a standard hash function, AES encryption hash, or variable-length binary encoding, and the generated bit stream is used to control the selection of specific tokens in the subsequent watermark embedding process.
[0059] Further, the global information of the macro layer in step 4.3 includes user ID and authorization code; the task-level information of the micro layer includes the context identifier or version number of the specific generated content.
[0060] Beneficial effects: The present application has the following beneficial effects: The present application designs a unified framework combining reinforcement learning control mechanism and multi-strategy text watermark embedding, dynamically intervenes in the token level of the large language model generation process through the introduction of a policy network based on a reward function, simultaneously fuses multiple embedding methods such as vocabulary replacement, syntax disturbance, and structure addition, and provides feedback signals with a watermark detector to realize end-to-end closed-loop optimization, thereby significantly improving the naturalness, robustness, and detectability of the text watermark without modifying the main structure of the language model, and having good deployment flexibility and traceability. BRIEF DESCRIPTION OF DRAWINGS
[0061] Figure 1 The overall flowchart of the present application is shown in the figure;
[0062] Figure 2 The overall framework of the present application is shown in the figure. DETAILED DESCRIPTION
[0063] The present application will be further illustrated below in conjunction with the drawings and specific embodiments, which are implemented on the premise of the technical scheme of the present application, and it should be understood that these embodiments are only used to illustrate the present application and not to limit the scope of the present application.
[0064] As shown in Figure 1 and Figure 2 , the controllable text watermark embedding method based on the reinforcement learning strategy model has the following specific steps:
[0065] Step 1: Model the language model generation process as a sequential decision-making process, with the state being the current generated text segment, the action being all the available tokens in the space, and the policy function taking the current token distribution as input and outputting a reordering probability distribution over the candidate token set. To train this policy model, a multi-objective reward function is introduced that considers both language generation quality (e.g., KL divergence from the base LLM output) and watermark embedding effectiveness (e.g., watermark detector recognition confidence), and a policy gradient-like method is used to optimize the policy network. The specific steps are as follows:
[0066] Step 1.1: Propose a multi-objective fusion reward function design method that considers both text generation quality and watermark embedding feasibility. Specifically, the reward function is designed as follows:
[0067] where (1)
[0068] where represents the confidence score of the embedded text being recognized as "watermarked" by the watermark detector, usually taking a probability value in [0, 1], which can be obtained by the watermark detector after inputting the text; represents the naturalness score of the text, which measures the language style and semantic consistency between the generated text and the original language model output, defined as the KL divergence between the two distributions:
[0069] (2)
[0070] Here represents the token probability distribution generated by the current policy model, is the output distribution of the original language model (e.g., GPT or BERT); the smaller the KL divergence, the more similar the two, i.e., the better the naturalness; therefore, take its opposite as the naturalness score item; at the same time, set appropriate weight parameters and to adjust the balance between naturalness and embedding ability according to the actual application scenario.
[0071] Step 1.2: Construct and optimize a reinforcement learning policy network to learn the strategy of selecting the most suitable token- under the given current generation context , so as to maximize the pre-defined reward function . The training method uses the standard policy gradient (Policy Gradient) optimization algorithm, and its gradient expression is as follows:
[0072] (3)
[0073] wherein, wherein, denotes the time step the token selected at time step t, denotes the current context state, such as the text sequence already generated; are the trainable parameters of the policy network, denotes the gradient operator, i.e. the vector of partial derivatives, with respect to the parameters ; denotes the objective function of the policy optimization with respect to the parameters
[0074] Step 1.3: Deploy the reinforcement learning policy network trained in step 1.2 as a token reranker, which is used to adjust the token selection logic in real time during the actual text generation process, so as to realize the accurate control of embedding behavior at the token level, and select the optimal token according to the following formula:
[0075] (4)
[0076] wherein, y denotes any token in the candidate set, V t denotes the candidate token set (such as top-k candidate set) output by the large language model at time step t, which contains all possible tokens that can be selected, denotes the token finally selected, is the scoring probability of the candidate token by the policy network under the current state;
[0077] Step 2: According to the watermark information (usually encoded as a binary bit string) required to be embedded, select the token corresponding to the target bit from the candidate token output by the language model for embedding, so as to realize information injection under the condition of semantic preservation; The specific steps are as follows:
[0078] Step 2.1: Obtain the top-k candidate token set of the language model, and extract the top-k tokens with the highest probability from the distribution output by the language model at each generation step to form the candidate set , i.e.
[0079] (5)
[0080] This set is the input basis for subsequent watermark bit and token mapping;
[0081] Step 2.2, assuming the watermark information to be embedded is represented as a bit sequence where is the i-th bit; the system at each step t follows the current bit Select a suitable token from the candidate set:
[0082] (6)
[0083] This mapping function can be implemented through position bias (e.g. even bits are encoded as 1), hash mapping (e.g. , or look-up table (e.g. preset dictionary index) to flexibly adapt to different language model output distributions;
[0084] Step 2.3: Use semantic-preserving syntactic perturbation techniques to perturb the sentence structure, including subject-predicate-object order replacement, phrase insertion, use of attributive clauses, etc., so that the original sentence and the perturbed sentence remain close in semantic space:
[0085] (7)
[0086] where represents the semantic embedding representation, the sentence vector extracted by models such as BERT, ensuring that the sentence content remains consistent after perturbation; represents the original sentence, i.e. the large language model generated text before any syntactic perturbation; represents the perturbed sentence, i.e. the sentence generated by syntactic transformation under the premise of keeping the semantics unchanged;
[0087] Step 2.4: Add a structured token sequence as a watermark carrying area, when the context length allows, a structured token sequence can be added at the end of the text to encode the remaining watermark bits:
[0088] (8)
[0089] where, is a function used to map the bit string to the token sequence, which can be implemented based on a template library; represents the original generated text sequence, represents the complete text after adding the structured watermark token sequence, which can be implemented based on a template library, for example, embedding in natural language expression such as: "This output was generated at [ID:0101101]".
[0090] Step 2.5: To adapt to the embedding needs of different contexts, use a strategy controller to control the embedding strategy according to the current context state with the embedded state output the policy score, select the optimal embedding strategy :
[0091] (9)
[0092] wherein, is a set of candidate strategies (such as lexical replacement, syntactic perturbation, structural addition), is a strategy scoring function that comprehensively evaluates embeddability, semantic preservation, and contextual adaptability, ultimately achieving dynamic embedding path planning;
[0093] Step 3: Train a watermark detector to identify whether watermark information is embedded in the generated text, and further use the identification result as a feedback signal for the reinforcement learning strategy to train the strategy model, thereby building a closed-loop optimization system of generation-detection-feedback; The specific steps are as follows:
[0094] Step 3.1: Build a watermark detector based on the Transformer architecture as a discrimination module into the entire watermark training system. The detector judges any input text x and outputs the probability score of containing watermark, which is normalized to confidence score by sigmoid function:
[0095] (10)
[0096] wherein, is the logit value output by the Transformer encoder, denotes the sigmoid activation function, and the output value is used to measure the strength of the current text watermark signal;
[0097] Step 3.2: Use watermarked and non-watermarked corpora to build a training set for supervised training, and the loss function is cross-entropy:
[0098] (11)
[0099] wherein denotes the label of whether the sample contains watermark, is the confidence score predicted by the model, and the training process enables the detector to distinguish between watermarked and non-watermarked texts;
[0100] Step 3.3: To form a joint optimization mechanism for watermark embedding and detection, introduce the watermark confidence score output by the detector into the reward function of the reinforcement learning strategy as an evaluation index for watermark embedding effect:
[0101] (12)
[0102] Step 4: Encode and manage the information required for user or platform identification, such as user ID, platform ID, timestamp, task number, etc., generate structured watermark load, and embed it as a binary bit sequence into the text generation process; the specific steps are as follows:
[0103] Step 4.1: First extract the core metadata from the generation task, including user identity (UserID), generation timestamp (Timestamp), and task number (TaskID), and pass them through the encoding function Encode the user identity UserID, generation timestamp Timestamp, and task number TaskID, and convert them into a bit stream ;
[0104] The encoding method can be a standard hash function, AES encryption hash, or variable-length binary encoding. The generated bit stream is used to control the selection of specific tokens in the subsequent watermark embedding process;
[0105] Step 4.2: Based on the candidate set at each token generation moment , map the bit stream to the corresponding token at the corresponding position; the mapping function determines the corresponding candidate token based on the bit value; for example, bit 0 is mapped to the natural token at the front of the distribution, and bit 1 is mapped to the token subset after disturbance or offset, so as to complete information embedding without significantly damaging language quality;
[0106] Step 4.3: A hierarchical embedding mechanism is introduced; the watermark bit stream can be divided into macro and micro layers; the macro layer is used to embed global information such as user ID and authorization code; the micro layer is used to embed task-level information such as context identification or version number of specific generated content;
[0107] Step 5: Integrate the entire embedding process into the existing large language model system in the form of a modular plug-in, as follows:
[0108] Step 5.1: Encapsulate the watermark embedder as an independent "post-processing module" that intercepts the output sequence after the LLM completes the token probability distribution output and performs strategic reordering and selection. This "post-processing module" does not modify any network parameters or architecture of the main model, but rather acts as a token selection reorderer attached to the decoding process (such as sampling or beam search) to achieve fine-tuning embedding of the output token;
[0109] Step 5.2: The watermark module is integrated through a standardized API interface call, and the user loads the watermark control strategy through an HTTP request or a local call method, inputting content including generated text, user identity information, task configuration parameters, etc. The module automatically completes watermark encoding, embedding, and control output.
[0110] Embodiment 1
[0111] To verify the effectiveness of the scheme of the present application, the following experiment is performed in this embodiment, using LLaMA-7B as the underlying language model, and GPT-2 as the baseline model. Without modifying the network structure of the baseline model, the sampling output is reranked by the strategy network.
[0112] The strategy network is of a Transformer structure, guided by a reward function with a reward coefficient set to , where the naturalness score and the watermark detectability come from a jointly trained detector. The decoding strategy is Top-k (k=8) and reranking based on reinforcement learning.
[0113] As shown in Table 1, the performance comparison data of the watermark embedding method based on the reinforcement learning strategy model of the present application and the traditional fixed strategy watermark in terms of text naturalness and detectability; where the evaluation index is perplexity, which measures the "naturalness" of the text to the language model, and the lower the value, the better; watermark detection accuracy (Detect Accuracy), which measures the classification accuracy of whether the text contains watermark (such as whether it contains system-embedded identity information); bit recovery rate (BRR), which measures the proportion of successfully recovered watermark bits from the embedded text; the following table is the performance comparison of the watermark embedding method based on reinforcement learning in terms of text naturalness and detectability:
[0114] Table 1
[0115]
[0116] Analysis of experimental results: From the experimental results in Table 1, it can be seen that the reinforcement learning watermark embedding strategy proposed in the present application effectively balances between text naturalness preservation and watermark detectability. Compared with the traditional rule-based watermark method (such as the fixed strategy watermark in the table), the PPL value is reduced from 20.6 to 20.0 (the lower the PPL value, the better); the DetectAccuracy value is improved from 94.2% to 96.7%; and the BRR value is improved from 85.5% to 89.1%; all the data comparisons are better than the fixed strategy watermark method in the prior art.
[0117] The adaptive embedding control mechanism formed by the training strategy of the application can automatically select a more covert embedding mode according to the context, and further enhance the concealment and attack resistance of the system.
[0118] The above specific embodiment is only one preferred embodiment of the application, and is not intended to limit the implementation and scope of claims of the application. Any equivalent changes and modifications made in accordance with the content of the application patent protection scope shall be included in the application patent protection scope.
Claims
1. A controllable text watermark embedding method based on a reinforcement learning policy model, characterized in that, Includes the following steps: Step 1: Construct a reinforcement learning policy model to control the token output of the large language model. Step 2: Under the policy control of the reinforcement learning policy model, watermark embedding is performed on the output text based on the watermark embedding strategy to form a candidate token sequence with hidden information. Step 3: Establish a watermark detector to determine whether watermark information has been successfully embedded in the text, and pass the detection result as a signal back to the reinforcement learning policy network of the reinforcement learning policy model to form a closed-loop training system. Step 4: Implement hierarchical encoding and mapping of identity information, encode specific identifiable information into watermark payload, and embed it into the text to achieve the functions of "traceability, verifiability, and attributability"; Step 5: Deploy the watermark embedder and complete model integration. Integrate the watermark embedding mechanism into the existing large language model framework with minimal modifications to support engineering deployment and large-scale application.
2. The controllable text watermark embedding method based on a reinforcement learning policy model according to claim 1, characterized in that, The specific steps for constructing the reinforcement learning policy model in step 1 are as follows: Step 1.1: A multi-objective fusion reward function design method is proposed to simultaneously consider text generation quality and watermark embedding feasibility. Specifically, the reward function is designed as follows: in, (1); in, This represents the confidence score of the embedded text being identified as "watermarked" by the watermark detector. The value is [0, 1] and is obtained by the watermark detector in step 3 after inputting the text. The text naturalness score, used to measure the stylistic and semantic consistency between the generated text and the original language model output, is defined as the KL divergence with the distribution of the underlying language model. (2); in, This represents the probability distribution of tokens generated by the current reinforcement learning policy model. The output distribution of the original language model is shown. A smaller KL divergence indicates greater similarity between the two models, meaning better naturalness. Therefore, its negative value is taken as the naturalness score. Simultaneously, corresponding weight parameters are set. and Adjust the balance between naturalness and embedding capability according to the actual application scenario; Step 1.2: Construct and optimize a reinforcement learning policy network To learn in a given current generation context Under the given conditions, select the most suitable token. The strategy is to maximize the predefined reward function. The training method employs the standard Policy Gradient optimization algorithm, whose gradient expression is as follows: (3); in, Indicates time step The token selected at the time Indicates the current context state, such as the already generated text sequence; These are the trainable parameters of the policy network. Indicates the parameter The gradient operator, i.e., the partial derivative vector; The objective function for policy optimization is represented by The gradient with respect to the parameter θ; Step 1.3: Deploy the reinforcement learning policy network trained in Step 1.2 as a token reranker. This reranker is used to adjust the token selection logic in real time during the actual text generation process, thereby achieving precise control over the embedding behavior at the token level. The optimal token is selected according to the following formula: (4); Where y represents any token in the candidate set, V t This represents the set of candidate tokens output by the large language model at time step t, which contains all possible tokens that can be selected. This indicates the finally selected token. For the policy network to select candidate tokens in the current state - The probability of scoring.
3. The controllable text watermark embedding method based on a reinforcement learning policy model according to claim 1, characterized in that, The specific method for embedding a watermark in the output text based on the watermark embedding strategy in step 2 is as follows: Step 2.1: Obtain the language model The candidate token set at each generation step Extract the top-ranked words with the highest probability from the distribution output by the language model. Each token constitutes a candidate set. ,Right now: (5); This set serves as the input basis for subsequent watermark bit-to-token mapping; where LLM stands for Large Language Model. Step 2.2: Assume the watermark information to be embedded is represented as a one-bit sequence. ,in For the i-th bit; the system follows the current bit in each step t. Select a suitable token from the candidate set: (6); This mapping function is implemented through positional bias, hash mapping, or table lookup to adapt to the output distribution of different language models. Step 2.3: Employ semantically preserving syntactic perturbation techniques to perturb the sentence structure. Methods include subject-verb-object substitution, phrase insertion, and the use of relative clauses, ensuring that the original sentence and the perturbed sentence remain semantically close. (7); in The semantic embedding representation is obtained by extracting sentence vectors from the BERT model, ensuring that the sentence content remains consistent with the original meaning after perturbation. This represents the original sentence, generated by a large language model before any syntactic perturbations. This refers to the sentence generated through syntactic transformations after perturbation, while preserving its semantics. Step 2.4: Use the appended structured token sequence as the watermark carrying area. When the context length allows, add the remaining watermark bits to the end of the text using the structured token sequence encoding. (8); in, It is a function used to map bit strings to token sequences, implemented based on a template library; This represents the original generated text sequence. This represents the complete text following the appended structured watermark token sequence; Step 2.5: To adapt to the embedding requirements of different contexts, a reinforcement learning policy model is used based on the current context state. With embedded state Output the policy score and select the optimal embedding policy. : (9); in, For the set of candidate strategies, The strategy scoring function comprehensively evaluates embeddability, semantic preservation, and contextual adaptability, ultimately achieving dynamic embedding path planning.
4. The controllable text watermark embedding method based on a reinforcement learning policy model according to claim 1, characterized in that, The specific method for establishing the watermark detector and participating in joint training feedback in step 3 is as follows: Step 3.1: Build a watermark detector based on the Transformer architecture As a discrimination module added to the entire watermark training system, the watermark detector judges any input text x and outputs its probability score of containing the watermark, which is then normalized to a confidence score using the sigmoid function. (10); in, The logit value output by the Transformer encoder. This represents the sigmoid activation function and its output value. This is used to measure the strength of the current text watermark signal; Step 3.2: Construct a training set using both watermarked and unwatermarked corpora, and perform supervised training. The loss function is cross-entropy. (11); in, A label indicating whether a sample contains a watermark. It is the confidence level of the model's prediction; this training process enables the watermark detector to distinguish between watermarked and non-watermarked text. Step 3.3: To form a joint optimization mechanism for watermark embedding and detection, the watermark confidence score output by the watermark detector is... It is incorporated into the reward function of the reinforcement learning policy model as an evaluation metric for the watermark embedding effect: (12)。 5. The controllable text watermark embedding method based on a reinforcement learning policy model according to claim 1, characterized in that, The specific method for implementing hierarchical bit encoding of identity information in step 4 is as follows: Step 4.1: First, extract core metadata from the generated task, including the user ID, generation timestamp, and task ID, and then use an encoding function. The UserID, Generation Timestamp, and TaskID are uniformly encoded and converted into a bitstream. ; Step 4.2: Candidate set based on each token generation time , bit stream Bitwise mapping to the token at the corresponding position; The mapping function determines the corresponding candidate token based on the bit value, maps bit 0 to the natural tokens that are distributed earlier, and maps bit 1 to the perturbed or offset subset of tokens, thereby completing the information embedding without significantly compromising the language quality. Step 4.3: A layered embedding mechanism is introduced, and the watermark bitstream is divided into a macro layer and a micro layer: the macro layer is used to embed global information; the micro layer is used for task-level information.
6. The controllable text watermarking embedding method based on a reinforcement learning policy model according to claim 1, characterized in that, The specific method for step 5 is as follows: Step 5.1: The watermark embedder is encapsulated as an independent "post-processing module". After the large language model LLM completes the output of the token probability distribution, the output sequence is reordered and selected in an interception manner. This "post-processing module" does not modify any network parameters or architecture of the main model, but is attached to the decoding process as a token selection and reordering unit to achieve fine-tuning of the embedded output token. Step 5.2: The watermark embedder is integrated through a standardized API interface call. Users load the watermark control strategy through HTTP request or local call. The input includes the generated text content, user identity information, and task configuration parameters. The watermark encoding, embedding, and control output are completed automatically.
7. The controllable text watermark embedding method based on a reinforcement learning policy model according to claim 5, characterized in that, The encoding method in step 4.1 is set to a standard hash function, AES encrypted hash, or variable-length binary encoding. The generated bit stream is used to control the selection of specific tokens in the subsequent watermark embedding process.
8. The controllable text watermarking embedding method based on a reinforcement learning policy model according to claim 5, characterized in that, The global information at the macro level in step 4.3 includes the user ID and authorization code; the task-level information at the micro level includes the context identifier or version number of the specific generated content.
Citation Information
Patent Citations
A method for embedding and extracting text watermarks
CN110414194B
Text watermark embedding and detecting method based on model context learning
CN118349970A
Watermark generation method and system based on large language model
CN119577707A
Watermark processing
US20250086257A1
Generation and detection of watermark for real-time voice conversion
WO2021030759A1
Cited By
Dynamic watermark embedding method and device based on text statistical characteristics and optimization strategy
CN121637465A
Dynamic watermark embedding method and device based on text statistical features and optimization strategy
CN121637465B
Language model multi-bit embedding-based imperceptible layered watermark embedding method
CN121997303A