A closed-loop error feedback multi-bit watermark embedding and extraction method for large language model short text generation
Patent Information
- Application Number
- CN202611035618.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-29
AI Technical Summary
当生成文本仅包含几十个词元时,每个比特维度所能获得的统计观测极为有限,使得提取过程对随机采样噪声高度敏感
本发明所记载的上述技术方案,针对大语言模型在短文本生成场景下难以稳定嵌入并恢复多比特水印信息的技术瓶颈,提出了一种基于闭环误差反馈的嵌入与提取机制。具体而言,该方法将水印嵌入过程构建为一个持续的状态跟踪与误差修正闭环:在每一生成步中,系统根据当前上下文和密钥为每个候选词元生成双极性的特征向量,同时实时维护所有已生成词元在各载荷维度上的全局累积状态,并将其与按生成进度增长的目标轨迹进行比较,从而获得每一维度上当前欠编码或过编码的实时误差。基于该误差向量与各候选词元特征向量的内积计算反馈得分,并用该得分动态调制候选词元的概率分布,使得后续词元的选择能够主动补偿尚未达到目标的载荷维度。
Smart Images

Figure CN122839348A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of natural language processing, artificial intelligence security, and content tracing of large language models, and particularly relates to a closed-loop error feedback multi-bit watermark embedding and extraction method for short text generation of large language models. Background Technology
[0002] With the widespread application of large language models in scenarios such as writing assistance, dialogue systems, machine translation, code generation, and information retrieval, the amount of text content generated by artificial intelligence is increasing dramatically. To achieve source tracking, attribution of responsibility, copyright protection, and abuse detection of generated content, an effective technical approach is to embed watermark information—which is difficult for humans to perceive but can be detected and extracted by algorithms—when the model generates text.
[0003] Currently, text watermarking technologies for large language models are mainly divided into two categories: zero-bit watermarking and multi-bit watermarking. Zero-bit watermarking is typically used only to determine whether a piece of text was generated by a specific watermarking model, and its output is a binary classification (presence or absence of watermark). Although this type of method has high detection efficiency, it is difficult to carry structured traceability information such as user identifiers, session numbers, timestamps, or model version numbers. This limitation has directly driven research into multi-bit watermarking technology.
[0004] In multi-bit watermarking, the watermark is not only a detectable marker but also a recoverable payload. Typically, Yoo et al. proposed the MPAC scheme, employing a time-division multiplexing strategy to assign different text positions or subsets of markers to different payload bits, ultimately recovering the complete message by aggregating the voting data from each bitstream. Feng et al. proposed the BiMark method, utilizing an unbiased multi-layer watermark structure and using bit flipping to assign specific subsets of the generated sequence to each bit for encoding. Jiang et al. proposed StealthInk, a method based on optimization to design a watermark embedding strategy, focusing on the watermark's concealment and the fluency of the generated text, embedding information through probability interval distribution, but sacrificing the reliability of watermark extraction to some extent. Park et al. proposed WaterMod, sorting the vocabulary in descending order according to the model's output probability, partitioning the vocabulary based on the remainder after modulo the probability ranking, and then embedding watermark information by applying statistical bias to different partitions. Xu et al. proposed XMark, which provides multiple observation opportunities for each lexical unit by repeatedly partitioning the vocabulary and taking the intersection of all resulting "green lists" as the final watermark candidate set. In addition, Qu et al. attempted to introduce error-correcting codes to enhance the robustness of the decoding end, but their improvements were mainly focused on the decoding stage and did not substantially optimize the embedding mechanism in the generation process.
[0005] While the aforementioned methods demonstrate certain advantages in their respective areas of focus, they generally face severe reliability bottlenecks in short text generation scenarios. These methods essentially allocate watermark capacity locally or in an open-loop manner, meaning each generation step encodes information independently based on the current context, lacking awareness and control over the global encoding state. When the generated text contains only a few dozen tokens, the statistical observations obtainable per bit dimension are extremely limited, making the extraction process highly sensitive to random sampling noise. More critically, due to the open-loop design, once an early encoding deviation occurs, the system cannot measure the cumulative difference between the expected message and the actual watermark state, nor can it reallocate the token budget in subsequent generation to correct the lagging encoding dimensions.
[0006] In summary, existing multi-bit watermarking schemes for short texts suffer from structural technical bottlenecks: the system must simultaneously encode multiple bits within an extremely limited lexical budget, and existing designs lack the ability to adapt to the randomness of the generation process, leading to a significant decrease in watermark extraction accuracy under high-load, short-text scenarios. Therefore, there is an urgent need for a new multi-bit watermark embedding and extraction technique that enables large language models to stably and reliably embed and recover multi-bit watermark information under the stringent conditions of short texts and high loads. Summary of the Invention
[0007] To address the aforementioned technical problems, this invention provides a closed-loop error feedback multi-bit watermark embedding and extraction method for short text generation using large language models, comprising the following steps: Obtain the multi-bit binary message payload to be embedded, where the message payload is used to characterize the source information of the generated text; In the process of autoregressive text generation by a large language model, a feature vector with the same dimension as the message payload is generated for each candidate word in the vocabulary, based on the key, the current context and candidate words. Maintain the global state vector of the feature vectors of the generated words; Calculate the target trajectory vector based on the currently generated number of steps and the target message vector obtained by converting the multi-bit binary message payload; Calculate the error vector based on the difference between the target trajectory vector and the global state vector; The feedback score of the candidate word is calculated based on the inner product of the feature vector and the error vector, and the probability distribution of the candidate word is modulated based on the feedback score. The current word is sampled based on the modulated probability distribution, and the global state vector is updated with the feature vector of the current word to iteratively execute subsequent generation steps until a complete watermarked text is generated. During watermark extraction, the feature vectors of each word in the watermarked text are recalculated using the same key and linearly accumulated. The multi-bit binary message payload is then recovered based on the sign of the accumulated values in each dimension.
[0008] Optionally, obtaining the multi-bit binary message payload to be embedded specifically includes converting one or more of the following into a binary sequence: user identifier, generation session identifier, timestamp, model version number, service instance identifier, or content authorization information.
[0009] Optionally, calculating the target trajectory vector specifically includes multiplying the preset trajectory slope parameter, the currently generated number of steps, and the target message vector to obtain the target trajectory vector.
[0010] Optionally, calculating the feedback score of the candidate lexical unit specifically includes: performing a dot product operation between the feature vector of the candidate lexical unit and the error vector, then dividing by the dimension of the message payload, and using the resulting quotient as the feedback score.
[0011] Optionally, modulating the probability distribution of candidate lexical units based on feedback scores specifically includes: multiplying the feedback score of each candidate lexical unit by the watermark intensity parameter, adding it to the original logits to obtain the modulated logits, and then inputting the modulated logits into a normalized exponential function to obtain the modulated probability distribution.
[0012] Optionally, the watermark strength parameter adopts a dynamic scheduling strategy: when the number of generation steps is less than a preset threshold, an initial strength value is used; when the number of generation steps is greater than or equal to the preset threshold, a gradually decreasing strength value is used until a preset minimum strength value is reached.
[0013] Optionally, recovering the multi-bit binary message payload based on the sign of the accumulated value of each dimension specifically includes: for each dimension, if the accumulated value is positive, it is determined to be +1, and if the accumulated value is negative, it is determined to be -1, thus obtaining a bipolar message, and then converting the bipolar message back into a binary message payload.
[0014] Optionally, linear accumulation during watermark extraction specifically includes: initializing the integral vector as a zero vector, extracting feature vectors for each word and accumulating them into the integral vector to obtain the accumulated value for each dimension.
[0015] On the other hand, the present invention also provides a closed-loop error feedback multi-bit watermark embedding and extraction system for short text generation of large language models, used to implement the method, including: The payload acquisition module is used to acquire the multi-bit binary message payload to be embedded, wherein the message payload is used to characterize the source information of the generated text; The feature generation module is used to generate a feature vector with the same dimension as the message payload for each candidate word in the vocabulary during the autoregressive text generation process of the large language model, based on the key, the current context and candidate words. The state maintenance module is used to maintain the global state vector of the feature vectors of the generated lexical units; The trajectory calculation module is used to calculate the target trajectory vector based on the currently generated number of steps and the target message vector converted from the multi-bit binary message payload; The error calculation module is used to calculate the error vector based on the difference between the target trajectory vector and the global state vector. The feedback modulation module is used to calculate the feedback score of the candidate word based on the inner product of the feature vector of the candidate word and the error vector, and to modulate the probability distribution of the candidate word based on the feedback score. The sampling update module is used to sample the current word element according to the modulated probability distribution, and update the global state vector with the feature vector of the current word element to iteratively execute the subsequent generation steps until the complete watermarked text is generated. The extraction and decoding module is used to recalculate the feature vectors of each word in the watermarked text using the same key and perform linear accumulation during watermark extraction, and recover the multi-bit binary message payload based on the sign of the accumulated values of each dimension.
[0016] On the other hand, the present invention also provides an electronic device including a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the processor implements the method when executing the computing program.
[0017] Compared with the prior art, the present invention has the following advantages and technical effects: The technical solution described in this invention addresses the technical bottleneck of large language models struggling to stably embed and recover multi-bit watermark information in short text generation scenarios. It proposes an embedding and extraction mechanism based on closed-loop error feedback. Specifically, this method constructs the watermark embedding process as a continuous state tracking and error correction closed loop: In each generation step, the system generates a bipolar feature vector for each candidate word based on the current context and key. Simultaneously, it maintains the global cumulative state of all generated words in each load dimension in real time and compares it with the target trajectory that grows according to the generation progress, thereby obtaining the real-time error of undercoding or overcoding in each dimension. A feedback score is calculated based on the inner product of this error vector and the feature vectors of each candidate word. This score is then used to dynamically modulate the probability distribution of candidate words, enabling the selection of subsequent words to proactively compensate for load dimensions that have not yet reached the target.
[0018] Through the aforementioned closed-loop feedback mechanism, this invention overcomes the limitations of traditional open-loop watermarking methods, where each generation step is independently encoded and early biases cannot be corrected. It enables all generated tokens to contribute to the joint observation across all payload dimensions, significantly improving the reliability and accuracy of multi-bit watermark extraction even with extremely limited text length. Furthermore, since the decoding end only needs to perform linear accumulation and symbol decision based on the same key for the token feature vectors in the text to be detected, without iterative search or complex post-processing, the decoding process has low computational complexity and low latency, facilitating practical system deployment. In addition, this method exhibits good robustness against common text perturbations such as paraphrasing attacks and random token substitution, effectively supporting applications such as source tracing, copyright protection, and abuse detection of generated text. Attached Figure Description
[0019] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is an overall flowchart of the closed-loop error feedback watermarking framework according to an embodiment of the present invention; Figure 2 This is a comparison chart (50 words) of the bit extraction accuracy of the present invention and the baseline method under different payload capacities. Figure 3 This is a schematic diagram comparing the average extraction time under different payload capacities in an embodiment of the present invention (50 words). Figure 4 This is a comparison diagram of the robustness of the present invention and the baseline method under the DIPPER paraphrase attack in an embodiment of the present invention (100 words, 8 bits). Figure 5 This is a comparison diagram (100 words, 8 bits) of the robustness of the present invention and the baseline method under random substitution attacks in an embodiment of the present invention. Detailed Implementation
[0020] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0021] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0022] Example 1 This embodiment provides a closed-loop error feedback multi-bit watermark embedding and extraction method for short text generation of large language models, including the following steps: Obtain the multi-bit binary message payload to be embedded, where the message payload is used to characterize the source information of the generated text; In the process of autoregressive text generation by a large language model, a feature vector with the same dimension as the message payload is generated for each candidate word in the vocabulary, based on the key, the current context and candidate words. Maintain the global state vector of the feature vectors of the generated words; Calculate the target trajectory vector based on the currently generated number of steps and the target message vector obtained by converting the multi-bit binary message payload; Calculate the error vector based on the difference between the target trajectory vector and the global state vector; The feedback score of the candidate word is calculated based on the inner product of the feature vector and the error vector, and the probability distribution of the candidate word is modulated based on the feedback score. The current word is sampled based on the modulated probability distribution, and the global state vector is updated with the feature vector of the current word to iteratively execute subsequent generation steps until a complete watermarked text is generated. During watermark extraction, the feature vectors of each word in the watermarked text are recalculated using the same key and linearly accumulated. The multi-bit binary message payload is then recovered based on the sign of the accumulated values in each dimension.
[0023] Compared to existing open-loop multi-bit watermarking methods, this embodiment no longer simply assigns fixed lexical positions to fixed bits. Instead, it ensures that each generated lexical contributes features across all payload dimensions and dynamically allocates subsequent generation budgets to the currently under-encoded watermark dimension through real-time error feedback. Here, "large language model" refers to a neural network language model capable of generating lexical sequences based on contextual autoregression. A lexical is a generation unit in the model's vocabulary, which can be a word, subword, byte fragment, or other word segmentation unit. The message payload refers to the multi-bit information that needs to be embedded and recovered at the watermark extraction end. Closed-loop error feedback refers to a control mechanism that continuously compares the accumulated watermark state with the target trajectory during the generation process and adjusts the distribution of subsequent candidate lexical logits based on this error.
[0024] As a possible implementation method, such as Figure 1 As shown, the specific steps include: (1) Bipolar conversion of payload messages: Obtain the multi-bit binary message payload to be embedded, wherein the message payload is used to characterize the source information of the generated text, including one or more of the following: user identifier, generation session identifier, timestamp, model version, service instance identifier, and content authorization information.
[0025] The binary message payload to be embedded is converted into a bipolar representation to obtain the target message vector, which facilitates vector accumulation and inner product feedback control.
[0026] ; in The original binary load is given, where L is the load length. This is a bipolar target message vector.
[0027] (2) Feature mapping: In the large language model In the generation step, the basic logits of candidate nouns are output based on the current context. Based on the key, the current context, and the candidate nouns, pseudo-random feature vectors based on the key and context are generated for the candidate nouns in the vocabulary to ensure that the detector is reproducible but unpredictable without the key.
[0028] ; Where K is the key shared by watermark embedding and extraction. This refers to the preceding information of the current generation step. The current candidate lexical unit is denoted as . Mapping the payload message and feature vector to the bipolar domain transforms the original binary data into a symbolic domain suitable for vector accumulation and direction determination, enabling subsequent target trajectory, error vector, and linear integral decoding to be expressed within the same vector space.
[0029] This step enables the watermark extraction end to reconstruct the same feature vector without knowing the error adjustment history during generation, and makes it difficult for a third party without holding the key to predict the feature vector.
[0030] (3) Closed-loop state tracking and trajectory error vector calculation: (3.1) Global State Maintenance: Maintain the global state vector and accumulate the feature vectors of the generated words: ; in, This represents the token sampled in the k-th step, so that each generated token contributes to the parallel observations of the L payload dimensions. The global state maintained in this step will be compared with the subsequent target trajectory. (3.2) Target trajectory definition: Calculate the target trajectory vector based on the current generated step number and the target message vector: ; in The trajectory slope parameter represents the expected cumulative rate of the watermark signal per step along the specified message payload direction. For the current number of steps generated, The target message payload vector; (3.3) Error vector calculation: Calculate the deviation between the target trajectory vector and the global state vector: ; in, This is the global watermark state vector at the end of the previous step. In the above process, the global cumulative watermark state represents the actual cumulative contribution of the actually sampled words in each payload dimension up to the current generation step; the target trajectory represents the ideal cumulative state that the system expects the watermark signal to reach along the bipolar message payload direction at the current generation progress. The difference between the two is the closed-loop error vector. If the closed-loop error vector is positive in a certain dimension, it means that the accumulation in that payload dimension is insufficient relative to the target trajectory, and subsequent generation should prioritize words that can increase the contribution in that dimension; if the closed-loop error vector is negative in a certain dimension, it means that the payload dimension has exceeded the target trajectory, and subsequent generation should suppress words that continue to accumulate along that direction, so that the selection of subsequent words can be dynamically adjusted according to the current watermark embedding state.
[0031] (4) Error-driven logits modulation mechanism: (4.1) Feedback-driven scoring: The score for each candidate word is calculated based on the inner product of the error vector and the feature vectors of each candidate word. ; in For candidate noun v in the current context The eigenvectors below, Given the current error vector, the inner product formula can be interpreted as the projection of the candidate word feature vector onto the error direction. Specifically, the feedback score of the candidate word is used to measure its ability to compensate for the current closed-loop error. The larger the score, the closer the watermark contribution direction of the candidate word is to the current error direction that needs to be compensated. (4.2) Dynamic logits modulation: The modulated logits are obtained by using the scoring modulation base logits as follows: ; in The watermark strength parameter is a piecewise function that is adaptively determined by a function that decreases with each generation step. When the absolute value of the error in a certain dimension is large, the scoring function is dominated by that dimension, so that the large language model will preferentially select the lexical units that correct the lag bits. (4.3) Watermark strength scheduling: A step-decay watermark strength scheduling strategy is adopted. in =0.8, which is the initial watermark strength; γ=0.99, which is the attenuation coefficient; =0.1, which is the minimum watermark strength. A higher intensity is used in the early stages of generation to help the system quickly align the state trajectory, while the intensity is gradually reduced in the later stages to mitigate text distortion and maintain the naturalness of the final text.
[0032] (5) State update and iteration: Add the feature vectors of the actual sampled words to the global state: ; in, The feature vector of the sampled word in the current step is used to update the context, and steps (2) to (5) are repeated until the predetermined generation length or termination condition is reached. Therefore, steps (2) to (5) form a closed loop logic: step (2) generates the corresponding feature vector for the candidate word according to the context and key; step (3) calculates the current error according to the current global state and the target trajectory; step (4) modulates the logits distribution of the candidate word according to the current error; and step (5) updates the global state according to the actual sampling results and feeds the updated state back to the next round of step (2). This loop enables the system to continuously correct early sampling deviations during the generation process, avoiding the problem of uncompensated errors in open-loop watermarking methods.
[0033] (6) Linear integral decoding mechanism: The progressive error modulation process at the watermark extraction end and the watermark embedding end is decoupled. Although the embedding end dynamically adjusts the logits distribution of candidate words using a closed-loop error vector during generation, the observable result ultimately reflected in the watermarked text is the cumulative direction of the pseudo-random feature vectors of each generated word across all payload dimensions. Therefore, the extraction end only needs to use the same key and the same context construction rules to reconstruct the corresponding feature vectors for each word in the text to be detected and perform linear accumulation.
[0034] (6.1) Linear accumulation of feature vectors: For a given text of length N, the decoder at the watermark extraction end extracts the feature vector of each word and calculates the global cumulative integral: ; (6.2) Symbol Decision Extraction: The values of the extracted multi-bit payload messages are determined based on the signs of each dimension of the integral vector. ; (6.3) Binary Conversion: Converting bipolar messages back to binary payload domain: .
[0035] Compared to existing technologies, the method proposed in this embodiment constructs the watermark embedding process as a continuous state tracking and error correction closed loop: In each generation step, the system generates a bipolar feature vector for each candidate word based on the current context and key, while simultaneously maintaining the global cumulative state of all generated words in each payload dimension in real time, and comparing it with the target trajectory that grows according to the generation progress, thereby obtaining the real-time error of undercoding or overcoding in each dimension. A feedback score is calculated based on the inner product of this error vector and the feature vectors of each candidate word, and this score is used to dynamically modulate the probability distribution of candidate words, enabling the selection of subsequent words to proactively compensate for payload dimensions that have not yet reached the target.
[0036] Through the aforementioned closed-loop feedback mechanism, this embodiment overcomes the limitations of traditional open-loop watermarking methods, where each generation step is independently encoded and early biases cannot be corrected. It enables all generated tokens to contribute to joint observations across all payload dimensions, significantly improving the reliability and accuracy of multi-bit watermark extraction even with extremely limited text length. Furthermore, since the decoding end only needs to perform linear accumulation and symbol decision based on the same key for the token feature vectors in the text to be detected, without iterative search or complex post-processing, the decoding process has low computational complexity and low latency, facilitating practical system deployment. In addition, this method exhibits good robustness to common text perturbations such as paraphrasing attacks and random token substitutions, effectively supporting applications such as source tracing, copyright protection, and abuse detection of generated text.
[0037] On the other hand, this embodiment also provides a closed-loop error feedback multi-bit watermark embedding and extraction system for short text generation of large language models, used to implement the method, including: The payload acquisition module is used to acquire the multi-bit binary message payload to be embedded, wherein the message payload is used to characterize the source information of the generated text; The feature generation module is used to generate a feature vector with the same dimension as the message payload for each candidate word in the vocabulary during the autoregressive text generation process of the large language model, based on the key, the current context and candidate words. The state maintenance module is used to maintain the global state vector of the feature vectors of the generated lexical units; The trajectory calculation module is used to calculate the target trajectory vector based on the currently generated number of steps and the target message vector converted from the multi-bit binary message payload; The error calculation module is used to calculate the error vector based on the difference between the target trajectory vector and the global state vector. The feedback modulation module is used to calculate the feedback score of the candidate word based on the inner product of the feature vector of the candidate word and the error vector, and to modulate the probability distribution of the candidate word based on the feedback score. The sampling update module is used to sample the current word element according to the modulated probability distribution, and update the global state vector with the feature vector of the current word element to iteratively execute the subsequent generation steps until the complete watermarked text is generated. The extraction and decoding module is used to recalculate the feature vectors of each word in the watermarked text using the same key and perform linear accumulation during watermark extraction, and recover the multi-bit binary message payload based on the sign of the accumulated values of each dimension.
[0038] It should be understood that the closed-loop error feedback multi-bit watermark embedding and extraction system for short text generation of large language models provided in the embodiments of the present invention has all the advantages of the methods provided in the above embodiments.
[0039] Example 2 This embodiment demonstrates the specific implementation process in a short text high-load scenario, using Meta-Llama-3.1-8B as the base model, with a temperature of 1.0, top-50 sampling, and 200 prompt words. The specific process is as follows: (1) System initialization and parameter configuration: Payload length L: configured as 8, 16, 32, or 64 bits; Generate length N: 50 words and 100 words; Trajectory slope α: set to 0.4; Initial watermark strength 0.8; Attenuation coefficient γ: 0.99; Minimum watermark strength : 0.1; (2) Payload message encoding: Binary payload message Converted to bipolar representation .
[0040] (3) Closed-loop feedback watermark embedding: Sample cue prefixes from the C4 dataset and perform the embedding process according to the encoding flow described in the invention. (3.1) Initialize the state vector It is a vector of all zeros; (3.2) In each generation step t: a. The basic logits of candidate lexical units output by the large language model; b. Calculate the target trajectory ; c. Calculate the error vector ; d. Generate feature vectors for each candidate noun. ; e. Calculate the score for each word. ; f. Modulating logits: ; g. Sample words from the Softmax distribution; h. Update status ; (4) Watermark message extraction: (4.1) Initialize the integration vector All zeros; (4.2) Extract feature vectors for each word and accumulate them; ; (4.3) Determine the values of the extracted multi-bit message payload based on the signs of each dimension of the integral vector: ; (4.4) Convert to binary payload: ; (5) Performance evaluation: like Figure 2 As shown, in a 50-word short text scenario, this embodiment achieves the highest bit extraction accuracy compared to existing baselines for different message payloads. Furthermore, as... Figure 3 As shown, when performing message payload extraction on 50-word text, the watermark extraction end of this embodiment only performs linear accumulation, and the average extraction time can be kept in the millisecond range, which is better than the existing baseline.
[0041] Implementable robustness verification against attacks: To verify the robustness of this embodiment under text perturbation, both semantic parsing attacks and random substitution attacks were conducted under the condition of 100-word embedding with 8 bits. In the semantic parsing attack, the DIPPER parser was used to rewrite the 100-word generated text, where (lexical diversity, order diversity) represent the attack strength. The lexical diversity parameter controls the word substitution strength, while order diversity refers to the degree of syntactic and word order rearrangement during the rewrite. In the random substitution attack, the 100-word text was also substituted, with attack strengths of 0.1, 0.2, and 0.3 representing the proportion of words replaced in the generated text.
[0042] like Figure 4 As shown, under the most intense parsing attack, the present invention achieves a bit extraction accuracy of 75.87%, which is superior to the performance of existing baselines. Meanwhile, as... Figure 5 As shown, with a replacement rate of 30%, the accuracy of the present invention at 80.47% is also superior to the existing baseline.
[0043] If feasible, conduct ablation experiments to verify: This embodiment verifies the contribution of each core component through ablation experiments, testing was conducted with 50 and 100 term lengths and 32 and 64-bit message payload configurations. Without a global state, the controller uses the target direction but does not accumulate implemented feature states. Without a feedback mechanism, the modulation direction is fixed to the payload vector M, and the score is obtained without using real-time error. The corresponding calculations were performed under the given conditions, and the specific results are shown in Table 1 below: Table 1 Experimental results show that: (1) After removing the feedback mechanism, the extraction accuracy collapsed to near random levels, verifying that dynamic error feedback is the main source of performance improvement.
[0044] (2) The performance of the variant without global state tracking is between that of the full model and the model without feedback mechanism, indicating that the trajectory guidance itself still provides weak directional priors, but lacks accurate state accumulation and cannot perform accurate correction.
[0045] In summary, the technical solutions proposed in this invention have the following advantages compared with the prior art: (1) High accuracy in short text extraction: This invention uses a closed-loop error feedback mechanism to enable each generated lexical unit to contribute to all payload dimensions simultaneously, increasing the nominal observation opportunity per bit from T / L to T. It achieves a bit extraction accuracy of 99.69% when embedding an 8-bit message payload with 50 lexical units, and maintains an accuracy of 72.84% even in the extreme scenario of a 64-bit high-load 50-lexical unit, outperforming existing baseline methods.
[0046] (2) Strong adaptive error correction capability: This invention can correct past errors and effectively alleviate error propagation by tracking the global embedding state in real time and dynamically adjusting the word selection according to the error vector. Ablation experiments show that after removing the feedback mechanism, the extraction accuracy drops to a near-random level, with only 51.46% bit extraction accuracy at 50 words, verifying that dynamic error feedback is the main source of performance improvement.
[0047] (3) Robustness advantage: The linear accumulator decoder of the present invention can resist moderate word-level perturbations. Under the DIPPER paraphrase attack (40,40) strength, the 8-bit payload still maintains an extraction accuracy of 75.87%; under the 30% random substitution attack, the 8-bit payload maintains an accuracy of 80.47%, both of which are better than the baseline.
[0048] (4) The decoding is lightweight and efficient. The decoding process only requires linear accumulation of word feature vectors, without iterative search or complex post-processing. The computational complexity is O(N), and it has good scalability.
[0049] On the other hand, this embodiment also provides an electronic device, including a memory, a processor, and a computing program stored in the memory and executable on the processor, wherein the processor implements the method when executing the computing program.
[0050] On the other hand, this embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method.
[0051] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A closed-loop error feedback multi-bit watermark embedding and extraction method for short text generation of large language models, characterized in that, Includes the following steps: Obtain the multi-bit binary message payload to be embedded, where the message payload is used to characterize the source information of the generated text; In the process of autoregressive text generation by a large language model, a feature vector with the same dimension as the message payload is generated for each candidate word in the vocabulary, based on the key, the current context and candidate words. Maintain the global state vector of the feature vectors of the generated words; Calculate the target trajectory vector based on the currently generated number of steps and the target message vector obtained by converting the multi-bit binary message payload; Calculate the error vector based on the difference between the target trajectory vector and the global state vector; The feedback score of the candidate word is calculated based on the inner product of the feature vector and the error vector, and the probability distribution of the candidate word is modulated based on the feedback score. The current word is sampled based on the modulated probability distribution, and the global state vector is updated with the feature vector of the current word to iteratively execute subsequent generation steps until a complete watermarked text is generated. During watermark extraction, the feature vectors of each word in the watermarked text are recalculated using the same key and linearly accumulated. The multi-bit binary message payload is then recovered based on the sign of the accumulated values in each dimension.
2. The method according to claim 1, characterized in that, Obtaining the multi-bit binary message payload to be embedded specifically includes converting one or more of the following into a binary sequence: user identifier, session identifier, timestamp, model version number, service instance identifier, or content authorization information.
3. The method according to claim 1, characterized in that, The calculation of the target trajectory vector specifically includes multiplying the preset trajectory slope parameter, the number of steps currently generated, and the target message vector to obtain the target trajectory vector.
4. The method according to claim 1, characterized in that, The specific steps for calculating the feedback score of candidate lexical units include: performing a dot product operation between the feature vector of the candidate lexical unit and the error vector, then dividing by the dimension of the message payload, and using the resulting quotient as the feedback score.
5. The method according to claim 1, characterized in that, The specific steps for modulating the probability distribution of candidate words based on feedback scores include: multiplying the feedback score of each candidate word by the watermark strength parameter, adding it to the original logits to obtain the modulated logits, and then inputting the modulated logits into the normalized exponential function to obtain the modulated probability distribution.
6. The method according to claim 5, characterized in that, The watermark strength parameter adopts a dynamic scheduling strategy: when the number of generation steps is less than a preset threshold, an initial strength value is used; when the number of generation steps is greater than or equal to the preset threshold, a gradually decreasing strength value is used until the preset minimum strength value is reached.
7. The method according to claim 1, characterized in that, The recovery of the multi-bit binary message payload based on the sign of the accumulated value of each dimension specifically includes: for each dimension, if the accumulated value is positive, it is determined to be positive 1, and if the accumulated value is negative, it is determined to be negative 1, thus obtaining a bipolar message, and then converting the bipolar message back into a binary message payload.
8. The method according to claim 1, characterized in that, The linear accumulation process during watermark extraction specifically includes: initializing the integral vector to a zero vector, extracting feature vectors for each word and accumulating them into the integral vector to obtain the accumulated value for each dimension.
9. A closed-loop error feedback multi-bit watermark embedding and extraction system for short text generation of large language models, characterized in that, For implementing the method according to any one of claims 1-8, comprising: The payload acquisition module is used to acquire the multi-bit binary message payload to be embedded, wherein the message payload is used to characterize the source information of the generated text; The feature generation module is used to generate a feature vector with the same dimension as the message payload for each candidate word in the vocabulary during the autoregressive text generation process of the large language model, based on the key, the current context and candidate words. The state maintenance module is used to maintain the global state vector of the feature vectors of the generated lexical units; The trajectory calculation module is used to calculate the target trajectory vector based on the currently generated number of steps and the target message vector converted from the multi-bit binary message payload; The error calculation module is used to calculate the error vector based on the difference between the target trajectory vector and the global state vector; The feedback modulation module is used to calculate the feedback score of the candidate word based on the inner product of the feature vector of the candidate word and the error vector, and to modulate the probability distribution of the candidate word based on the feedback score. The sampling update module is used to sample the current word element according to the modulated probability distribution, and update the global state vector with the feature vector of the current word element to iteratively execute the subsequent generation steps until the complete watermarked text is generated. The extraction and decoding module is used to recalculate the feature vectors of each word in the watermarked text using the same key and perform linear accumulation during watermark extraction, and recover the multi-bit binary message payload based on the sign of the accumulated values of each dimension.
10. An electronic device comprising a memory, a processor, and a computing program stored in the memory and executable on the processor, characterized in that, When the processor executes the computing program, it implements the method of any one of claims 1-8.