Watermarking for large language models

WO2026183100A1PCT designated stage Publication Date: 2026-09-03GEORGE MASON UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/016434
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-26
Filing Date
2026-02-24
Publication Date
2026-09-03

Smart Images

  • Figure US2026016434_03092026_PF_FP_ABST
    Figure US2026016434_03092026_PF_FP_ABST
Patent Text Reader

Abstract

Watermarking of AI-generated text, and detection thereof. Watermarking algorithms for Large Language Models (LLMs) are presented that achieve watermarking with random initialization as well as achieving distortion-free embedding of multiple information bits into watermarks. Such watermarking does not compromise original functionality or quality in a watermarking structure and methodology that achieve multi-bit distortion-free watermarking with efficient information decoding.
Need to check novelty before this filing date? Find Prior Art

Description

Docket: GMU-24-035WATERMARKING FOR LARGE LANGUAGE MODELSRELATED APPLICATIONS

[0001] This application claims the benefit of provisional application serial number 63 / 763,671 filed February 26, 2025, titled “Multi-Bit Distortion-Free Watermarking for Large Language Models,” the entire content of which is hereby incorporated by reference.BACKGROUND

[0002] The emergence of large language models (LLMs), such as ChatGPT, Chat GPT-3, Chat GPT-4, Claude 3, Gemini, Grok, Mistral 7B / Mixtral 8x7B, or the like, has ushered in the dawn of general artificial intelligence due to their remarkable capabilities in understanding and generating human-like text. LLMs can be fine-tuned for a wide range of natural language processing tasks, such as text summarization, translation, question-answering, and more. Their versatility allows for the development of more generalized Al systems. The natural language generation capabilities of these models contribute to more natural and human-like interactions between machines and humans, making Al systems more accessible and user-friendly.

[0003] Unfortunately, along with these attractive features, LLMs also create opportunities for malicious use, such as mis-information spreading, inappropriate content generation, unethical activities engagements, academic cheating, etc. To bolster Al account-ability, it is crucial to be able to ascertain the provenance of a given text, i.e., whether the text was generated by an LLM or crafted by a human, and if the text was generated by an LLM, which LLM was used.

[0004] Post-hoc detectors, which were proposed as an initial approach to address this concern, are based on the use of statistical outliers or training a binary classifier over the human-generated and LLM-generated texts. However, these methods tend to be rendered ineffective as LLM-generated texts have become increasingly similar to the human-generated texts with advances in LLMs.

[0005] Another approach to Al accountability is to embed a watermark in a generated text. Watermarks are embedded by means of intentional and hopefully imperceptible modifications to an LLM. An issue with recent watermarking schemes for LLMs is that the distribution of the generated watermarked text deviates from that of the original LLM distribution, which results in distortion of the text quality.Docket: GMU-24-035

[0006] A good watermarking scheme should be distortion-free with watermarked text having the same output distribution as the original LLM; have a low probability of false alarm so that human-generated text should be detected as Al-generated with negligible probability; and have a high probability of correct detection such that Al-generated text should have a high probability of detection.BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The accompanying drawings provide visual representations which will be used to describe various representative embodiments more fully and can be used by those skilled in the art to understand better the representative embodiments disclosed and their inherent advantages. In these drawings, like reference numerals identify corresponding or analogous elements.

[0008] FIG. 1 illustrates a watermarking mapping rule.

[0009] FIG. 2 illustrates a watermarking mapping rule.

[0010] FIG. 3 illustrates a multi-bit watermarking mapping rule in a distribution interval shift coding (DISC), in accordance with embodiments of the present disclosure.

[0011] FIG. 4 illustrates the score function of a DISC algorithm, in accordance with embodiments of the present disclosure

[0012] FIG. 5 illustrates a block diagram of an encoder, in accordance with embodiments of the present disclosure.

[0013] FIG. 6 illustrates a block diagram of a decoder, in accordance with embodiments of the present disclosure.

[0014] FIGs. 7-9 illustrates example flowcharts of encoding and detecting watermarks for Al-generated text, in accordance with embodiments of the present disclosure.DETAILED DESCRIPTION

[0015] New watermarking algorithms for Large Language Models (LLMs) are presented herein with a primary focus on achieving watermarking with random initialization as well as achieving distortion-free embedding of multiple information bits into the watermark. A key advantage lies in providing embedding power without compromising original functionality or quality in a watermarking structure and methodology that is the first to achieve multi-bit distortion-free watermarking with efficient information decoding. This advancement opens up avenues for many applications in content authentication and communication security.Docket: GMU-24-035

[0016] With regard to generating and detecting a multiple-bit watermark in Al-generated text, consider the following encoder, decoder, and methods of encoding and detecting.Therefore, in accordance an encoder is operable to create a watermark indicative of AI-generated text, the encoder encodes multiple bits of meta-information in the watermark in accordance with a multi-bit distortion-free watermark mapping rule, the watermarked AI-generated text having the same distribution as the Al-generated text. Further in a related method of encoding meta-data in watermarks of Al-generated text, when generating text by artificial intelligence (Al) the encoder encoding multiple bits of meta-information in a watermark of the Al-generated text in accordance with a multi-bit distortion-free watermark mapping rule, thereby generating a watermarked Al-generated text, the watermarked Al-generated text having the same distribution as the Al-generated text. A decoder is provided that, responsive to receiving text and a secret key shared with an encoder, the decoder operable to determine whether the received text is watermarked and thus watermarked Al-generated text. A watermark of a determined watermarked Al-generated text in accordance with a multi-bit distortion-free watermark mapping rule based on a distribution interval shift coding (DISC) of the multi-bit distortion-free watermark mapping rule. A method of decoding a watermark of Al-generated text includes, responsive to receiving text and a secret key shared with an encoder, a decoder determining whether the received text is watermarked and thus watermarked Al-generated text, a watermark of a determined watermarked Al-generated text in accordance with a multi-bit distortion-free watermark mapping rule based on a distribution interval shift coding (DISC) of the multi-bit distortion-free watermark mapping rule. Further, either such encoder or a decoder, or related methods, need only consider one time any ngram context plus current token in accordance with a shifted score function.With regard to generating and detecting watermarking with random initialization, consider the following encoder, decoder, and methods of encoding and detecting. Therefore, an encoder is operable to create a watermark indicative of Al-generated text in accordance with a distortion-free watermark mapping rule with random initialization to generate a watermarked Al-generated text. The encoder responsive to a prompt and in accordance with the distortion-free watermark mapping rule generates a sequence of uniform random numbers and the watermarked Al-generated text based upon an accumulated empirical entropy of a sampled set of tokens of an artificial language model that generated the Al-generated text exceeding a threshold. An initial random chunk of tokens and length thereof sampled from the artificial language model is not dependent on the output of a pseudorandom function.Docket: GMU-24-035Further in a related method, when generating text by artificial intelligence (Al) an encoder generating a watermark of the Al-generated text in accordance with a distortion-free watermark mapping rule with random initialization and thereby generating a watermarked Al-generated text. Responsive to a prompt and in accordance with the distortion-free watermark mapping rule the encoder generating a sequence of uniform random numbers and the watermarked Al-generated text based upon an accumulated empirical entropy of a sampled set of tokens of an artificial language model that generated the Al-generated text exceeding a threshold, in which an initial random chunk of tokens and length thereof sampled from the artificial language model is not dependent on the output of a pseudorandom function. A decoder is provided that, responsive to receiving text and a secret key shared with an encoder, is operable to determine whether the received text is watermarked and thus watermarked Al-generated text. In accordance with a shifted score function the decoder is operable to only considers one time any ngram context plus current token and where the decoder is operable to compare a probability value to a threshold false positive rate in determining whether the received text is watermarked. A method of decoding a watermark of Al-generated text includes a decoder, responsive to receiving text and a secret key shared with an encoder, determining whether the received text is watermarked and thus watermarked Al-generated text. Further, either such encoder or a decoder, or related methods, need only consider one time any ngram context plus current token in accordance with a shifted score function.

[0017] Random variables are denoted by boldfaced letters, e.g., w, whereas, specific realizations of a random variable are denoted by non-boldfaced letters, e.g., w. In this case, w = w means that the random variable w gets value w. For brevity, in conditional probability expressions, we abbreviate expressions such as w = w by simply w. We use uppercase letters to denote vectors and calligraphic font to denote sets. We use: = (%j, xi+1, Xj) and X)n]: = (xlt, xn) to denote subsequences of a sequence {xn}.

[0018] Definition 1 (Language Model). A vocabulary V = {vltv2,..., V\V\] is a set of tokens. Given a generated sequence of tokens VF[t-i] 6 Vt-1and a prompt a =VF[_(-Wp— i):o] 6 VNp, where Npis the length of the prompt, a Language Model M is specified by a conditional distributionPt.i = PM(G I ^[t-i]<«)fi (1): = PMtwt = G1Docket: GMU-24-035where i = 1, | V | for the tthtoken wtG V. We use Dt= (pt,i< Pt,2< ■■■ > Pt,|v|) to denote this conditional probability distribution over the set V and define Dt[Vj] = ptj, for j = 1, |V|.

[0019] Note that wtis a nonstationary discrete-time random process with conditional probability distribution given by (1). The response of a language model M to a prompt a is a random vector M(a): = W[L], where L is a random variable, with conditional distribution given byP{M(a) = Ww} = nf=i PM (™t I W-ifa)=PM( W-i] Ia) ’ PM( W I W-i]’a)-{=r’2’ - -

[0020] where PM(W[O] Ia) = 1, and pM(-| kk[o]<a)=PM(’IA)- This random response is generated by sampling from Dtuntil a special terminating token done 6 V is generated. Therefore, M(a) = W[L] implies that W[L] = done. When we talk about altering a language model, we mean Dt-> £)t', where the prompt a and the sequence generated tokens so far, i.e., are fixed. The entropy of the response of a language model M to a prompt a is defined asH(tz):= EM(a){-lnP{M(tz)}} (3)

[0021] Having established a language model definition, we now define a distortion-free watermarking algorithm.

[0022] Definition 2. A watermarking algorithm is distortion-free if for any prompt a and a sequence of watermarked generated textwe haveP{wt= wtI= pM(wtI iy[t-1];a) = £>t[wt],for all tE {1,

[0023] In other words, a watermarking algorithm is distortion-free if the watermarked text has the same distribution Dtas the non-watermarked text.

[0024] Building upon Christ, M., Gunn, S., and Zamir, O. Undetectable Watermarks for Language Models. 2023. URL http: / / arxiv.org / abs / 2306.09194, hereinafter Christ et al. 2023, the watermarking algorithms disclosed herein use pseudorandom functions (PRFs), such as hashes, to generate random numbers. PRFs are defined as follows,

[0025] Definition s (PseudoRandom Functions (PRF)). Let F = {Fsk: {0,l}Z1G) -> {O.ljFG) |sk 6 {0,l}A} be a family of functions. F is a PRF if Fskis efficiently computable and for all probabilistic polynomial-time di stingui shers D,Docket: GMU-24-035[00261 IP^(».i)4c"‘k(W) = 1}& - = 1}|< negl(A),where / : {0, 1}Z1-> {0, 1}*2denotes a random function. A function ^(A) is negligible, denoted by negl (A), if #(A) e 0 for every poly(-).

[0027] Distortion-free watermarking embeds the watermark in the correlation between {wt]=1and a specially crafted i.i.d. sequence {yt} =i, where is the length of the generated text. yt~ Uniform [0,1]. In the watermarking process, the token wtis generated such that it has high correlation with yt, for t = 1,..., £. On the other hand, for the human-generated text or non-watermarked text, the token wtis independent of yt, for t = 1,..., £. For a given sequence of tokens VF = VF[^ and a sequence of generated random numbers Y = Y^, the detection algorithm is a statistical testY) that checks whether the correlation between the two sequences exceeds a threshold.

[0028] In the watermarking process, the correlation between {wt} and {yt} is introduced using a watermarking mapping rule.

[0029] Definition 4 (Watermarking mapping rule). For a random variable y defined over the sample space fl with distribution Py, a watermarking mapping rule F(fl, V; Py) is a process of mapping a partition of a sample space fl into the token set V of the language model M. This process is done in two steps: First, a partition U =... |y|] on the sample space fl is formed. Then each partition part. Ay is mapped to r, for j = 1,..., |V|.

[0030] Watermarking mapping rule T (fl, V).

[0031] The distortion-free property of a watermarking algorithm following a watermarking mapping rule Ft(fl, V; Py) is characterized as follows.

[0032] Proposition 5. A watermarking algorithm following a watermarking mapping rule rt(fl, V; Py) is distortion-free if and only if for every prompt a and the past generated tokens

[0033] P{yte Ajit) = pM(vj I VFft-!], a) = Dt[vj]. (4)

[0034] ZERO-BIT DISTORTION-FREE WATERMARKING

[0035] To achieve the requirements that a good watermarking scheme should be distortion- free, have a low probability of false alarms and a high probability of correct detection, watermarking may rely on binarization of the token set and token generation based on the value of a pseudorandom function (PRF). Zero-bit watermarking schemes discussed hereinDocket: GMU-24-035reference a watermark that does not embed any information beyond differentiating an AI-generated text from a human-generated one.

[0036] Binarization of language model

[0037] Watermarking follows a watermarking mapping rule and is done on binary tokens. Language model M with token set V is converted into a language modelwith binary token set Vb= {0,1}. This conversion is done as follows: First, each token v G V is represented as a distinct binary string in {0,l}loglvl, where log(-) denotes the base-2 logarithm. In this way, sampling log|V | times from the language modelis equivalent to sampling one token from language model M. Henceforth, we will assume binarization has been applied.

[0038] Watermarking without random initialization

[0039] The watermarking algorithm decides the value for each binary token according to a watermarking mapping rule, in accordance with Christ et al. 2023, as shown in FIG. 2.Given a prompt aband the past generated tokensthe watermarking mapping rule Fj(£l, Vb\ Py) is specified as follows:£1 = [0,1], yi ~ Uniform[0,l],

[0040] nCH) = {Au= [Pi(l),l], A2,i = [O,pf(l))}, (5)HOM = o, = i,where pj(l) is the probability of the i-th token being 1 according to M6. According to (5), P{u G A1;J = Pi (0), P{jz, G A2, J = pi(l).

[0041] By Proposition 5, a watermarking scheme following this watermarking mapping rule is distortion-free. The random variablecan be generated using a PRFFsk: {0,l}polyiW -> {0,l}poly2«, with a secret key sk G {0,l}A, shared between the encoder and the decoder. Here, is the security parameter of the watermarking algorithm.

[0042] Unlike previous approaches, index i is not used as the input to this PRF, as this choice of input will make the watermark very prone to a simple deletion attack, and removing one word in a watermarked text can render the detection process useless. Instead we use the context ngram Si h=as the input to the PRF for the i-th token. Usually h is chosen such that [ / i / log|V| ] < 8. As an example, set [ / i / log|V| ] = 5. Here, polyxis chosen such that the ngram Si his not too long for the PRF, and if it is too short it will be padded. On the other hand, let z be the integer representation of the output of PRF; then taking2poly2(Z)results in a real number in [0,1]. In our notation, we assume these steps are included and Fsk(i) G [0,1].Docket: GMU-24-035The watermarking encoding scheme can be summarized as follows:v_ p (c Awb - i1’ 0 < yi< pi(.l\yi ~F^Si-h)’wi - (o, Pi(l) < yt < 1.(6)For Wb= and Y = Y^, the decoder uses a statistical test ip(Wb, F) to check whether the correlation between Y and Wbexceeds a certain threshold. The null hypothesis is dCy. text is non-watermarked while the alternative hypothesis istext is watermarked. First, for each binary token a score value is calculated. This score value depends on the binary token value, wb, and the uniform random number generated, y. Define C(Wb, F) as the sum of the score values for all (Wi >yi), i = 1,(?)

[0043] The p-value for the observed value z of a random variable z is defined as the probability of observing a value at least as extreme as the observed value z under0, i.e., p-value(z) = P{ z > z | dC0}. (8)

[0044] The p-value for C(Wb, F) is compared to a threshold FPR, where FPR is the maximum tolerable false positive rate for detecting a non-watermarked text as watermarked. If p-value(C(VFi,<F)) < FPR, the text is detected as watermarked. This leads to a critical region Dc= {C(Wb, F) > 0], such that J-Cois rejected if C(Wb, F) G Dc, and a region of acceptance = {C(Wb, F) < 0], where J-Cois accepted if C(Wb, F)G. In other words, for a non-watermarked text l / F^wand the constructed Y. the threshold 0 is chosen such that, P{C(VF^W, F) > 0} < FPR, for all l / F^wwith length ■£. This choice of threshold 0, on the other hand, affects the false negative rate. If in addition to FPR, there is also a maximum tolerable false negative rate FNR, this bound will yield a minimum length of watermarked text such that both false positive rate and false negative rate are bounded by FNR and FPR, respectively.

[0045] The score function is calculated as follows,fin—, w? = 1,

[0046] s(w / ,,y() = l,yi „ (’)In -, w,- = 0.I l-y / 1

[0047] This score function is designed such that, given a prompt aband the past generated tokens the expected value of the score function for the i -th token, if it is watermarked, is greater than its expected value if it is non-watermarked.

[0048] Using the central limit theorem (CLT), we can approximate the distribution of C(l / Fw- ^)- There can be correlation between the s(wb,yiys. This correlation can stem fromDocket: GMU-24-035the correlation between the tokens generated by an LLM or from similar ngrams with length h appearing throughout the text. Not much can be done for the first cause of correlation, however, by adopting the idea from (Fernandez et al. 2023), in the detection process we can eliminate the second cause. Namely, we consider any ngram “context plus current token,”i.e., only once. In other words, if for different values of i, the corresponding context plus current token ngram is repeated before, we will not include its corresponding score value in C(Wb, Y Therefore, to simplify our derivation, in applying the CLT we shall assume that the s(w,yj)’s are independent. Our experimental studies have shown that this assumption does not adversely affect the detection algorithm.

[0049] Either the encoder and decoder can check for repeated ngrams; if the encoder checks for repeated ngrams, the decoder will not check and vice-versa. If a unit in the encoder searches for repeated ngrams, it looks at the last h tokens. If they are already repeated previously in the text, then the next token will not be watermarked and will be generated according to the language model itself.

[0050] In the decoder, the unit checks whether the last h tokens plus the current token is already repeated. If it is not repeated, then the score value for the current token will be added to the accumulated score to be compared with the threshold; otherwise, if the ngram is repeated the score of current token is ignored.

[0051] Consider in a simplified example that a user instructs the LLM, “Write a story about Alice.” The prompt is “Write a Story about Alice” and the context may be “about Alice.”, in the example where h = 4, and the current token might be “girl.” An example probability distribution for what the next word might be may include:Alice 0.5She 0.1One 0.2School 0.0001In this example probability distribution, there is a 50% chance that the next generated word by the LLM is “Alice”, a 10% probability that the next generated word by the LLM is “She”, a 20% probability that the next generated word by the LLM is “One”, but only a 0.01% probability that the next generated word by the LLM is “School.”Docket: GMU-24-035

[0052] The threshold 6 and the critical region Dcare chosen such that for a nonwatermarked text IVNWand the constructed Y, P{C(I / Fr)’w, Y) > 6} < FPR. Using the CLT approximation for a non-watermarked text l / F^wand the constructed Y as,However, for non-watermarked token i, s(w,yi) ~ Exp(l). Hence, using the assumption of independence of ^(w^y^’s, we can obtain a more accurate result as CCC LCI).where Erk(l) denotes the Erlang distribution with shape parameter k and rate.

[0053] Therefore, for a given L and FPR, we can derive 6 as follows,0= ^(DCFPR) = Q-1(L, FPR), (10)where FErL^(x') is the tail distribution function of a random variable x ~ ErL(l) and ErL( ) denotes the Erlang distribution with shape parameter k and rate:FErL(i)(%) = P{x > x] = = Q(L,x), (H)where F(L, x) is the upper incomplete gamma function and Q(L, x) is the regularized gamma function.

[0054] We have derived an approximation for Lmin: / 2(FPR, FNR)min~ ’where1 1 1 1 / i(FPR, FNR) = 21n— — + 4.41n— — (13)\ 2FPR AJ 2FNRv’

[0055] The approximation in (12) is similar to the lower bound approximation, i.e.,1 1O( In-) with FPR=FNR= J], We can express Lminin (12) in terms of the average s )conditional entropy, conditioned on the past tokens, per token of response of the language model M to the prompt a, i.e., ('(a) as| | / 2(FPR, FNR) >Sl 1<(a)2log|V|V’

[0056] Based on (14), the number of required tokens to achieve a desired probability of error for a watermarking scheme, on a binarized language modelis log|V | times than what it would have been under ML The exact and approximate values of Lminderived by the numerical method mentioned here and by using (12), for different values of £(0^) for vocabulary size |V| = 50272. The bounds on false positive rate and false negative rate areDocket: GMU-24-035considered equal. As the average conditional entropy per token for the generated text increases, fewer tokens are required to achieve the same false negative and positive rates.

[0057] Watermarking with random initialization

[0058] In order to address the problem of always generating the same response for a prompt a by the watermarking encoding in (6), Christ et al. 2023 proposed to initiate the watermarked text with a chunk of tokens R that is randomly sampled from language model M6. The empirical entropy for R is defined asHe(Mb, ab, R~) = (15)

[0059] After sampling tokens from language modeluntil He(b, ab, R) exceeds a threshold A. for the sampled set of tokens R = wbm^ the watermarking encoding starts. The sampled binary tokens in R and the context ngram Si h=are used as the input to a PRF Fsk: {0,l}polyi(71 -> {0,l}poly2^, with a secret key sk 6 {0,l}A, to generate a random variable Kj. Following the same steps as previous section, the watermarking encoding can be summarized using (6), with yt= Fsk(F,5jft). We can impose a minimum length h for R. Therefore, following this procedure for a prompt ab, a sequence of uniform random numbers, pi+i: L]andawatermarked text ( / ?, Wbn+1. Lj) will be generated, where the initial chunk of tokens, R, is sampled from the language modeland the rest of the tokens are decided based on the watermarking encoding rule in (6) with y;- = Fsk(F, Si h). We then obtain the encoder in Algorithm 1, below:Algorithm 1 Zero-bit encoder with random initializationInput: A prompt a and a secret key skOutput: Watermarked text WbL^1: t <- 1; H <- 0;2: while done wbt-1 do3: Pr(l) PM»(1 I Pr(0) 1 - pt(l)4: if 11 / . then5: Sample wbwith (pt(0),pt(l));6: H <- H — lnpt(w^)7: if H > k and t > h then8: R - 9: endif10: else11: Establish mapping ruleyt«- Fsk(F,5t ft);{Eq. (5)}12: yt<- Fsk( / ?,5t,ft); wt<- l[yte A2,t];13: end if14: t <- t + 1;Docket: GMU-24-03515: end while

[0060] The initial random chunk R, including its length, is random but does not depend on the output of the PRF. If the accumulated empirical entropy never surpasses, then the text generated by Algorithm 1 will contain no watermark. For example, when the answer to the prompt is a deterministic answer, such a case happens. Given a prompt aband a fixed initial chunk R, the conditional distribution for response of a language modelcan be derived from (2), with M and a replaced byand (ab, R), respectively. In other words, (ab, R) is treated as the prompt.

[0061] As in the previous section, we can assume the decoder has access to the reconstructed sequence F[n+1: L], For a given text Wb, a statistical test ip(Wb, F) is used to test hypothesis0against J-C1. First, an initial chunk of Wbwith length m, is considered as / ? = Then for that specific initial chunk R, Y(R) = F[m+1: L] is constructed and we defineC(Wb, Y;m): = Zf=m+iS (16)

[0062] Again, in calculating C(Wb, Y; tn), any ngram “context plus current token” is considered only once. The estimated n in the decoder is defined asn*\ = argminft<m<L-i{p-valuem(C(VFi,<Y; tn))}= min Q(L — m, C(Wb, Y; tn)).h<m< L-l

[0063] Then the global p-value(C(VFi,<F)), defined as,p-value(C(VFi), F)): = l -(l - p-valuen.(C(PFft, F;n*)))L-\1 Jis calculated. If the global p-value(C(VFi,<F)) < FPR, the text is detected as watermarked, where similar to the previous section, FPR is the maximum tolerable false positive rate. We then obtain the watermark detector given in Algorithm 2, below:Algorithm 2 Zero-bit detector with random initializationInput: Text WbL^ and a secret key skOutput: true or false;1: for m <— h,..., £ - 1 do2: < Am= 0; C(WbL], Y,-m) ^ Q, R ^ Wbm]3: for i «— m + 1,..., £ do4: if £ c / Zmthen5: Add VFp_h:q to yt<— Fsk(R,6: Vi wb• yt+ (1 - wf ) • (1 - yf)7: s(wb, yf) InDocket: GMU-24-0358: C(^?Y;m) - C(^, Y;m) + s(w?, yt)9: endif 10: end forll: pm- Q(|^m|, C(<],, Y;rn))12: end for13: n* argminh^L-i pm{Eq. (17)}14:if 1 - (1 - pn* fr~h< FPR then15: return true Else return false16: end if

[0064] As in the previous section, we can derive a lower bound Lminon the number of required watermarked tokens such that the false positive rate and false negative rate are bounded by FPR and FNR, respectively. At first we start with an estimate of Lminas Lmin= Lo, for a small value Lo. Then for this estimate Lmin= Lo, we derive ft and 6nand also derive an upper bound for false negative rate. If the derived upper bound on false negative rate is less than FNR, the estimated value of Lmin, is correct. However, if the derived upper bound on false negative rate is greater than FNR, we increase the estimate Lmin, i.e Lmin<-Lmin+ 1, and repeat this process until for Lmin= L*, the upper bound is bounded by FNR, and we derive Lmin= log| V\ [y^].

[0065] MULTI-BIT DISTORTION-FREE WATERMARKING

[0066] Methods for watermarking LLMs may distinguish Al-generated text from humangenerated text by slightly altering the model output distribution, but they also distort the quality of the text, exposing the watermark to adversarial detection. Distortion-free watermarking methods may require a secret key to detect the watermark. These approaches generally embed zero-bit water-marks that do not provide additional information beyond tagging a text as being Al-generated. As previously mentioned, zero-bit watermarking schemes reference a watermark that does not embed any information beyond differentiating an Al-generated text from a human-generated one.

[0067] A zero-bit distortion-free water-marking method is extended and improved by embedding multiple bits of meta-information as part of the watermark. We also develop a computationally efficient decoder that extracts the embedded information from the watermark with low bit error rate (BER).

[0068] In practice, it is crucial to encode meta-information such as the language model name, model version, and generation time within a watermark. For example, encoding metainformation in the watermark supports forensic analysis in case of mis-use. It helps trace back the origin of content and assists in determining whether a specific model or versionDocket: GMU-24-035was in-volved. Moreover, the incorporation of meta-information in the watermark aligns with the need for responsible and accountable Al practices.

[0069] Therefore, watermarking is extended to include the following additional properties:• Multi-bit embedding: The watermark encodes multiple bits of meta-information. • High probability of correct decoding: The bit error rate (BER) in decoding the embedded information bits should be low.• Efficient decoding: The decoding algorithm should be efficient and not require exhaustive search over all the possible embedded information bits.

[0070] For a multi-bit distortion-free watermarking scheme, Definition 4 above is extended to a multi-bit watermarking mapping rule suitable.

[0071] Definition 6 (Multi-bit watermarking mapping rule). For a random variable y defined over the sample space £1 with distribution Py, a multi-bit watermarking mapping rule T (£1, V,; Py) maps a partition of a sample space £1 into the token set V of the language model M. This process is done in two steps: First a partition U(M)= { / ^(M),...,i4|V|(M)}, depending on the message M G JC, on the sample space £1 is formed. Then each partition part,(M) is mapped to Vj for j = 1, |V|.

[0072] The multi -bit watermarking mapping rule is depicted in FIG. 1. The mapping rules in red and black depict two distinct mapping rules based on two different embedded messages. We assume = {0,1,..., 2m— 1], such that each message M conveys m bits of information. Let F(£l, M) denote the partition on £1 depending the embedded message M, i.e. F(£1, M) = UfJTy On the other hand, F(.4j(M)), denoted the mapped token, i.e. F(.4j(m)) = Vi. Finally, F(£l, V, M) denotes the the process of partitioning the sample space based on the embedded message M and then mapping the partition parts into the corresponding tokens.

[0073] Proposition 5 is extended for a multi-bit watermarking mapping rule as follows in Proposition 7.Proposition 7. A watermarking algorithm following a multi-bit watermarking mapping rule rt(£l, V, J f; Py) is distortion-free if and only if for every prompt a and the past generated tokens and for every message M G JVC,P{ytG A;,t(M)} = pM(v;I a) = Dt[vjJ. (19)

[0074] Next, we propose a new multi-bit and distortion-free watermarking algorithm called Distribution Interval Shift Coding (DISC), which follows a multi -bit watermarking mappingDocket: GMU-24-035rule as depicted in FIG. 3. Given a prompt a and the past generated tokensthe multibit watermarking mapping rule Fj(£l, V, Af; Py) is specified as follows:£1 = [0,1], yi ~ Uniform[0,l],ri(n, M) = {^1,i(M)M2,i(M)}, (20)(M)) = 0, rfG42.f(M)) = 1.For (1) + 8M< 1,^i,i(W) = {0 < yt< <5M} U {pi(l) + 8M< yt< 1},A2,im = {5M<YI< Pi(l) + 8M}, whereas for Pf(l) + 8M> 1,^i,i(W) = {<5M+ Pi(l) - 1 < yt < <5M},= {0 < yt< 8M+ pi(l) - 1} U {8M< yt< 1},where SM= M3, M E M = {0,1,....2m— 1], and 6 = 2~m. As above, yt= FskR, Si h) E [0,1] and the DISC scheme can be shown to be distortion-free.

[0075] For a prompt aband a message M, a sequence of uniform random numbers, V[n+1:£] and a watermarked text Wb= ( / ?, VV|n+1:£]) will be generated, where the initial chunk of tokens, R is sampled from the language modeland the rest of the tokens are decided based on the watermarking encoding rule in (20) and (22) with yt= FS\ R, Si h). Therefore, we extend the encoding Algorithm 1 in the Input and lines 13 and 15 to obtain the DISC encoder given in Algorithm 3, here:Algorithm 3 DISC encoderInput: prompt a, secret key sk, message M E M13: Establish mapping rule Ft(£l, V, M);15: wt<- l[yte A2,t(M)];

[0076] Similar to above, for a given text {Wb}, the statistical test ip(Wb, K) is used to test hypothesis J~C0against J-C1. This statistical test is performed as follows. First, an initial chunk of Wbwith length m, is considered as R =Then for that specific initial chunk R, Y(R) = f[m+i:£] is constructed. Then for 6M> = M'6 and M' E JVC as the assumed message by the decoder,C(Wb, Y; m, 8MC) = Sf=m+i $ (wf, yf, 8MC), (23) is calculated. Any ngram “context + current token” is considered only once. The score function is given as follows:Docket: GMU-24-035fin — - Yi G [0,< w = 1,yt~sM,+1Yi G (< V< 1-Lwf = 1,lnyi-sM'(24)insM,-yiYi G [0,< V),wf = o,In- — - - Yi G [<5M', l],w? = 0.Reference FIG. 4, which illustrates an example score function of the DISC algorithm.

[0077] Note that for a fixed m and different values of M' E M, C(Wb, Y;m, 6M>) are correlated with each other. In fact, according to (24), for a fixed m and all different values of M' G, s(w,yi;<5Mare functions of s(wb,yi; <50). Hence, for a fixed m, by having Wb, Y and C(Wb, Y;m, <50) allC(Wb, Y; m, can be uniquely determined for all M' E M. The score function in (24) is a shifted version of a score function.

[0078] Similar to above, after calculating C (Wb, Y; m, 6MY), the estimated n and M in the decoder, i.e., n* and M”, defined as,= argminft <m< L-i{p-valuem M>(C(Wb, Y;m, <5M'))}- M'eM= min Q(L — m, C(Wb, Y; m, 8M'Y). (25)M'EJVCThen the global p-value C W6, K)), defined as,p-valuefCfW6, K))(26): =! - (! - |M|p-valuen*.M*(C(lVi,, y;n*, M*)))L-ft,is calculated. If the global p-value C W6, K)) < FPR, the text is detected as watermarked. Therefore, the detecting Algorithm 2 is extended in the Output and lines 4, 12-15, 17, 19, 20-21 to obtain the DISC decoder given in Algorithm 4, here:Algorithm 4 DISC decoderOutput: true or false and message M if text is water-marked4: C Wb, Y;m, M') 0;12: for M' E M do13: Calculate s(wb,yi 8MY) {Eq. (24)}14: C(Wb, Y;m, M') <- C(Wb, Y; m, M') + s^w^yp, 6M>);15: EndFor19: 'i'. W M'eM 20: If 1 (1 - |M[p„* ), h< FPR then 21: return true and M' Else return false EndifDocket: GMU-24-035

[0079] For a watermarked text with length L and initial chunk R, and the constructed F( / ?), generated as response to prompt ab, we can derive a lower bound on the number of required watermarked tokens, Lmin, such that the false positive rate and false negative rate are bounded by FPR and FNR, respectively.

[0080] For a watermarked text Wbvwith length L and initial chunk R =and the constructed F( / ?), generated as response to prompt ab, we can avoid the exhaustive search over all M' E JVC in (25) to find M*. For m = n, C(Wb, Y; m, A) follows a pattern, i.e.,having maximum at A = 0 and two valleys on each side. Because we know the pattern of C(Wb, Y; m, A), in order to find M*, we can simply calculate C(Wb, Y; m, 6M>) for a small set of equally spaced 8M' E JvCs, with |S| « | |. After calculating C(Wb, Y; m, 8M8) at 8M' E JVCS, because of the patterns of C(Wb, Y; m, A) are known, we can obtain a rough estimate of M* and then by a finer search we can derive the exact M*. Therefore, for a textwith length L, the complexity of decoding Algorithm 4 can be reduced to O(L2), and does not depend on the number of possible embedded information bits.

[0081] Embedding capacity of DISC

[0082] With regard to an estimate of the embedding capacity of DISC, the embedding capacity depends not only on the average conditional entropy per token for the generated text, but also on the distribution of the probability of each binary token being 1, i.e. pdf of pi(l). The embedding capacity can be fixed in advance. As mentioned in Algorithm 4 above, in a watermarked text length of R, i.e., n, and the embedded message, i.e., M, are estimated according to (25) as «*and M*. Therefore, for a watermarked text the embedded message can be detected incorrectly if either nfJn or M* J M. The case in which a watermarked text is detected as non-watermarked may be considered a false negative case.

[0083] The efficacy of DISC in embedding and extracting the watermark may be assessed by simulating the binary sequences. Consider the following example, intended for illustration and therefore not to be considered limiting to the embodiments envisioned herein, in which a real token is represented by 17 bits. For m-bits watermark, there exist 2mdistinct information options M for watermarking, i.e., M E JVC = {0,1,..., 2m— 1}. With M as the watermarking information to be conveyed, 6M= M3 should be embedded during text generation. In this context, we experiment with different values for m ranging from 1 to 4, as an example.

[0084] For the text with L bits (i.e., L I 17 real tokens), we randomly generate the probability of each bit being 1 and a corresponding random value u for that bit.Subsequently, the text is generated using the DISC encoder. In the process of watermarkDocket: GMU-24-035decoding, we investigate various 6M, values to pinpoint the one exhibiting the highest score within the text. The BER may be considered at a logarithmic scale when extracting watermarks of varying lengths across different numbers of real tokens in the text. For each length of text, the DISC algorithm in certain experimentation was executed 10,000 times to compute BER. Notably, the BER for four different m exhibits a significant decrease initially. Specifically, as a non-limiting example, the extraction of a 1 -bit watermark achieves a 0 BER over a text of merely 6 tokens, while a 4-bit watermark attains 0 BER at 20 tokens, just as an example.

[0085] The above embodiments of the disclosure may be illustrated by example encoder and decoder arrangements. Referring now to encoder block diagram 500 of FIG. 5, how watermarking may be applied to Al-generated text by an encoder is illustrated. A prompt and a context 505 are provided to a large language model (LLM) 510 which generates a probability distribution which is binarized as a binarized probability distribution at 520. A vector of binarized probability distribution is created at 530. The next chosen word at 570 is determined by the generated number Y' or Y" depending upon the value of an accumulated empirical entropy of the initial chunk of tokens R, defined by (15) above, relative to a threshold, thereby providing a switch of the encoder that when triggered switches from a deterministic (random) model 540’ to a probabilistic one 540” in which the generation of watermarked tokens begins.

[0086] When the accumulated empirical entropy is less than or equal to threshold, random number generator (RNG) 550’ of block 540’ randomly generates Y' which together with the vector of binarized probability distribution at 530 is used to select the next token at 570; this token is not watermarked. Once the accumulated empirical entropy is greater than threshold, random number generator (RNG) 550” of block 540” randomly generates Y" which together with the vector of binarized probability distribution at 530 is used to select the next token at 570; this token is watermarked. Unlike RNG 550’ of 540’, RNG 550” of 540” generates Y" depending on the context, the secret key sk G {0,l]Ashared with a decoder, and index z of the t-th binarized token. These three inputs to RNG 550” causes the tokens generated by LLM 510 to now be watermarked.

[0087] Referring now to FIG. 6, decoder block diagram 600 is illustrated. While decoder 600 doesn’t have access to the LLM and its generated probabilities it assumes an initial length h+1 in a context window as shown. The inputs 605 provided to RNG 610 include context information, the secret key shared with the encoder and the index of binary tokens.Docket: GMU-24-035The generated Yj 630 together with binarized tokens Wjb620 are input to 640 where the score function of all the tokens are summarized and compared to a threshold, as described in example equation (24), though other score functions may be used.

[0088] As used herein, PRFs such as hashes may seed the RNGs described above.

[0089] With regard to a unit that checks for the repeated ngrams, this unit of the RNG may reside in either the encoder 500 or the decoder 600. As noted above, ngram “context plus current token,” i.e.,is considered only once. If it is used in the encoder, then it will not be used in the decoder, in accordance with embodiments herein.

[0090] In the encoder unit, the unit in the RNG 550’ or 550” looks at the last h tokens. If they are already repeated (previously at some point of the text we have the same sequence of tokens) then the next token will not be watermarked and will be generated according to the language model itself. In the decoder unit, the unit of RNG 610 checks whether the last h tokens together with the current token is already repeated. If it is not repeated then the score value for the current token will be added to the accumulated score to be compared with the threshold; otherwise if the ngram is repeated the score of current token is ignored.

[0091] In decoder 600 the score function for each token for all values of possible messages are calculated and then the best probable message and the length of initial chunk is determined according to line 19 in Algorithm 4, above. Then the score value for the best probable message is compared to the threshold, thereby defining an argmin operation before comparing to the threshold.

[0092] Referring now to FIGs. 7-9, flowcharts illustrate the methodology of embodiments described herein. In flow 700 of FIG. 7, the methodology for embedding multiple bits in a watermark when an LLM generates Al-generated text is illustrated. In Block 710, multiple bits of meta-information are embedded in a watermark of the Al-generated text in accordance with a multi-bit distortion-free watermark mapping rule to generate a watermarked Al-generated text, with the watermarked Al-generated text having the same distribution as the Al-generated text. In Block 720, on the decoder side, responsive to receiving text and a secret key shared with the encoder, a decoder determines whether the received text is watermarked and thus watermarked Al-generated text, in accordance with a shifted score function, either the encoder or the decoder only considers one time any ngram context plus current token, the decoder compares a probability value to a threshold false positive rate in determining whether the received text is watermarked, indicative of Al-generated text.Docket: GMU-24-035

[0093] Referring now flowchart 800 of FIG. 8, with regard to embedding multiple bits as meta-data more specially to the encoder, at Block 810, responsive to receiving a prompt and a text message, an encoder generates a sequence of uniform random numbers and watermarked Al-generated text. An initial chunk of tokens sampled from an artificial language model that generated the Al-generated text and a plurality of tokens of the watermarked Al-generated text is determined based on a deterministic distribution interval shift coding (disc) of the multi-bit distortion-free watermark mapping rule. The multi-bit distortion-free watermark mapping rule includes:• Generating randomly the tokens of the initial chunk until an accumulated empirical entropy thereof does not exceed a threshold at Block 810; and• Generating deterministically the plurality of tokens of the watermarked Al-generated text responsive to the accumulated empirical entropy of the tokens of the initial chunk exceeding the threshold at Block 820.

[0094] Considering random initialization, consider flowchart 900 of FIG. 9 in which when generating text by artificial intelligence (Al) an encoder generates a watermark of the Al-generated text in accordance with a distortion-free watermark mapping rule with random initialization to generate a watermarked Al-generated text at block 910. Responsive to a prompt and in accordance with the distortion-free watermark mapping rule, the encoder generates a sequence of uniform random numbers and the watermarked Al-generated based upon an accumulated empirical entropy of a sampled set of tokens of an artificial language model that generated the Al-generated text exceeding a threshold. An initial random chunk of tokens and length thereof sampled from the artificial language model is not dependent on the output of a pseudorandom function. At block 920, a decoder, responsive to receiving text and a secret key shared with the encoder, determines whether the received text is watermarked. In accordance with a shifted score function, one the encoder or the decoder considers one time any ngram context plus current token. The decoder compares a probability value to a threshold false positive rate in determining whether the received text is watermarked.

[0095] The illustrated examples shown in FIGs. 7-9 sets forth computer-implemented methods 700, 800, 900 of generating and detecting watermarking for Al-generated text, in accordance with one or more embodiments. In accordance with one or more embodiments, these methods may be implemented, for example, in logic instructions (e.g., software),Docket: GMU-24-035configurable logic, fixed-functionality hardware logic, etc., or any combination thereof executed by one or more processors of a computing device.

[0096] A new watermarking algorithm for Large Language Models (LLMs) with a primary focus on achieving distortion-free embedding of multiple information bits into the watermark is disclosed herein. A key advantage lies in providing embedding power without compromising original functionality or quality in a watermarking structure and methodology that is the first to achieve multi-bit distortion-free watermarking with efficient information decoding. This advancement opens up avenues for many applications in content authentication and communication security.

[0097] Embodiments are operational with numerous other general purpose or special purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations that may be suitable for use with various embodiments include, but are not limited to, embedded computing systems, personal computers, server computers, mobile devices, hand-held or laptop devices, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, medical device, network PCs, minicomputers, mainframe computers, cloud services, telephonic systems, distributed computing environments that include any of the above systems or devices, and the like.

[0098] Embodiments may be described in the general context of computer executable instructions, such as program modules, being executed by computing capable devices.Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Some embodiments may be designed to be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in both local and global environments.

[0099] Further, various embodiments described herein are implemented using dedicated hardware, configurable hardware or programmed processors executing programming instructions that are broadly described in flow chart form that can be stored on any suitable electronic storage medium or transmitted over any suitable electronic communication medium. A combination of these elements may be used. Those skilled in the art will appreciate that the processes and mechanisms described above can be implemented in any number of variations without departing from the present disclosure. For example, the order of certain operations carried out can often be varied, additional operations can be added, orDocket: GMU-24-035operations can be deleted, without departing from the present disclosure. Such variations are contemplated and considered equivalent.

[0100] While implementations of the disclosure are susceptible to embodiment in many different forms, there is shown in the drawings and will herein be described in detail specific embodiments, with the understanding that the present disclosure is to be considered as an example of the principles of the disclosure and not intended to limit the disclosure to the specific embodiments shown and described. In the description above, like reference numerals may be used to describe the same, similar or corresponding parts in the several views of the drawings.

[0101] In this document, relational terms such as first and second, top and bottom, and the like may be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. The terms “comprises,” “comprising,” “includes,” “including,” “has,” “having,” or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element preceded by “comprises...a” does not, without more constraints, preclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0102] Reference throughout this document to “one embodiment,” “certain embodiments,” “an embodiment,” “implementation(s),” “aspect(s),” or similar terms means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, the appearances of such phrases or in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments without limitation.

[0103] The term “or” as used herein is to be interpreted as an inclusive or meaning any one or any combination. Therefore, “A, B or C” means “any of the following: A; B; C; A and B; A and C; B and C; A, B and C.” An exception to this definition will occur only when a combination of elements, functions, steps or acts are in some way inherently mutually exclusive. Also, grammatical conjunctions are intended to express any and all disjunctive and conjunctive combinations of conjoined clauses, sentences, words, and the like, unless otherwise stated or clear from the context. Thus, the term “or” should generally be understood to mean “and / or” and so forth. References to items in the singular should beDocket: GMU-24-035understood to include items in the plural, and vice versa, unless explicitly stated otherwise or clear from the text.

[0104] Recitation of ranges of values herein are not intended to be limiting, referring instead individually to any and all values falling within the range, unless otherwise indicated, and each separate value within such a range is incorporated into the specification as if it were individually recited herein. The words “about,” “approximately,” or the like, when accompanying a numerical value, are to be construed as indicating a deviation as would be appreciated by one of ordinary skill in the art to operate satisfactorily for an intended purpose. Ranges of values and / or numeric values are provided herein as examples only, and do not constitute a limitation on the scope of the described embodiments. The use of any and all examples, or exemplary language (“e.g.,” “such as,” “for example,” or the like) provided herein, is intended merely to better illuminate the embodiments and does not pose a limitation on the scope of the embodiments. No language in the specification should be construed as indicating any unclaimed element as essential to the practice of the embodiments.

[0105] For simplicity and clarity of illustration, reference numerals may be repeated among the figures to indicate corresponding or analogous elements. Numerous details are set forth to provide an understanding of the embodiments described herein. The embodiments may be practiced without these details. In other instances, well-known methods, procedures, and components have not been described in detail to avoid obscuring the embodiments described. The description is not to be considered as limited to the scope of the embodiments described herein.

[0106] In the following description, it is understood that terms such as “first,” “second,” “top,” “bottom,” “up,” “down,” “above,” “below,” and the like, are words of convenience and are not to be construed as limiting terms. Also, the terms apparatus, device, system, etc. may be used interchangeably in this text.

[0107] The many features and advantages of the disclosure are apparent from the detailed specification, and, thus, it is intended by the appended claims to cover all such features and advantages of the disclosure which fall within the scope of the disclosure. Further, since numerous modifications and variations will readily occur to those skilled in the art, it is not desired to limit the disclosure to the exact construction and operation illustrated and described, and, accordingly, all suitable modifications and equivalents may be resorted to that fall within the scope of the disclosure.

Claims

Docket: GMU-24-035WHAT IS CLAIMED IS:

1. An encoder, comprising:when generating text by artificial intelligence (Al) the encoder operable to encode multiple bits of meta-information in a watermark of the Al-generated text in accordance with a multi-bit distortion-free watermark mapping rule to generate a watermarked Al-generated text, the watermarked Al-generated text having the same distribution as the Al-generated text.

2. The encoder of claim 1, where responsive to receiving a prompt and a text message, the encoder generates a sequence of uniform random numbers and the watermarked Al-generated text, where an initial chunk of tokens is sampled from an artificial language model that generated the Al-generated text and a plurality of tokens of the watermarked Al-generated text is determined based on a deterministic distribution interval shift coding (DISC) of the multi-bit distortion-free watermark mapping rule with the plurality of tokens of the watermarked Al-generated text generated deterministically responsive to an accumulated empirical entropy of the tokens of the initial chunk exceeding a threshold.

3. The encoder of claim 2, where the tokens of the initial chunk of tokens generated randomly prior to the accumulated empirical entropy of the tokens of the initial chunk exceeding the threshold.

4. The encoder of claim 2, where an embedding capacity of the DISC of the multi-bit distortion-free watermark mapping rule is determined by one or more of an average conditional entropy per token for the Al-generated text and a distribution of the probability of a plurality of tokens sampled from the artificial language model / Docket: GMU-24-0355. The encoder of claim 4, where the distribution of the probability of the plurality of tokens sampled from the artificial language model is the distribution of the probability of each of the tokens sampled from the artificial language model6. The encoder of claim 2, where a decoder, responsive to receiving text and a secret key shared with the encoder, is operable to determine whether the received text is watermarked and thus watermarked Al-generated text, where in accordance with a shifted score function one or more of the encoder and the decoder only considers one time any ngram context plus current token and where the decoder compares a probability value to a threshold false positive rate in determining whether the received text is watermarked and thus watermarked Al-generated text7. The encoder of claim 6, where the decoder is operable to determine one or more of a language model name, a model version of language model, generation time, AND USER ID, of the Al-generated text from the multiple bits of meta-information encoded in the watermark of the watermarked Al-generated text.

8. The encoder of claim 1, the encoder encoding meta-information in the watermark of the Al-generated text includes encoding one or more of a language model name, a model version of language model, and generation time of the Al-generated text.

9. The encoder of claim 1, where one or more of where the distribution of the generated text does not deviate from an output distribution of the original LLM and where for the watermarked Al-generated text, a token has a high correlation with a sequence associated with the watermark.

10. A method of encoding meta-data in watermarks of Al-generated text, comprising:Docket: GMU-24-035when generating text by artificial intelligence (Al) an encoder encoding multiple bits of meta-information in a watermark of the Al-generated text in accordance with a multi-bit distortion-free watermark mapping rule and thereby generating a watermarked Al-generated text, the watermarked Al-generated text having the same distribution as the Al-generated text.

11. The method of claim 10, where responsive to the encoder receiving a prompt and a text message, the encoder generating a sequence of uniform random numbers and the watermarked AI-generated text, further including sampling an initial chunk of tokens from an artificial language model that generated the Al-generated text and determining a plurality of tokens based on a deterministic distribution interval shift coding (DISC) of the multi-bit distortion-free watermark mapping rule including deterministically generating the plurality of tokens of the watermarked Al-generated text responsive to an accumulated empirical entropy of the tokens of the initial chunk exceeding a threshold.

12. The method of claim 11, further comprising randomly generating the tokens of the initial chunk of tokens prior to the accumulated empirical entropy of the tokens of the initial chunk exceeding the threshold.

13. The method of claim 11, further comprising determining an embedding capacity of the DISC of the multi-bit distortion-free watermark mapping rule based on one or more of an average conditional entropy per token for the Al-generated text and a distribution of the probability of a plurality of tokens sampled from the artificial language model.

14. The method of claim 13, where the distribution of the probability of the plurality of tokens sampled from the artificial language model is the distribution of the probability of each of the tokens sampled from the artificial language modelDocket: GMU-24-03515. The method of claim 11, further comprising a decoder, responsive to receiving text and a secret key shared with the encoder, determining whether the received text is watermarked and thus watermarked Al-generated text, one or more of the encoder and the only considering one time any ngram context plus current token in accordance with a shifted score function and the decoder comparing a probability value to a threshold false positive rate in determining whether the received text is watermarked and thus watermarked Al-generated text.

16. The method of claim 15, further comprising the decoder determining one or more of a language model name, a model version of language model, and generation time of the Al-generated text from the multiple bits of meta-information encoded in the watermark of the watermarked Al-generated text.

17. The method of claim 10, where the encoder encoding meta-information in the watermark of the Al-generated text includes encoding one or more of a language model name, a model version of language model, and generation time of the Al-generated text.

18. The method of claim 10, including one or more of where the distribution of the generated text does not deviate from an output distribution of the original LLM and where for the watermarked Al-generated text, a token has a high correlation with a sequence associated with the watermark.

19. A decoder, comprising:responsive to receiving text and a secret key shared with an encoder, the decoder operable to determine whether the received text is watermarked and thus watermarked Al-generated text, a watermark of a determined watermarked Al-generated text in accordanceDocket: GMU-24-035with a multi-bit distortion-free watermark mapping rule based on a distribution interval shift coding (DISC) of the multi-bit distortion-free watermark mapping rule.

20. The decoder of claim 19, where in accordance with a shifted score function one or more of the decoder and an encoder is operable to only consider one time any ngram context plus current token and where the decoder is operable to compare a probability value to a threshold false positive rate to determine whether the received text is watermarked and thus watermarked Al-generated text.

21. The decoder of claim 19, where the decoder is operable to determine one or more of a language model name, a model version of language model, and generation time of the AI-generated text from the multiple bits of meta-information encoded in the watermark of the watermarked Al-generated text.

22. A method of decoding a watermark of Al-generated text, comprising:responsive to receiving text and a secret key shared with an encoder, a decoder determining whether the received text is watermarked and thus watermarked Al-generated text, a watermark of a determined watermarked Al-generated text in accordance with a multibit distortion-free watermark mapping rule based on a distribution interval shift coding (DISC) of the multi-bit distortion-free watermark mapping rule.

23. The method of claim 22, further comprising one or more of the decoder and an encoder only considering one time any ngram context plus current token in accordance with a shifted score function and the decoder comparing a probability value to a threshold false positive rate to determine whether the received text is watermarked and thus watermarked Al-generated text.Docket: GMU-24-03524. The method of claim 22, further comprising the decoder determining one or more of a language model name, a model version of language model, and generation time of the AI-generated text from the multiple bits of meta-information encoded in the watermark of the watermarked Al-generated text.

25. An encoder, comprising:when generating text by artificial intelligence (Al) the encoder operable to generate a watermark of the Al-generated text in accordance with a distortion-free watermark mapping rule with random initialization to generate a watermarked Al-generated text, where the encoder responsive to a prompt and in accordance with the distortion-free watermark mapping rule generates a sequence of uniform random numbers and the watermarked Al-generated text based upon an accumulated empirical entropy of a sampled set of tokens of an artificial language model that generated the Al-generated text exceeding a threshold, in which an initial random chunk of tokens and length thereof sampled from the artificial language model is not dependent on the output of a pseudorandom function.

26. The encoder of claim 25, where the encoder responsive to the prompt and a secret key generates the sequence of uniform random numbers and the watermarked Al-generated text.The encoder of claim 25, where the watermarked Al-generated text has the same distribution as the Al-generated text.

27. The encoder of claim 25, further comprising a decoder, responsive to receiving text and a secret key shared with the encoder, operable to determine whether the received text is watermarked and thus watermarked Al-generated text, where in accordance with a shifted score function one or more of the encoder and the decoder only considers one time any ngramDocket: GMU-24-035context plus current token and where the decoder compares a probability value to a threshold false positive rate in determining whether the received text is watermarked.

28. The encoder of claim 27, where the threshold false positive rate is a maximum tolerable false positive rate for detecting a non-watermarked text as watermarked.

29. A method of encoding watermarks of Al-generated text, comprising:when generating text by artificial intelligence (Al) an encoder generating a watermark of the Al-generated text in accordance with a distortion-free watermark mapping rule with random initialization and thereby generating a watermarked Al-generated text, where responsive to a prompt and in accordance with the distortion-free watermark mapping rule the encoder generating a sequence of uniform random numbers and the watermarked Al-generated text based upon an accumulated empirical entropy of a sampled set of tokens of an artificial language model that generated the Al-generated text exceeding a threshold, in which an initial random chunk of tokens and length thereof sampled from the artificial language model is not dependent on the output of a pseudorandom function.

30. The method of claim 29, where responsive to the prompt and a secret key, the encoder generating the sequence of uniform random numbers and the watermarked Al-generated text.

31. The method of claim 29, where the watermarked Al-generated text having the same distribution as the Al-generated text.

32. The method of claim 29, further comprising:a decoder, responsive to receiving text and a secret key shared with the encoder, determining whether the received text is watermarked and thus watermarked Al-generated text, one or more of the encoder and the decoder only considering one time any ngramDocket: GMU-24-035context plus current token in accordance with a shifted score function and the decoder comparing a probability value to a threshold false positive rate in determining whether the received text is watermarked.

33. The method of claim 32, where the threshold false positive rate is a maximum tolerable false positive rate for the decoder detecting a non-watermarked text as watermarked.

34. A decoder, comprising:responsive to receiving text and a secret key shared with an encoder, the decoder operable to determine whether the received text is watermarked and thus watermarked AI-generated text, where in accordance with a shifted score function the decoder is operable to only considers one time any ngram context plus current token and where the decoder is operable to compare a probability value to a threshold false positive rate in determining whether the received text is watermarked.

35. The method of claim 34, where the watermarked Al-generated text is based on a distribution interval shift coding (DISC) of a multi-bit distortion-free watermark mapping rule.

36. The method of claim 34, where the threshold false positive rate is a maximum tolerable false positive rate for detecting a non-watermarked text as watermarked.

37. The method of claim 34, where the probability value compared by the decoder is a high probability that the decoder detects a text generated by Al as Al-generated text compared to a threshold value.

38. The method of claim 34, where the probability value compared by the decoder is a negligible probability that the decoder detects a text generated by Al as Al-generated text compared to a threshold value.Docket: GMU-24-03539. A method of decoding a watermark of Al-generated text, comprising:a decoder, responsive to receiving text and a secret key shared with an encoder, determining whether the received text is watermarked and thus watermarked Al-gen erated text, one or more of the decoder and an encoder only considering one time any ngram context plus current token in accordance with a shifted score function and the decoder comparing a probability value to a threshold false positive rate in determining whether the received text is watermarked.

40. The method of claim 39, where the watermarked Al-generated text is based on a distribution interval shift coding (DISC) of a multi-bit distortion-free watermark mapping rule.

41. The method of claim 39, where the threshold false positive rate is a maximum tolerable false positive rate for detecting a non-watermarked text as watermarked.