A text watermark embedding and detection method based on model context learning

Through a model-based context learning method and the use of a large language model for text watermark embedding and detection, the problems of low capacity and insufficient robustness of Chinese text watermarks in traditional methods are solved, and a watermark technology with high concealment and high capacity is achieved.

CN118349970BActive Publication Date: 2025-09-16TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410499472.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-24
Publication Date
2025-09-16
Estimated Expiration
2044-04-24

AI Technical Summary

Technical Problem

Traditional text watermarking technology has shortcomings in terms of large capacity and high concealment, especially watermarks based on semantic features have poor robustness and cannot meet the needs of copying, anti-counterfeiting and traceability, and existing methods have low synonym replacement capacity in Chinese texts.

Method used

A model-based context learning method is adopted. Through word segmentation, part-of-speech analysis, synonym library construction and context learning, a large language model is used to generate rewritten text and embed watermark information. The block cipher encryption algorithm is combined for embedding and detection, and the part-of-speech matching pattern is used to recover the watermark.

Benefits of technology

It achieves the goal of improving watermark capacity and robustness while maintaining semantic consistency. It is applicable to Chinese text and significantly improves the concealment and detection and recovery effects of watermarks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118349970B_ABST
    Figure CN118349970B_ABST
Patent Text Reader

Abstract

A text watermark embedding and detection method based on model context learning includes the following steps: S1. Segmenting a given text and performing part-of-speech analysis to form a vocabulary set; S2. Selecting suitable replacement words from the vocabulary set; S3. Selecting synonyms similar to the selected words from a synonym library to construct a new synonym library; S4. Creating a series of synonymous rewriting examples as corpus to form a text watermark synonym rewriting corpus; using the corpus to train a language model for contextual learning guided by thought chains; then using the language model to generate rewritten texts, selecting the rewritten text that is most semantically similar to the given text based on overall sentence semantic similarity and contextual word semantic similarity; S5. Encoding the watermark information into a binary sequence, selecting synonyms to replace the synonyms at corresponding positions in the rewritten text, and thus embedding the binary sequence into the rewritten text. This method maintains semantic consistency and concealment while improving watermark capacity and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a text watermark technology in the field of information security, and in particular to a text watermark embedding and detection method based on model context learning. Background Art

[0002] In the digital age, copyright protection and data security of text information have become particularly important. Traditional text watermarking technologies typically rely on structural features of text, such as fonts and spacing, but these methods often struggle to meet the requirements of large capacity and high confidentiality. Text watermarks can be categorized into three types, depending on the embedding method: text format-based watermarks, character feature-based watermarks, and linguistic semantic feature-based watermarks. While the first two methods offer strong confidentiality and watermark capacity, they are unable to resist re-entry attacks, nor can they achieve anti-counterfeiting traceability after the text has been copied, resulting in weak watermark robustness. Therefore, considering the requirements for anti-copying and strong robustness in specific application scenarios, semantic feature-based watermarks have become a research hotspot and application direction in recent years.

[0003] Traditional semantic-based watermarks primarily replace specific words with synonyms through a synonym dictionary. This can lead to problems such as semantic incoherence, inconsistent context, and difficulty disambiguating word meanings and parts of speech. Some studies have used BERT-based masked language models to perform context-sensitive synonym replacement and watermark embedding detection. They also achieve watermark embedding and detection by computationally generating binary encodings of words during large language model generation and selectively replacing synonyms to alter their distribution. However, because BERT can only ensure relatively coherent semantics and has poor control over semantic consistency, it is also unable to generate multiple valid synonym replacements at the lexical level for Chinese text, resulting in low capacity and therefore limited practical application scenarios.

[0004] It should be noted that the information disclosed in the above background technology section is only used to understand the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention

[0005] The main purpose of the present invention is to overcome the defects of the above-mentioned background technology and provide a text watermark embedding and detection method based on model context learning, which maintains semantic consistency and concealment while improving the watermark capacity and robustness. The watermark detection has a high degree of recovery and is more robust.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A text watermark embedding method based on model context learning includes the following steps:

[0008] S1. Segment the given text and perform part-of-speech analysis to form a vocabulary set;

[0009] S2, selecting suitable words for replacement from the vocabulary set;

[0010] S3, selecting synonyms similar to the selected words from the synonym database to construct a new synonym database;

[0011] S4. Using the constructed synonym library, combined with the original text and the selected replacement words, a series of synonymous rewriting examples are created as corpus to form a text watermark synonymous rewriting corpus; using the corpus as input, and providing task description guides and thought chain prompts, allowing the pre-trained language model to perform contextual learning under the guidance of thought chain; then, using the language model that has learned the context to generate rewritten texts, considering the embedding requirements of watermark information when generating the rewritten texts; calculating the overall sentence similarity and contextual word similarity between the generated rewritten texts and the original carrier texts, screening and sorting the rewritten texts according to the calculated similarities, and selecting the rewritten texts that are closest in semantics to the given text;

[0012] S5. Encode the watermark information into a binary sequence, select a suitable synonym for each binary bit in the binary sequence according to the synonym library constructed in step S3, and replace the synonym at the corresponding position in the rewritten text with the selected synonym, thereby embedding the binary sequence into the rewritten text.

[0013] Further:

[0014] In step S1, a word segmenter is used to segment the given carrier text and perform part-of-speech analysis to obtain a vocabulary set;

[0015] In step S2, according to the mask position screening algorithm, words suitable for replacement are screened from the word segmentation results, and these words are masked according to the priority order to form a masked text;

[0016] In step S3, synonym selection is performed using a cosine similarity calculation function;

[0017] In step S4, a language model trained with contextual learning examples is used to generate synonymous replacement rewritten text for the masked text.

[0018] In step S4, context learning specifically includes: selecting N examples closest to the text to be embedded according to the sentence encoding distance to form prompt examples of the language model, and inputting the language model to perform context learning.

[0019] In step S5, the candidate word set in the rewritten text that meets the preset similarity threshold is used as the final candidate word set for watermark embedding. For each bit of the binary sequence, the word corresponding to the bit is determined from the final candidate word set, and then a suitable synonym is selected from the constructed synonym library to replace the corresponding word in the rewritten sentence.

[0020] In step S5, the candidate word set is encoded into binary information, and the watermark is encrypted using a block cipher encryption algorithm to obtain an encrypted watermark; each code in the binary information of the candidate word set is traversed, and if it matches the encrypted watermark, the word at the corresponding position in the rewritten text is replaced with the selected synonym to complete the watermark embedding.

[0021] A text watermark detection method based on model context learning includes the following steps:

[0022] T1. Perform word segmentation on the text to be tested and identify the part of speech of the vocabulary;

[0023] T2, apply the position filter to determine the location of suspected watermark embedding in the text;

[0024] T3, using part-of-speech matching patterns to identify and extract synonymous replacement words;

[0025] T4. Query a pre-set synonymous replacement code table based on the extracted synonymous replacement vocabulary to confirm the location of the replaced vocabulary;

[0026] T5. Decode the binary code corresponding to the replacement vocabulary to restore the original embedded watermark information.

[0027] Further:

[0028] In step T2, the watermarked text is processed sentence by sentence, and a position filter is applied to determine the position in each sentence that is rewritten by synonymous substitution;

[0029] In step T3, adjacent part-of-speech tags are used to assist in matching, and the sentence is searched for the existence of <left adjacent part-of-speech t l + word component + right adjacent part of speech t r >structure.

[0030] In step T5, the binary code is decoded using the block cipher decoding function and the key to restore the original embedded watermark information.

[0031] A computer-readable storage medium stores a computer program, which implements the method when executed by a processor.

[0032] The embodiments of the present invention have the following beneficial effects:

[0033] The present invention provides a text watermark embedding and detection method based on model context learning, which can effectively utilize the large language model's ability to understand semantics and context in detail to achieve watermark embedding and detection, especially while maintaining semantic consistency and concealment, improving watermark capacity and robustness, and minimizing parameter updates to reduce training complexity.

[0034] The text watermarking method based on a large language model in the embodiment of the present invention introduces the context learning of a large language model into the text watermarking method based on linguistics for the first time, and realizes the embedding of multi-bit information in Chinese text using a large language model, which not only ensures the synonym replacement of the specified local position, but also gives full play to the maintenance of semantic fluency of the language model, thereby achieving strong concealment of the watermark and a higher watermark capacity. The method of the embodiment of the present invention utilizes the context understanding ability of the large language model, and through context learning prompts and thinking chain strategies, guides the model to automatically inject Chinese watermarks based on the understanding of the semantics of the original text, effectively solving the problem that the traditional Latin-based language model masking method based on character-level processing is not suitable for Chinese scenarios, and can significantly improve the semantic consistency and fluency before and after watermark embedding, as well as the interpretability of watermark injection. In terms of watermark detection, the embodiment of the present invention proposes watermark recovery through part-of-speech matching mode, which has a low average bit error rate, a high degree of watermark detection recovery, and stronger robustness.

[0035] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a flow chart of a text watermark embedding method based on model context learning according to an embodiment of the present invention.

[0037] Figure 2 This is a flow chart of a text watermark detection method based on model context learning according to an embodiment of the present invention.

[0038] Figure 3 This is the large language model context learning thought chain reasoning process of an embodiment of the present invention. DETAILED DESCRIPTION

[0039] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope of the present invention and its application.

[0040] See Figure 1 The embodiment of the present invention provides a text watermark embedding method based on model context learning, comprising the following steps:

[0041] S1. Segment the given text and perform part-of-speech analysis to form a vocabulary set;

[0042] S2, selecting suitable words for replacement from the vocabulary set;

[0043] S3, selecting synonyms similar to the selected words from the synonym database to construct a new synonym database;

[0044] S4. Using the constructed synonym library, combined with the original text and the selected replacement words, a series of synonymous rewriting examples are created as corpus to form a text watermark synonymous rewriting corpus; using the corpus as input, and providing task description guides and thought chain prompts, allowing the pre-trained language model to perform contextual learning under the guidance of thought chain; then, using the language model that has learned the context to generate rewritten texts, considering the embedding requirements of watermark information when generating the rewritten texts; calculating the overall sentence similarity and contextual word similarity between the generated rewritten texts and the original carrier texts, screening and sorting the rewritten texts according to the calculated similarities, and selecting the rewritten texts that are closest in semantics to the given text;

[0045] S5. Encode the watermark information into a binary sequence, select a suitable synonym for each binary bit in the binary sequence according to the synonym library constructed in step S3, and replace the synonym at the corresponding position in the rewritten text with the selected synonym, thereby embedding the binary sequence into the rewritten text.

[0046] See Figure 2 The embodiment of the present invention further provides a text watermark detection method based on model context learning, comprising the following steps:

[0047] T1. Perform word segmentation on the text to be tested and identify the part of speech of the vocabulary;

[0048] T2, apply the position filter to determine the location of suspected watermark embedding in the text;

[0049] T3, using part-of-speech matching patterns to identify and extract synonymous replacement words;

[0050] T4. Query a pre-set synonymous replacement code table based on the extracted synonymous replacement vocabulary to confirm the location of the replaced vocabulary;

[0051] T5. Decode the binary code corresponding to the replacement vocabulary to restore the original embedded watermark information.

[0052] In some embodiments, a text watermarking method based on large model context learning utilizes domain knowledge-driven word replacement selection and selects replacement words from a synonym library through cosine similarity.

[0053] In some embodiments, the thought chain reasoning-assisted watermark embedding process includes text semantic analysis, synonym selection, replacement, and semantic consistency checking.

[0054] In some embodiments, the watermark information is converted into a binary code and embedded through a synonym replacement vocabulary.

[0055] In some embodiments, a text watermark detection method utilizes replacement information and contextual cues to extract and recover watermark information from text.

[0056] The text watermarking method based on a large language model in the embodiment of the present invention introduces the context learning of a large language model into the text watermarking method based on linguistics for the first time, and realizes the embedding of multi-bit information in Chinese text using a large language model, which not only ensures the synonym replacement of the specified local position, but also gives full play to the maintenance of semantic fluency of the language model, thereby achieving strong concealment of the watermark and a higher watermark capacity. The method of the embodiment of the present invention utilizes the context understanding ability of the large language model, and through context learning prompts and thinking chain strategies, guides the model to automatically inject Chinese watermarks based on the understanding of the semantics of the original text, effectively solving the problem that the traditional Latin-based language model masking method based on character-level processing is not suitable for Chinese scenarios, and can significantly improve the semantic consistency and fluency before and after watermark embedding, as well as the interpretability of watermark injection. In terms of watermark detection, the embodiment of the present invention proposes watermark recovery through part-of-speech matching mode, which has a low average bit error rate, a high degree of watermark detection recovery, and stronger robustness.

[0057] Specific embodiments of the present invention are further described below.

[0058] Given a text T to be watermarked, which contains a series of words V, let C(V) represent the set of replaceable words selected from V, and define S(V) as the set of synonyms similar to the words in V. In the context learning process, the cosine similarity calculation function CosSim(·,·) is used to measure the similarity between the text T and the example text T i The similarity between them is expressed as D(T,T′). The thought chain prompt CoT(V) is designed to assist the pre-trained language model M in generating the rewritten text T″. The watermark information W is converted into an encoded binary sequence B through the encoding function E(W) and embedded into the text T to obtain the text T embedded with the watermark watermarked In the watermark detection and recovery phase, the replacement coding table and decoding function D(B) recorded when the watermark is embedded are used to obtain the watermark from T watermarked Extract the binary sequence B and pass it through the decoding function E -1 (B) Restore the original watermark information W.

[0059] 1. Watermark embedding method based on context learning

[0060] Domain knowledge-driven vocabulary replacement selection

[0061] During the text watermark embedding process, the target text T is first segmented to form a vocabulary set V. Then, for each word in set V, domain knowledge is used to select suitable replacement words, primarily adjectives, adverbs, and conjunctions, as these words have high semantic replacement flexibility and have minimal impact on the overall meaning of the sentence. This selection ensures the concealment of the watermark embedding and the readability of the text.

[0062] Example selection and ranking based on sentence encoding distance

[0063] To improve watermark embedding accuracy, this solution employs a sentence encoding distance-based example selection and ranking technique. By calculating the cosine similarity between the target text T and a series of candidate example texts, we identify the examples that are semantically closest to the target text. These example texts serve as contextual cues, guiding the language model to generate semantically consistent rewrites of the target text, thus providing favorable conditions for watermark embedding.

[0064] Watermark embedding assisted by thought chain reasoning

[0065] The watermark embedding process in this solution not only relies on the automatic generation capabilities of the language model but also incorporates thought chain reasoning. By designing a series of thought chain prompts, the model is guided to consider the embedding requirements of watermark information when generating rewritten text. This increases the watermark capacity when generating multiple rewritten sentences, ensuring coherence while also providing interpretability.

[0066] Figure 3 The diagram shows the reasoning process of the large language model context learning thinking chain in an embodiment of the present invention.

[0067] The example shows how to use contextual learning to rewrite a sentence.

[0068] Original sentence: "George Bernard Shaw is a prominent example."

[0069] Mask location: "A case of Shaw [mask] protruding."

[0070] Rewrite sentence 1: "George Bernard Shaw is a prominent example."

[0071] Rewrite sentence 2: "George Bernard Shaw is a prominent example."

[0072] This example of designing contextual learning involves a task of synonym replacement in Chinese text. The task requires that in a given original sentence, the positions marked with [mask] be replaced with synonyms to form a new sentence. The details are as follows:

[0073] Original sentence: "Establish posts according to demand, hire according to posts, and those who need to be transferred must be transferred."

[0074] The position that needs to be replaced is "must be", which is indicated here by two [mask] marks.

[0075] The design thinking chain is guided as follows:

[0076] The replacement approach includes the following steps:

[0077] 1. Analyze the semantics of the sentence.

[0078] 2. Choose appropriate synonyms to replace the sentence based on its semantics, and select no more than 4 synonyms as alternatives.

[0079] 3. Check whether the replaced sentence is fluent and whether it is a common expression, and explain which options were excluded when checking for fluency.

[0080] 4. Get the final rewritten sentences after screening, and provide as many rewritten sentences as possible, but no more than 3.

[0081] The results output under the guidance of the thinking chain are as follows.

[0082] The instructions for the Chinese text synonym replacement task are as follows:

[0083] Given the sentence "Establish posts based on demand, hire according to posts, and those who need to be transferred must be transferred," we need to find two appropriate words to fill the blanks in "[mask][mask] who need to be transferred must be transferred." Based on the context and meaning of the sentence, we need to choose words that express "necessity" or "certainty."

[0084] During the replacement process, the following synonyms are considered:

[0085] Must: Directly expresses necessity, but is repeated in the original sentence and is not considered a replacement.

[0086] ·Must: It emphasizes the necessity of doing something and fits the context.

[0087] Required: Usually used to express necessary requirements, but the context is not very suitable.

[0088] Necessary: ​​While it can express a need, it may not be strong enough in this context.

[0089] ·Fei De: A colloquial expression with a certain mandatory meaning.

[0090] The final rewritten sentence is as follows:

[0091] Rewrite sentence 1: Establish posts according to need, hire people according to posts, and those who need to be transferred must be transferred.

[0092] In general, the watermark embedding algorithm can be expressed as:

[0093] (1) For a given carrier text X = {x1, x2, ..., x i ,…x N}, use the word segmenter to perform word segmentation and complete part-of-speech analysis, use the mask position screening algorithm for adjectives, adverbs, and conjunctions suitable for replacement in the sentence, select words according to priority for masking, and obtain the masked sentence: X′ i ={x1,…,x {i-1} ,[MASK][MASK],x {i+1} ,…,x N}

[0094] (2) Apply the language model L learned from the context learning example C c Generate synonymous replacement sentences for the carrier text and pass the overall sentence similarity S sts Similarity S with contextual word meaning cts The candidate word set T′ that is finally included in the watermark embedding is obtained by screening and sorting.

[0095]

[0096] 2. Watermark detection method based on part-of-speech pattern matching

[0097] Watermark detection is the reverse process of watermark embedding. Since the candidate word replacement table has been stored in the embedding stage, the detection stage does not need to reapply the large language model to generate replacement words sentence by sentence, which greatly improves the efficiency of watermark decoding. In order to extract the watermark information, the detection stage first performs a watermark on the text X′ that has been embedded with the watermark. w ={x′1,x′2,…,x′ N}Apply position filter F sentence by sentence p , determine the position of each sentence to be replaced by synonyms. In this process, the adjacent part-of-speech tags are used to assist in matching, and the presence of the left adjacent part-of-speech tag in the sentence is searched. l + word component + right adjacent part of speech t r > structure, exists and the word component can obtain the word set corresponding to the part of speech position according to the synonym replacement coding table T, and then the replaced word position can be confirmed. Then query the binary code B corresponding to the replacement word at the current position and use the block cipher decoding function D {cipher} The original embedded binary sequence is obtained by using the key K, thereby recovering the watermark information W.

[0098]

[0099] In the performance test, the text watermark detection method proposed in the present invention was applied to the Baichuan model, Tongyi Qianwen, and GPT-4 language models, and comparative tests were conducted to verify the effectiveness of the method of the present invention.

[0100] The text watermark embedding and detection method based on model context learning provided by the embodiment of the present invention effectively utilizes the large language model's ability to understand semantics and context in detail to achieve watermark embedding and detection, especially while maintaining semantic consistency and concealment, improving watermark capacity and robustness, and minimizing parameter updates to reduce training complexity.

[0101] The text watermark embedding and detection method based on model context learning in this invention overcomes the limitations of traditional linguistic text watermarking technology in Chinese processing. Specific technical advantages and benefits include:

[0102] 1. Strong concealment and high capacity: Using a large language model to embed multi-bit information in Chinese text, synonym replacement of specified local locations is achieved, enhancing the concealment of the watermark and increasing the watermark capacity.

[0103] 2. Semantic Fluency: The language model’s ability to maintain semantic fluency is fully utilized, ensuring the semantic coherence and readability of the text after watermark embedding.

[0104] 3. Adapting to Chinese scenarios: Through carefully designed contextual learning prompts and thought chain strategies, the model is guided to automatically inject Chinese watermarks, solving the problem that the Latin-based language model masking method is not applicable in Chinese processing.

[0105] 4. Robustness: A watermark detection and recovery method based on part-of-speech matching pattern is proposed, which has a low average bit error rate and high recovery degree, enhancing the robustness of the watermark technology.

[0106] 5. Technological leadership: For the first time, large-scale model context learning generation is introduced into the field of text watermarking, providing a new direction for subsequent research.

[0107] This invention has broad application prospects in the information security industry, and specific application scenarios include (but are not limited to):

[0108] 1. Publishing and media industries: It can be used for copyright protection and prevention of unauthorized content copying, helping publishers and content creators track and protect their digital works.

[0109] 2. Legal Document Management: Ensure the originality and integrity of legal documents, provide secure document management and verification services, and prevent document tampering and forgery.

[0110] 3. Government and public security: Government agencies can use this technology to mark and track confidential documents to prevent the leakage and unauthorized access of sensitive information.

[0111] 4. Enterprise data protection: Enterprises can use this technology to protect business secrets and internal communications, and prevent corporate espionage and data leaks.

[0112] An embodiment of the present invention further provides a storage medium for storing a computer program, which at least performs the above method when executed.

[0113] An embodiment of the present invention further provides a control device, comprising a processor and a storage medium for storing a computer program; wherein the processor is configured to execute at least the method described above when executing the computer program.

[0114] An embodiment of the present invention further provides a processor, which executes a computer program and at least performs the method described above.

[0115] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a magnetic disk memory or a magnetic tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0116] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0117] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0118] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0119] Those skilled in the art will appreciate that all or part of the steps of the above-mentioned method embodiments may be implemented by hardware associated with program instructions, and the aforementioned program may be stored in a computer-readable storage medium. When the program is executed, the program executes the steps of the above-mentioned method embodiments. The aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0120] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0121] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0122] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.

[0123] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0124] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. Those skilled in the art will recognize that, without departing from the scope of the present invention, several equivalent substitutions or obvious variations can be made, and the performance or use of the same should be considered to fall within the scope of protection of the present invention.

Claims

1. A text watermark embedding method based on model context learning, characterized in that: The steps include: S1. Segment the given text and perform part-of-speech analysis to form a vocabulary set; S2, selecting suitable words for replacement from the vocabulary set; S3, selecting synonyms similar to the selected words from the synonym database to construct a new synonym database; S4. Using the constructed synonym library, combined with the original text and the selected replacement words, a series of synonymous rewriting examples are created as corpus to form a text watermark synonymous rewriting corpus; using the corpus as input, and providing task description guidance and thought chain prompts, the pre-trained language model is allowed to perform contextual learning guided by thought chain; then, the language model that has undergone contextual learning is used to generate rewritten text, taking into account the need to embed watermark information when generating the rewritten text; Calculate the overall sentence similarity and contextual word similarity between the generated rewritten text and the original carrier text, filter and sort the rewritten texts based on the calculated similarity, and select the rewritten text that is closest to the given text semantics; S5. Encode the watermark information into a binary sequence, select a suitable synonym for each binary bit in the binary sequence according to the synonym library constructed in step S3, and replace the synonym at the corresponding position in the rewritten text with the selected synonym, thereby embedding the binary sequence into the rewritten text.

2. The text watermark embedding method based on model context learning according to claim 1, characterized in that In step S1, a word segmenter is used to segment the given carrier text and perform part-of-speech analysis to obtain a vocabulary set; In step S2, according to the mask position screening algorithm, words suitable for replacement are screened from the word segmentation results, and these words are masked according to the priority order to form a masked text; In step S3, synonym selection is performed using a cosine similarity calculation function; In step S4, a language model that has been learned through context is used to generate a synonymous replacement rewritten text for the masked text.

3. The text watermark embedding method based on model context learning according to any one of claims 1 to 2, characterized in that: In step S4, context learning specifically includes: selecting N examples closest to the text to be embedded according to the sentence encoding distance to form prompt examples of the language model, and inputting the language model to perform context learning.

4. The text watermark embedding method based on model context learning according to any one of claims 1 to 2, characterized in that: In step S5, the candidate word set in the rewritten text that meets the preset similarity threshold is used as the final candidate word set for watermark embedding. For each bit of the binary sequence, the word corresponding to the bit is determined from the final candidate word set, and then a suitable synonym is selected from the constructed synonym library to replace the corresponding word in the rewritten sentence.

5. The text watermark embedding method based on model context learning according to claim 4, characterized in that: In step S5, the candidate word set is encoded into binary information, and the watermark is encrypted using a block cipher encryption algorithm to obtain an encrypted watermark; each code in the binary information of the candidate word set is traversed, and if it matches the encrypted watermark, the word at the corresponding position in the rewritten text is replaced with the selected synonym to complete the watermark embedding.

6. A text watermark detection method based on model context learning, detecting a text watermark embedded by the text watermark embedding method based on model context learning according to any one of claims 1 to 5, characterized in that: The steps include: T1. Perform word segmentation on the text to be tested and identify the part of speech of the vocabulary; T2, apply the position filter to determine the location of suspected watermark embedding in the text; T3, using part-of-speech matching patterns to identify and extract synonymous replacement words; T4. Query a pre-set synonymous replacement code table based on the extracted synonymous replacement vocabulary to confirm the location of the replaced vocabulary; T5. Decode the binary code corresponding to the replacement vocabulary to restore the original embedded watermark information.

7. The text watermark detection method based on model context learning according to claim 6, characterized in that: In step T2, the text with the embedded watermark is processed sentence by sentence, and a position filter is applied to determine the position in each sentence that is rewritten by synonymous substitution.

8. The text watermark detection method based on model context learning according to claim 6 or 7, characterized in that: In step T3, adjacent part-of-speech tags are used to assist in matching and find the existence of structure.

9. The text watermark detection method based on model context learning according to any one of claims 6 to 7, characterized in that: In step T5, the binary code is decoded using the block cipher decoding function and the key to restore the original embedded watermark information.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Method and device for generating and reading watermark based on deformation

    CN114817873A

  • Natural language watermarking method

    CN116522298A