Text processing method, electronic device and storage medium

By performing structural detection and embedding bit selection on the text, the problems of insufficient watermark concealment and robustness in the existing technology are solved, and effective copyright protection in the text is achieved.

CN120409463BActive Publication Date: 2025-09-09INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510875946.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-09
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Existing text watermarking technology is difficult to achieve the protection of text copyright while ensuring concealment and robustness, especially when it comes to format adjustment, multilingual text and text content modification.

Method used

By performing text structure detection on the processed text, determining the structural information of sentences, paragraphs and chapters, and selecting embedding bits based on contextual semantic relevance, watermark information is added to achieve hierarchical protection.

Benefits of technology

It ensures that the watermark information does not destroy the semantic coherence and structural integrity in the text, improves the concealment and robustness of the watermark, and effectively protects the copyright of the text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409463B_ABST
    Figure CN120409463B_ABST
Patent Text Reader

Abstract

The present application provides a text processing method, electronic device, and storage medium, relating to the fields of natural language processing technology and large model technology. The text processing method comprises: performing text structure detection on a text to be processed to obtain text structure information of the text to be processed, wherein the text structure information comprises at least one of the following: a grammatical structure to which multiple words in a sentence of the text to be processed belong, a logical structure to which multiple sentences in a paragraph of the text to be processed belong, and a chapter structure to which multiple paragraphs in a chapter of the text to be processed belong; determining at least one embedding position for adding watermark information in the text to be processed based on the correlation between the text structure information and contextual semantics in the text to be processed; and adding the target watermark information to be added to the at least one embedding position.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of natural language processing technology and large model technology, and specifically to a text processing method, electronic device and storage medium. Background Art

[0002] With the increasing application of Large Language Models (LLMs) in various fields, including documents, academic papers, and business reports, copyright ownership and content traceability of generated texts have become urgent issues. A common approach is to add watermarks to the text to protect copyright.

[0003] However, traditional text watermarking technologies have numerous limitations, making it difficult to achieve robustness while ensuring watermark concealment. For example, certain techniques based on format perturbation can easily expose the watermark due to format adjustments and are ineffective when dealing with punctuation-free or multilingual text. Techniques based on statistical features have low embedding density and are susceptible to modifications to the text content, leading to watermark signal attenuation or even loss. Techniques based on semantic perturbation can destroy the semantic coherence of the text and lack deep protection of the text structure. Summary of the Invention

[0004] In view of the above problems, the present application provides a text processing method, an electronic device and a storage medium.

[0005] According to a first aspect of the present application, a text processing method is provided, comprising:

[0006] A text structure detection is performed on the text to be processed to obtain text structure information of the text to be processed, wherein the text structure information includes at least one of the following: a grammatical structure to which multiple words in a sentence of the text to be processed belong, a logical structure to which multiple sentences in a paragraph of the text to be processed belong, and a chapter structure to which multiple paragraphs in a chapter of the text to be processed belong; based on the correlation between the text structure information and contextual semantics in the text to be processed, at least one embedding position for adding watermark information in the text to be processed is determined; and the target watermark information to be added is added to the at least one embedding position.

[0007] A second aspect of the present application provides a text processing device, including a detection module, a determination module and an adding module.

[0008] a detection module configured to perform text structure detection on the text to be processed to obtain text structure information of the text to be processed, wherein the text structure information includes at least one of the following: a grammatical structure to which multiple words in a sentence of the text to be processed belong, a logical structure to which multiple sentences in a paragraph of the text to be processed belong, and a chapter structure to which multiple paragraphs in a chapter of the text to be processed belong;

[0009] A determination module, configured to determine at least one embedding position for adding watermark information in the text to be processed based on the correlation between the text structure information and the context semantics in the text to be processed;

[0010] The adding module is used to add the target watermark information to be added to at least one embedding position.

[0011] The third aspect of the present application provides an electronic device comprising one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the steps of the above method.

[0012] The fourth aspect of the present application further provides a non-volatile computer-readable storage medium having a computer program or instruction stored thereon, which implements the steps of the above method when the computer program or instruction is executed by a processor.

[0013] The fifth aspect of the present application further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The above contents and other objects, features and advantages of the present application will become more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:

[0015] Figure 1 An application scenario diagram of the text processing method according to an embodiment of the present application is shown;

[0016] Figure 2 A flowchart of a text processing method according to an embodiment of the present application is shown;

[0017] Figure 3 A flowchart of a method for training a first detection model according to an embodiment of the present application is shown;

[0018] Figure 4 A flowchart of a sentence-level watermark embedding method according to an embodiment of the present application is shown;

[0019] Figure 5 A flow chart of a paragraph-level watermark embedding method according to an embodiment of the present application is shown;

[0020] Figure 6 A flow chart of a chapter-level watermark embedding method according to an embodiment of the present application is shown;

[0021] Figure 7 A flowchart of a method for generating chapter codes according to an embodiment of the present application is shown;

[0022] Figure 8A schematic diagram showing a watermark information embedding process according to an embodiment of the present application is shown;

[0023] Figure 9 A flow chart of a method for restoring watermark information according to an embodiment of the present application is shown;

[0024] Figure 10 A schematic diagram of the entire process of watermark embedding and watermark detection according to an embodiment of the present application is shown;

[0025] Figure 11 shows a structural block diagram of a text processing device according to an embodiment of the present application; and

[0026] Figure 12 A block diagram of an electronic device suitable for implementing a text processing method according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0027] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.

[0028] The terms used herein are only for describing specific embodiments and are not intended to limit the present application. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0029] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0030] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0031] In the technical solution of this application, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, application and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0032] The embodiment of the present application provides a text processing method, including:

[0033] A text structure detection is performed on the text to be processed to obtain text structure information of the text to be processed, wherein the text structure information includes at least one of the following: a grammatical structure to which multiple words in a sentence of the text to be processed belong, a logical structure to which multiple sentences in a paragraph of the text to be processed belong, and a chapter structure to which multiple paragraphs in a chapter of the text to be processed belong; based on the correlation between the text structure information and contextual semantics in the text to be processed, at least one embedding position for adding watermark information in the text to be processed is determined; and the target watermark information to be added is added to the at least one embedding position.

[0034] Figure 1 The following schematically shows an application scenario diagram of the text processing method according to an embodiment of the present application. Figure 1 As shown, the application scenario 100 according to this embodiment may include a large language model 101, a processor 102 and a client 103, wherein the processor 102 is used to execute the text processing method provided in the embodiment of the present application, and the user interacts with the large language model 101 and the processor 102 through the client 103. For example, the user inputs a prompt word to the large language model 101 through the client, and the large language model 101 generates relevant text based on the prompt word input by the user, such as professional documents, business reports, etc. In order to determine the copyright ownership of the generated text, the user can input the target watermark information to the processor 102 through the client 103, and the processor 102 executes the text processing method provided in the embodiment of the present application, embeds the target watermark information in the text generated by the large language model, and displays the text with the target watermark information added to the user through the client 103. Similarly, the processor 102 can also perform watermark restoration on the text embedded with the watermark information to trace the content and prevent the text from being tampered with.

[0035] The following will be based on Figure 1 The scene described by Figures 2 to 6 The text processing method of the application embodiment is described in detail.

[0036] Figure 2 The flowchart of the text processing method according to an embodiment of the present application is schematically shown.

[0037] like Figure 2 As shown, the method of this embodiment includes operations S210 to S230, and the text processing method can be executed by a host controller in a storage device.

[0038] In operation S210 , text structure detection is performed on the text to be processed to obtain text structure information of the text to be processed.

[0039] According to an embodiment of the present application, the text structure information includes at least one of the following: the grammatical structure to which multiple words in the sentences of the text to be processed belong, the logical structure to which multiple sentences in the paragraphs of the text to be processed belong, and the chapter structure to which multiple paragraphs in the chapters of the text to be processed belong.

[0040] In operation S220, at least one embedding position for adding watermark information in the text to be processed is determined based on the correlation between the text structure information and the contextual semantics in the text to be processed.

[0041] In operation S230, target watermark information to be added is added to at least one embedding position.

[0042] In one example, since the LLM-generated text has structural characteristics such as grammatical diversity, logical rigor, and textual standardization, watermark embedding is required to achieve hierarchical protection of the three-level self-similar features of sentences, paragraphs, and chapters without destroying semantic coherence and structural integrity. However, related technologies are difficult to meet the copyright protection needs in the above complex scenarios. Therefore, in an embodiment of the present application, a text structure detection is first performed on the text to be processed. The text structure detection may be performed using a detection model, for example, a text structure detection is performed based on a pre-trained Hidden Markov Model (HMM) to obtain text structure information of the text to be processed. Specifically, the text structure information includes the grammatical structure to which multiple words in the sentences of the text to be processed belong, and the grammatical structure may be, for example, subject-verb-object, subject-copula-table, adverbial-predicate, attributive-predicate, verb-object, and parallel, etc.; the text structure information also includes the logical structure to which multiple sentences in the paragraphs of the text to be processed belong, and the logical structure may be, for example, a general-specific structure, a specific-general structure, a progressive structure, a parallel structure, a cause-and-effect structure, a transitional structure, and a sequential structure, etc.; the text structure information also includes the chapter structure to which multiple paragraphs in the chapters of the text to be processed belong, for example, it may be an introduction, a conclusion, and other chapter structures. Through pre-training models and efficient parsing algorithms, we can quickly perform structure detection and watermark embedding on the processed text, thereby improving the efficiency of the entire watermark embedding process. Based on the detection model, we can perform text structure detection on the processed text, accurately identify the grammatical structure, logical structure and paragraph structure in the text, and ensure that watermark embedding is carried out without destroying the semantic coherence and structural integrity of the text, thus providing a basis for achieving semantically lossless copyright protection.

[0043] After determining the text structure information, one or more embedding locations for adding watermark information to the text to be processed are determined based on the correlation between the text structure information and the contextual semantics of the text to be processed. The embedding locations determined in operation S220 include at least one or more of the sentence level, paragraph level, and chapter level, providing a basis for implementing hierarchical protection of self-similar features at the sentence, paragraph, and chapter levels. The target watermark information to be added can be, for example, default watermark information or user-edited watermark information. The target watermark information to be added is added to at least one of the aforementioned embedding locations to complete the hierarchical watermark embedding at the sentence, paragraph, and chapter levels.

[0044] The text processing method provided by the embodiment of the present application can accurately identify the grammatical structure, logical structure and paragraph structure of the text by performing text structure detection on the text to be processed. This provides a basis for the subsequent determination of the watermark embedding position, ensuring that the embedding of the watermark does not destroy the semantic and structural integrity of the text. The target watermark information is added to the determined embedding position, thereby achieving effective embedding of the watermark information in the text. Compared with the watermark embedding methods in related technologies, such as the method of adding watermarks based on statistical feature technology (such as word frequency distribution deviation), it is susceptible to synonym replacement or sentence reorganization, resulting in signal attenuation; for example, the method of embedding watermarks through synonym replacement or word order adjustment based on semantic perturbation technology may destroy semantic coherence and lack association between levels, and local modifications are prone to global failure. The method of the embodiment of the present application adds watermark information to the semantic structure position in the document at various semantic structure positions of the text, such as the end of a word in an adverbial structure, the end of a sentence in a progressive relationship, the end of a paragraph in a chapter, etc. In this way, the watermark information does not destroy the original grammatical structure and does not affect semantic understanding and semantic parsing; in addition, because these watermarks are added to the semantic structure position, these positions are usually not modified when the text is modified. Therefore, text modification will not destroy the original watermark information, and has strong robustness.

[0045] Among them, in the process of determining at least one embedding position for adding watermark information in the text to be processed, the reference factors include: text structure information and the correlation of the contextual semantics in the text to be processed. For example, after detecting that the sentence in the text to be processed contains multiple grammatical structure positions, further based on the correlation between each grammatical structure position and the contextual semantics, the position with a low correlation with the contextual semantics (less than a preset correlation threshold) among the multiple grammatical structure positions is used as the embedding position for adding the watermark. Because the position with a low correlation with the contextual semantics belongs to the semantic segmentation position in the text, it has a weak correlation with the contextual semantics and has a strong independence. The text content here is usually not easily modified by the user. Therefore, adding watermark information here is not easy to be destroyed, which further improves the robustness of the watermark.

[0046] The text processing method provided in the embodiment of the present application can be executed by a processor, and the processor and the large language model can be deployed on the same server or on different servers, which is not limited here. For example, the processor can be deployed on a first server, and the large language model can be deployed on a second server. After the first server receives the creative text generated by the user using the language model, it performs text structure detection on the creative text to obtain text structure information of the creative text, such as the grammatical structure to which multiple words in the sentences of the creative text belong, the logical structure to which multiple sentences in the paragraphs of the creative text belong, and the chapter structure to which multiple paragraphs in the chapters of the creative text belong; based on the correlation between the text structure information and the contextual semantics in the creative text, at least one embedding position for adding watermark information in the creative text is determined; the first server receives the target watermark information input by the user through the client, and adds the target watermark information to at least one embedding position; the client is used to display the target text with the target watermark information added to the user. Since the watermark information is embedded in the embedding position determined based on the text structure information and the contextual semantic correlation, it is ensured that the embedding of the watermark does not destroy the semantics and structural integrity of the text, thereby achieving effective protection of the copyright of the user's creative text.

[0047] According to an embodiment of the present application, text structure detection is performed on the text to be processed to obtain text structure features of the text to be processed. Text structure detection can be performed using a detection model. Three separate detection models can be trained for sentence-level, paragraph-level, and chapter-level text structure predictions, and then predictions can be performed.

[0048] For example, it can be a sentence-level grammatical structure prediction. First, a first detection model is trained, and a first detection is performed on multiple words based on the first detection model to obtain multiple first probabilities that each word belongs to multiple candidate grammatical structures, and the grammatical structure to which the word belongs is determined based on the multiple first probabilities.

[0049] For example, for paragraph-level logical structure prediction, a second detection model may be first trained, and a second detection may be performed on multiple sentences based on the second detection model to obtain multiple second probabilities that each sentence belongs to multiple candidate logical structures, and the logical structure to which the sentence belongs may be determined based on the multiple second probabilities.

[0050] For example, for chapter-level text structure prediction, a third detection model may be first trained, and a third detection may be performed on multiple paragraphs based on the third detection model to obtain multiple third probabilities that each paragraph belongs to multiple candidate text structures, and the text structure to which the paragraph belongs may be determined based on the multiple third probabilities.

[0051] The above three detection models may adopt hidden Markov models. The construction processes of these three detection models are respectively introduced below.

[0052] After training the three detection models, the model parameters include state transition information (state transition matrix) and observation probability information (observation probability matrix).

[0053] Each element in the state transition matrix represents the probability of transitioning from the first text structure to the second. Specifically, if the current text object (word, sentence, paragraph) is in the first text structure, the probability that the next text object (word, sentence, paragraph) will be in the second text structure. Text structure refers to the grammatical structure of the preceding words, the logical structure of the sentences, and the discourse structure of the paragraphs.

[0054] Each element in the observation probability matrix represents the probability of occurrence of various textual expressions under the second text structure. Specifically, it represents the probability of occurrence of various textual expressions corresponding to the second text structure. For detection models at different levels, it represents the probability of occurrence of various textual expressions corresponding to an observed word under a certain grammatical structure, the probability of occurrence of various textual expressions corresponding to an observed sentence under a certain logical structure, and the probability of occurrence of various textual expressions corresponding to an observed paragraph under a certain textual structure.

[0055] Figure 3 The flowchart of the training method of the first detection model according to the embodiment of the present application is schematically shown. Figure 3 As shown, it includes operations S310 to S330.

[0056] In operation S310 , a training text is obtained, and word segmentation processing is performed on a plurality of sample sentences in the training text to obtain a plurality of sample words.

[0057] In operation S320 , the grammatical structures to which the plurality of sample words belong are determined to generate a first label of the training text.

[0058] In operation S330 , a basic model is trained using the plurality of sample words and the first label to obtain a first detection model.

[0059] According to an embodiment of the present application, the model parameters of the first detection model include first state transition information and first observation probability information, wherein the first state transition information is used to represent the probability of transitioning from the first candidate grammatical structure to the second candidate grammatical structure, and the first observation probability information is used to represent the probability of each of the multiple target sample words belonging to the second candidate grammatical structure, and the first candidate grammatical structure and the second candidate grammatical structure belong to at least one of the multiple grammatical structures to which the multiple sample words belong. The first candidate grammatical structure and the second candidate grammatical structure can be the same, and both are one of the multiple grammatical structures; the first candidate grammatical structure and the second candidate grammatical structure can also be different, and both are two different grammatical structures among the multiple grammatical structures.

[0060] In one example, the construction process of the sentence-level hidden Markov model in the embodiment of the present application is as follows: first, a training text containing multiple sample sentences is collected, and each sample sentence in the training text is segmented to obtain multiple sample words. For each sample word, the grammatical structure to which it belongs is marked, such as subject-verb-object, adverbial-predicate, parallel and other grammatical structures, to obtain a sentence-level dataset: the grammatical structure of each sample word (subject-verb-object, 、 , parallel etc.), generate sample words and their corresponding grammatical structure labels, which serve as training samples to train the basic model, resulting in the first detection model, namely the sentence-level hidden Markov model. After training, the model parameters include first state transition information and first observation probability information. The first state transition information is used to represent the probability of transitioning from the first candidate grammatical structure to the second candidate grammatical structure (assuming the current word is the first candidate grammatical structure, the probability of the next word being the second candidate grammatical structure). The first observation probability information is used to represent the probability of each target sample word appearing in the second candidate grammatical structure (the probability of each of the multiple textual expressions of the observation word corresponding to the second candidate grammatical structure appearing).

[0061] According to an embodiment of the present application, the process of constructing a paragraph-level hidden Markov model is as follows:

[0062] A training text is obtained and multiple sample paragraphs in the training text are separately processed to obtain multiple sample sentences. The logical structures to which each of the multiple sample sentences belongs are determined to generate a second label for the training text. A basic model is trained using the multiple sample sentences and the second labels to obtain a second detection model, wherein model parameters of the second detection model include second state transition information and second observation probability information, wherein the second state transition information is used to represent the probability of transitioning from a first candidate logical structure to a second candidate logical structure, and the second observation probability information is used to represent the probability of each of multiple target sample sentences belonging to the second candidate logical structure, wherein the first candidate logical structure and the second candidate logical structure belong to at least one of the multiple logical structures to which the multiple sample sentences belong.

[0063] In one example, similar to sentence-level hidden Markov model training, training text containing multiple sample paragraphs is first collected. Each sample paragraph in the training text is then segmented into sentences, yielding multiple sample sentences. Sentence segmentation can be performed using commas or periods as delimiters, with sentences ending with commas or periods considered as one sentence. For example, for the text "First, let's take an example...", "First," and "Let's take an example" are segmented into two sentences. Each sample sentence is annotated with its assigned logical structure, such as general-to-specific, progressive, or causal, generating a second label for the training text and obtaining a paragraph-level dataset. The sample sentences and their corresponding logical structure labels are then used to train a base model to obtain a second detection model, namely, a paragraph-level hidden Markov model. The parameters of the second detection model include second state transition information, which represents the probability of transitioning from the first candidate logical structure to the second candidate logical structure (i.e., the probability that the current sentence is the first candidate logical structure and the following sentence is the second candidate logical structure). The second observation probability information represents the probability of each target sample sentence occurring under the second candidate logical structure (i.e., the probability of each of the various textual representations of the observation sentence corresponding to the second candidate logical structure occurring).

[0064] According to an embodiment of the present application, the construction process of the chapter-level hidden Markov model is as follows: obtain a training text, and segment the training text to obtain a plurality of sample paragraphs. Determine the chapter structure to which the plurality of sample paragraphs belong to generate a third label for the training text. Use the plurality of sample paragraphs and the third label to train the basic model to obtain a third detection model, wherein the model parameters of the third detection model include third state transition information and third observation probability information, wherein the third state transition information is used to characterize the probability of transitioning from the first candidate chapter structure to the second candidate chapter structure, and the third observation probability information is used to characterize the probability of each of the plurality of target sample paragraphs belonging to the second candidate chapter structure, and the first candidate chapter structure and the second candidate chapter structure belong to at least one of the plurality of chapter structures to which the plurality of sample paragraphs belong.

[0065] In one example, a training text containing multiple sample paragraphs is first collected, and the training text is segmented to obtain multiple sample paragraphs. Each sample paragraph is labeled with the chapter structure to which it belongs, such as the introduction, main text, conclusion, etc., and a third label of the training text is generated to obtain a chapter-level dataset. It should be understood that since the choice of observation set does not affect the watermark embedding and detection effect, it is only used as an example here and is not fully listed. The basic model is trained using sample paragraphs and their corresponding chapter structure labels to determine the third state transition information and third observation probability information in the third detection model, which is a chapter-level hidden Markov model. The third state transition information represents the probability of transitioning from the first candidate chapter structure to the second candidate chapter structure (assuming that the current paragraph is the first chapter structure and the next paragraph is the second chapter structure). The third observation probability information represents the probability of each target sample paragraph appearing under the second candidate chapter structure (the probability of each of the multiple text expressions of the observation paragraph corresponding to the second chapter structure appearing).

[0066] When constructing the data set, we define the state set, observation set, state transition matrix and observation probability matrix of each level of the hidden Markov model. Taking the sentence-level hidden Markov model as an example, the state set is , where it is assumed Representing the subject-verb-object structure, Characterize the structure of the Characterizes the parallel structure. The observation set is , where, for example, the observation set can represent the corresponding multiple text expressions under the current grammatical structure, such as, in the adverbial structure, various prepositions usually appear, then It can represent a variety of prepositional expressions, such as It means "in", Means "from", Indicates "to". Define the state transfer information composed of the transition probability between states, that is, the state transfer matrix , as shown in formula (1):

[0067] ----(1);

[0068] in Indicates the slave state Transfer to state The probability of, for example Indicates that the current word is the subject, predicate, and object , the next word is transferred to the adverbial structure The probability is 0.2.

[0069] Define the observation probability information composed of the generation probability of the observed words in the current state, that is, the observation probability matrix , as shown in formula (2):

[0070] ----(2);

[0071] in Indicates status Generate observation words The probability of, for example Indicates that the current statement is a predicate structure , the next word is generated The probability of "from" is 0.7.

[0072] The state and observation dimensions of the paragraph-level and chapter-level hidden Markov models are consistent with those of the aforementioned datasets. The rest of the construction process is the same as that of the sentence-level hidden Markov model and will not be repeated here.

[0073] The method for training the three hidden Markov models mentioned above can be to adopt a variety of feasible training strategies for training.

[0074] One method is to use data statistics to calculate the probability of each semantic structure transitioning to another semantic structure when the sample data volume is large enough to obtain state transition information; and to calculate the probabilities of multiple text expressions corresponding to each semantic structure to obtain observation probability information.

[0075] For example, when training a sentence-level hidden Markov model, we calculate the probability that the words following each sample word belong to various grammatical structures (subject-verb-object, adverbial-predicate, parallel, etc.). This yields the probability of transitioning from the first candidate grammatical structure (the grammatical structure of the current word) to the second candidate grammatical structure (the grammatical structure of the following word), generating a state transition matrix. Furthermore, we calculate the probability of the corresponding textual representation of the observed word under each grammatical structure. For example, for the adverbial-predicate structure, we calculate the probability of various prepositions appearing in the sample to generate an observation probability matrix.

[0076] Another method may be to adopt the maximum likelihood estimation method, or to train the above three models in combination. For example, the first detection model is trained first, the second detection model is trained based on the model parameters of the first detection model, and the third detection model is trained based on the model parameters of the second detection model.

[0077] For example, the three hidden Markov models mentioned above are trained with the goal of maximizing the joint likelihood of the observation sequence , is the log-likelihood function of the model, which indicates that the observation sequence is The logarithm of the joint probability under . Maximize , that is, find the model parameters that maximize the probability of the observation sequence As shown in formula (3):

[0078] ----(3);

[0079] Where l is the level, and its values ​​are 1, 2, and 3, corresponding to sentence level, paragraph level, and chapter level, respectively. is the length of the feature sequence at the lth level. For the sentence level, is the number of words in the sentence; for paragraph level, is the number of sentences in the paragraph; for the chapter level, is the number of paragraphs in the chapter; Indicates that the lth level is at time step Observation values ​​(e.g., textual expressions of observation words, observation sentences, and observation paragraphs); Indicates that the lth level is at time step hidden state (such as the grammatical structure of words, the logical structure of sentences, and the textual structure of paragraphs); For the level The model parameters, is the state transition matrix, is the observation probability matrix. Refers to a given hidden state and model parameters When the observation value The emission probability of , in the hidden Markov model, this is usually determined by the observation probability matrix.

[0080] Objective function is the sum of the log-likelihood probabilities at all levels (sentence level, paragraph level, chapter level). Specifically, for each level l, the observation value at each time step t in that level is calculated At a given hidden state and model parameters The emission probabilities under these conditions are accumulated over time and levels to maximize the objective function. , using a forward-backward algorithm of a hidden Markov model, such as the Baum-Welch algorithm, to iteratively optimize , until the likelihood probability changes less than the acceptable error, such as , which is as follows (4):

[0081] ----(4);

[0082] Among them, n represents the number of iterations, that is, the current one is the nth iteration.

[0083] , representing the model parameters (including the state transition matrix and the observation matrix) at the n-th iteration;

[0084] , representing the model parameters after the (n + 1)-th iteration update;

[0085] , representing the value of the log-likelihood function of the model at the n-th iteration.

[0086] By performing the above Hidden Markov Model training on the sentence-level syntactic state, paragraph-level logical state, and chapter-level discourse state, the state transition matrix and the observation probability matrix at each level are generated. The embodiments of the present application do not limit the training method of the Hidden Markov Model.

[0087] After training the above three detection models, the semantic structure position detection is performed using the detection models. Hereinafter, taking the method of performing the first detection on multiple words based on the first detection model to obtain multiple first probabilities that each word belongs to multiple candidate syntactic structures as an example, the method of performing detection using the Hidden Markov Model is exemplarily described.

[0088] According to the embodiments of the present application, performing the first detection on multiple words based on the first detection model to obtain multiple first probabilities that each word belongs to multiple candidate syntactic structures includes: determining multiple first probabilities that each word belongs to multiple candidate syntactic structures by using the first state transition information and the first observation probability information.

[0089] In one example, after training the state transition information (state transition matrix) and the observation probability information (observation probability matrix) of each level of the detection model, first, the text to be processed is segmented to obtain a word sequence, and the word sequence is input into the first detection model. For each word, the first probability of it belonging to each candidate syntactic structure is calculated using the state transition matrix and the observation probability matrix of the model to determine the syntactic structure of the word. Taking the sentence-level Hidden Markov Model to detect the text to be processed as an example, assume detecting the syntactic structure of the sentence "He quickly solved the problem". After segmenting the sentence "He quickly solved the problem", the word sequence ["He", "quickly", "solved", "了", "problem"] is obtained.

[0090] The specific detection process is as follows for example:

[0091] Assume the initial state probability distribution: [0.6, 0.3, 0.1];

[0092] Word sequence: ["He", "quickly", "solved", "了", "problem"];

[0093] Assume reading the observation probability (B) of each vocabulary according to the observation probability matrix of the model as follows:

[0094] "He": The observation probability is 0.6 in the syntactic structure 0 - state and 0.1 in the syntactic structure 1 - state and 0.2 in the syntactic structure 2 - state and 0.2 in the syntactic structure 2 - state.

[0095] "Quickly": The observation probability is 0.2 in the state and 0.8 in the state and 0.1 in the state and 0.1 in the state <0oo0283>

[0096] "Solve": The observation probability is 0.7 in the state and 0.1 in the state and 0.2 in the state and 0.2 in the state

[0097] "Le": The observation probability is 0.2 in the state and 0.5 in the state and 0.1 in the state<os00292>and 0.1 in the state

[0098] "Problem": The observation probability is 0.1 in the state and 0.1 in the state and 0.8 in the state and 0.8 in the state

[0099] Assume that the transition probabilities (A) between states are read according to the state transition matrix of the model as follows: The probability from to is 0.3, and the probability from to <000os03>is 0.5... and so on. For each word and each state, calculate the maximum probability of reaching that state along the path and record the path. Finally, obtain the most likely state sequence: That is:

[0100] He ( , 0.6): The subject in the subject - predicate - object structure. It represents the main body of the sentence, performing actions or assuming states.

[0101] Quickly ( , 0.8): The adverbial in the adverbial - head structure. It modifies the verb "solve", describing the manner, time or condition of the action.

[0102] Solve ( , 0.7): The predicate in the subject - predicate - object structure. It represents the action or state issued by the subject.

[0103] ( , 0.5): An adverbial phrase in an adverbial-indirect structure. Here, it indicates the completed state of an action, supplementing the state of the action of "solving".

[0104] question( , 0.8): The object in the parallel structure. It indicates the object of the predicate verb "solve" and is the object component of the sentence.

[0105] By detecting the text to be processed, a basis is provided for the subsequent text watermark embedding, ensuring that the watermark is embedded in the grammatically correct text position, thereby improving the concealment and robustness of the watermark.

[0106] It should be understood that the principles and processes of performing a second detection on a sentence based on the second detection model to obtain the logical structure of the sentence and performing a third detection on a paragraph based on the third detection model to determine the chapter structure to which the paragraph belongs are similar to the aforementioned process of performing a first detection on multiple words based on the first detection model, and will not be repeated here.

[0107] After determining the text structure features of the text to be processed, at least one embedding position for adding watermark information in the text to be processed is further determined according to the correlation between the text structure information and the context semantics in the text to be processed.

[0108] According to an embodiment of the present application, based on the text structure information and the correlation between the contextual semantics in the text to be processed, determining at least one embedding position for adding watermark information in the text to be processed includes at least one of the following: based on the grammatical structure to which multiple words belong, the position of the target word that belongs to the predetermined grammatical structure and whose contextual semantic correlation meets the predetermined condition (for example, the correlation is less than a preset correlation threshold) is used as the first embedding position; based on the logical structure to which multiple sentences belong, the position of the target sentence that belongs to the predetermined logical structure and whose contextual semantic correlation meets the predetermined condition is used as the second embedding position; based on the chapter structure to which multiple paragraphs belong, the position of the target paragraph that belongs to the predetermined chapter structure and whose contextual semantic correlation meets the predetermined condition is used as the third embedding position.

[0109] In one example, while obtaining the text structure information according to operation S210, the observation probabilities corresponding to each word, sentence, and paragraph in the text to be processed are also determined. The threshold Q is used to screen out the positions with higher observation probabilities as candidate embedding positions. Only the positions with observation probabilities greater than or equal to Q will be determined as candidate embedding positions. The value of the threshold Q can be adjusted according to specific requirements and the experimental results of the validation set. A higher Q value indicates that only positions with very high observation probabilities are selected for embedding, which can improve the invisibility and robustness of the watermark, but may reduce the number of embeddable positions, thereby reducing the watermark capacity. On the contrary, a lower Q value can increase the number of embedding positions and improve the watermark capacity, but may reduce the invisibility and robustness. Suppose Q = 0.6, then only the positions with observation probabilities ≥ 0.6 will be selected as candidate embedding positions.

[0110] For example, taking the detection of a sentence using a sentence-level hidden Markov model as an example, when detecting the syntactic structure of "He solved the problem this month", it is "He" (subject-verb-object 0.6) → "In" (adverbial-middle 0.8) → "This month" (adverbial-middle 0.7) → "Solve" (subject-verb-object 0.5) → "Problem" (subject-verb-object 0.8), then it is determined that the positions where the words "He" ( , 0.6), "In" ( , 0.8), "This month" ( , 0.7), "Problem" ( , 0.8) are located are used as candidate embedding positions.

[0111] After determining the candidate embedding positions, it is also necessary to consider whether the context semantic relevance of the words at the candidate embedding positions meets the predetermined conditions. Specifically, for example, in the adverbial-middle structure, there will probably be some prepositions, time words, place nouns, etc. in this structure. These words are used as semantic segmentation positions, and their semantic relevance to their previous or next text is not strong. These words are used as the first embedding positions, and embedding watermark information at these positions does not damage the original semantics. For example, by detecting the context semantic relevance of the words at each candidate embedding position (for example, using a language model), it is found that the semantic relevance of "In" to the context (the relevance is 0.3) is lower than the preset relevance threshold of 0.5. Then, "In" is used as the embedding position.

[0112] After determining the logical structures to which multiple sentences belong, the positions of sentences belonging to the predetermined logical structure are selected as candidate embedding positions. The semantic relevance of the context is then determined to meet predetermined conditions, and the positions of sentences whose semantic relevance meets the predetermined conditions are determined as the second embedding positions. For example, in a causal structure, conjunctions such as "because," "so," and "therefore" may appear (here, these conjunctions are treated as independent sentences). After relevance testing, these sentences are generally less relevant to the context and, as semantic segmentation positions, have weak semantic relevance to either the preceding or following context. Therefore, the positions of the conjunctions are selected as the second embedding positions, and embedding watermark information at these positions does not destroy the original semantics.

[0113] Similarly, after determining the chapter structure to which each of the multiple paragraphs belongs, the paragraphs belonging to the predetermined chapter structure are selected as candidate embedding locations. Their contextual semantic relevance is then determined to determine whether it meets the predetermined criteria. The target paragraphs whose contextual semantic relevance meets the predetermined criteria are then selected as the third embedding location. For example, the final paragraph of a chapter may not have strong semantic relevance to the next chapter. The initial paragraph of a chapter may not have strong semantic relevance to the previous chapter. These locations are used as third embedding locations, and embedding watermark information in these locations does not destroy the original semantics.

[0114] According to the embodiments of the present application, watermarks are embedded at locations with low semantic relevance without destroying the original semantics, and no errors will occur when the language model is used to detect the text subsequently. Moreover, since the semantics of these locations are relatively independent, they are usually not easily modified when users fine-tune the text, and the watermark is not easily destroyed.

[0115] This embodiment of the present application determines the embedding location for adding watermark information based on text structure information, ensuring that the watermark is embedded in a relatively hidden location within the text with minimal semantic impact. This not only improves the watermark's concealment, but also reduces its interference with text readability, while ensuring that the watermark is stably embedded within the text. Determining the embedding location at different levels (sentence, paragraph, and chapter) allows the watermark to be flexibly embedded in the appropriate location based on the text's structural characteristics, enhancing the flexibility and adaptability of watermark embedding.

[0116] According to an embodiment of the present application, adding the target watermark information to be added to at least one embedding position includes: generating a J-bit watermark code (for example, a binary code) in a predetermined format based on the target watermark information, where J is a positive integer; converting the j-bit target code in the J-bit watermark code into an invisible j-bit predetermined character, where the predetermined character includes a zero-width character; and sequentially adding the j-bit predetermined character to j target embedding positions in a plurality of embedding positions, where j is a positive integer less than or equal to J.

[0117] In one example, the target watermark information is converted into an encoded sequence in a specific format (such as binary, UTF-8 encoding), where J represents the total number of bits of the encoded sequence. For example, after encoding "©2024 AI-generated" using UTF-8, it is converted into a binary bitstream. For example, a total of J bits are generated, such as a 16-bit binary bitstream "0010101101001001" for processing.

[0118] Convert the j-bit target encoding in the J-bit watermark encoding into an invisible j-bit predetermined character. For example, one of the characters "0" or "1" in the binary encoding can be converted into a zero-width character for embedding.

[0119] Specifically, sequentially add the j-bit predetermined characters to the j target embedding bits among the multiple embedding bits.

[0120] For example, when using "1" as the target encoding, the method of embedding the watermark encoding "0010101101001001" into the statement "Zhang San failed to fulfill the payment obligation stipulated in the contract within a reasonable period" is as follows:

[0121] First, the embeddable grammar bits selected after segmenting "Zhang San failed to fulfill the payment obligation stipulated in the contract within a reasonable period" are "Zhang San", "in", "within a reasonable period",....

[0122] The first bit "0" of the watermark encoding corresponds to the first embedding bit "Zhang San", the second bit "0" of the watermark encoding corresponds to the second embedding bit "in", the third bit "1" of the watermark encoding corresponds to the third embedding bit "within a reasonable period",.... Convert the third bit "1" of the watermark encoding into a zero-width character and embed it correspondingly into the third embedding bit "within a reasonable period" (the positions where the watermark encoding is "0" are not processed). That is, sequentially insert zero-width characters into the selected target embedding bits. The zero-width character can be, for example, a zero-width space (ZeroWidth Space, ZWS), whose Unicode encoding is U+200B, to complete the embedding of the watermark information. The embedded text is: Zhang San failed to fulfill the payment obligation stipulated in the contract within a reasonable period\u200B.

[0123] By converting the watermark encoding into invisible characters and sequentially adding them to the embedding bits, an efficient and concealed watermark embedding method is achieved, which can embed the watermark information into the text without affecting the readability of the text.

[0124] Next, it will be combined with Figures 4 to 7 Introduce the process of hierarchical watermark embedding provided according to the embodiments of the present application.

[0125] Figure 4 Schematically shows a flowchart of the statement-level watermark embedding method according to the embodiments of the present application. As Figure 4As shown, the method of this embodiment includes operation S410 and operation S420.

[0126] In operation S410, the target watermark information is encoded in a predetermined format to generate an M-bit initial code, where M is a positive integer.

[0127] In operation S420, at least a part of the M-bit initial code is added to the first embedding bit.

[0128] In one example, first, the target watermark information "©2024 Text AI Generated" is preprocessed, that is, the target watermark information is encoded in a predetermined format to generate an M-bit initial code. Assuming M = 16, the initial code (B Figure 5 , , , Figure 5 ,

[0134] ,

[0133] ) is 0010101101001001. When performing sentence-level watermark information embedding, at least a part of the M-bit initial code is added to the first embedding bit. For example, in the sentence "Zhang San failed to fulfill the payment obligation agreed in the contract within a reasonable period and still did not pay after being repeatedly urged by Li Si.", the first embedding bits (candidate bit 1, candidate bit 2, and candidate bit 3) with an observation probability ≥ 0.6 are detected and screened through a sentence-level hidden Markov model as follows:

[0129] Candidate bit 1: between "in" and "within a reasonable period" (adverbial-middle structure, probability 0.8);

[0130] Candidate bit 2: between "contractually agreed" and "of" (determinative-middle structure, probability 0.7);

[0131] Candidate bit 3: between "repeatedly" and "urged" (adverbial-middle structure, probability 0.9).

[0132] Since the first 3 bits of the 16-bit initial code 0010101101001001 are 001, the watermark information embedding operation in this sentence is: embed the 1st bit 0 (do not insert ZWS), embed the 2nd bit 0 (do not insert ZWS), embed the 3rd bit 1 (insert \u200B). Finally, the sentence after embedding the watermark information becomes: "Zhang San failed to fulfill the payment obligation agreed in the contract within a reasonable period and still did not pay after being repeatedly \u200B urged by Li Si." The actually embedded bit stream needs to match all candidate bits in order. This is only a simplified example here.

[0133] Figure 5 Schematically shows a flowchart of a paragraph-level watermark embedding method according to an embodiment of the present application. As Figure 5 shown, it includes operation 510 and operation 520.

[0134] In operation S510, the M-bit initial code generated based on the target watermark information is encoded based on the error correction mechanism to generate an (M+N)-bit error correction code, where the error correction code includes the M-bit initial code and an N-bit redundant check code, where N is a positive integer; the redundant check code is used to: in the case where the initial code is modified, encode the modified initial code to restore the modified initial code.

[0135] In operation S520, at least part of the N-bit redundancy check code is added to the second embedded bits.

[0136] In one example, the 16-bit initial code generated based on the target watermark information is encoded based on the error correction mechanism to generate a (16+4)-bit error correction code. For example, the Reed-Solomon (RS) coding technology can be used based on the 16-bit initial code. Generate 4-bit redundant check code (Assuming the redundant bits are 1100), the redundant check code can restore the modified initial code when the initial code is modified, and finally obtain a 20-bit error correction code consisting of 16-bit initial code and 4-bit redundant check code:

[0137] ;

[0138] When embedding paragraph-level watermark information, at least part of the N-bit redundant check code is added to the second embedding bit. Assume that the target paragraph contains 3 sentences:

[0139] Statement 1: First, Zhang San failed to fulfill his payment obligation.

[0140] Sentence 2: Secondly, Li Si made repeated demands but to no avail.

[0141] Statement 3: Finally, Zhang San committed a breach of contract.

[0142] Based on the paragraph-level hidden Markov model, the logical structure of the paragraph is detected as a progressive structure (connecting words: "first" (0.7), "second" (0.6), "last" (0.8)). The three second embedding positions are determined based on the observation probability:

[0143] Candidate 1: "First," after (probability 0.7).

[0144] Candidate 2: "Secondly," (probability 0.6).

[0145] Candidate position 3: "Finally," after (probability 0.8).

[0146] Due to the redundant check code is 1100, then the watermark information embedding operation in this paragraph is: embed Position 1 1 (insert \u200B), embedded Position 2 1 (insert \u200B), embedded Bit 3 is 0 (not inserted).

[0147] The final paragraph after embedding the watermark information becomes:

[0148] "First, Zhang San failed to fulfill his payment obligations.

[0149] Secondly, Li Si made repeated demands but to no avail.

[0150] Finally, Zhang San breached the contract.”

[0151] By encoding the target watermark information to generate an initial code, the watermark information can be embedded in the text in digital form. The encoding process can compress or transform the watermark information to accommodate the capacity limitations of the embedding bit, thereby improving the watermark's confidentiality and anti-interference capabilities. The completeness of the embedded watermark information can be flexibly selected based on the available embedding bit space and the importance of the watermark information, thereby increasing the flexibility of watermark embedding. By introducing an error correction coding mechanism to generate an error correction code and adding a redundant check code to the second embedded bit, the robustness of the watermark is enhanced. Even if the watermark information is partially damaged during transmission or storage, the redundant check code can be used to correct the error and restore it, ensuring the integrity and accuracy of the watermark information.

[0152] Figure 6 The flowchart of the chapter-level watermark embedding method according to an embodiment of the present application is schematically shown. Figure 6 As shown, it includes operations S610 to S630.

[0153] In operation 610 , a plurality of groups of reference information codes corresponding to a plurality of paragraphs are generated based on historical record information of each of a plurality of paragraphs included in each chapter of the text to be processed being added with target watermark information.

[0154] During the watermarking process, the watermark information added to each paragraph can be recorded, including the historical records of the initial code and redundant check code. After the adding is completed, the initial code and redundant check code added to each paragraph can be determined based on the historical records, that is, multiple groups of reference information codes corresponding to multiple paragraphs can be obtained.

[0155] In one example, the reference information code is intermediate information used to generate chapter codes and is related to the paragraph watermark information. Specifically, the reference information code is the watermark information added to the paragraph and is used in the subsequent generation of chapter codes to ensure the integrity of the text and the consistency of the watermark information. The reference information code can be understood as a watermark feature identifier for the paragraph, carrying the key characteristics of the paragraph watermark information and used in the subsequent generation of chapter-level codes, thereby achieving watermark protection and verification for the entire chapter.

[0156] In operation S620, chapter codes for the chapters are generated based on the multiple sets of reference information codes. For example, the chapter codes may be obtained by performing a hash calculation on the multiple sets of reference information codes of the multiple paragraphs included in each chapter.

[0157] In operation S630, at least part of the chapter code is added to the third embedding bit, for example, at the end of the chapter.

[0158] According to an embodiment of the present application, by generating chapter watermark information based on the sentence-level and paragraph-level watermark information added to the chapter paragraphs, associations at the sentence level, paragraph level, and chapter level can be established. Multi-level associated watermarks are not easy to crack and have higher security. Robustness is improved through the hierarchical association of watermarks.

[0159] Figure 7 The flowchart of the chapter code generation method according to the embodiment of the present application is schematically shown. Figure 7 As shown, the method for generating chapter codes for chapters based on multiple groups of reference information codes may further include operations S710 to S740.

[0160] In operation S710, hash operations are performed on the multiple groups of reference information codes to generate multiple groups of paragraph codes corresponding to the multiple paragraphs.

[0161] In operation S720, a structure of a hash tree used for performing a hash operation based on the plurality of groups of reference information codes is determined.

[0162] According to an embodiment of the present application, the hash tree includes multiple layers of nodes, each node representing a code value to be determined.

[0163] In operation S730, based on the multiple sets of reference information codes, the code values ​​of the nodes in each layer are determined layer by layer until the code value of the root node is determined.

[0164] According to an embodiment of the present application, the hash tree includes K layers; based on multiple groups of reference information codes, determining the coding values ​​of the nodes in each layer layer by layer includes: using the coding values ​​of the nodes in the k-1th layer to calculate the coding values ​​of the nodes in the kth layer, wherein the multiple groups of reference information codes serve as the coding values ​​of multiple 1st layer nodes; k=2,…K.

[0165] In operation S740, a chapter code is generated based on the code value of the root node.

[0166] According to an embodiment of the present application, generating a chapter code based on the code value of the root node includes: performing code bit compression on the code value of the root node to obtain the chapter code.

[0167] In one example, after generating multiple sets of paragraph codes corresponding to multiple paragraphs, a hash tree, such as a Merkle tree, is constructed. This construction process effectively integrates the watermark information of all paragraphs within a chapter. Determining the Merkle tree structure involves determining the hash tree's hierarchy and node organization. Typically, a Merkle tree is a binary tree, but other types of hash trees are also possible. The number of leaf nodes depends on the number of paragraphs. If the number of paragraphs is not a power of 2, the last level is padded with duplicate leaf nodes.

[0168] For a hash tree with K layers, starting from layer 1, multiple reference information codes are used as the encoding values ​​for the nodes in layer 1. For layer k (k = 2, ..., K), the encoding values ​​of the nodes in layer k-1 are concatenated and hashed to obtain the encoding value of the nodes in layer k. This calculation is continued layer by layer until the encoding value of the root node is obtained. Specifically, starting from the leaf nodes, the hash values ​​of the parent nodes are calculated layer by layer. For each parent node, the hash values ​​of its two child nodes are concatenated and hashed to obtain the hash value of the parent node.

[0169] For a chapter containing four paragraphs, assume that the reference information codes for each paragraph are H1, H2, H3, and H4. Calculate the hash values ​​of the multiple reference information codes. For example, generate a 128-bit hash value based on the multiple reference information codes, and take the first 32 bits to generate multiple paragraph codes H(H1), H(H2), H(H3), and H(H4) corresponding to the multiple paragraphs. These paragraph codes will serve as leaf nodes in the hash tree.

[0170] Take the 3-layer hash tree structure as an example (K=3):

[0171] Layer 1: Paragraph 1 paragraph code H (H1), paragraph 2 paragraph code H (H2), paragraph 3 paragraph code H (H3), paragraph 4 paragraph code H (H4).

[0172] Layer 2: Calculates the hash of the combined nodes in layer 1.

[0173] Layer 3 (root node): Calculates the hash of the combined nodes in layer 2.

[0174] The encoding values ​​of the nodes in the first layer are used to calculate the encoding values ​​of the nodes in the second layer. The encoding value of node 1 in the second layer is: H(H(H1)+H(H2)); the encoding value of node 2 in the second layer is: H(H(H3)+H(H4)). The encoding value of the nodes in the second layer is used to calculate the encoding value of the third layer (root node): the encoding value of the root node is: H(H(H(H1)+H(H2))+H(H(H3)+H(H4))). Finally, the hash value of the root node is calculated. The hash value of the root node represents a summary of the content of the entire chapter. The root hash value is unique, and any slight change in any paragraph will cause the root hash value to change significantly. After obtaining the encoding value of the root node, the encoding value of the root node can be directly used as the chapter code. As a preference, in order to improve embedding efficiency and reduce the length of the code, the encoding value of the root node can be compressed to obtain a shorter binary bit stream as the chapter code.

[0175] After obtaining the chapter code, the compressed chapter code is converted into invisible characters, such as zero-width characters, and embedded into the third embedding position. The specific embedding process is the same as the aforementioned sentence-level and paragraph-level watermark information embedding process, which will not be repeated here.

[0176] According to the embodiments of this application, by generating chapter codes based on a hash tree, the information of each paragraph can be implicitly stored in the text, eliminating reliance on external storage, ultimately achieving efficient and reliable copyright protection with semantic losslessness and structural integrity. Furthermore, because the information generated by the hash tree is self-verifying, it also facilitates subsequent verification of watermark information.

[0177] By generating reference information codes based on the watermarking history of a paragraph and then generating chapter codes based on these reference information codes, watermark information can be integrated and identified at the chapter level, further enhancing copyright protection and traceability of the text. Adding at least part of the chapter code to the third embedding position allows watermark information to be embedded at the chapter level, achieving identification and protection of the watermark on a wider scale.

[0178] Figure 8 The following is a schematic diagram showing a watermark information embedding process according to an embodiment of the present application.

[0179] like Figure 8As shown in an example, a watermark is embedded into text (e.g., generated by an LLM model) via a processor. The processor performs semantic structure bit detection, which may involve invoking a trained multi-level Hidden Markov Model (HMM), such as one comprising a first detection model, a second detection model, and a third detection model, to perform semantic structure bit detection. The processor then determines a final embedding bit from the multiple candidate embedding bits detected, including a first embedding bit for a target word in a sentence, a second embedding bit for a target sentence in a paragraph, and a third embedding bit for a target paragraph in a chapter.

[0180] In the watermark embedding stage, the watermark information input by the user is first obtained, and the watermark information is converted into an M-bit binary bit stream. The M+N bit stream is generated through RS encoding. The M-bit binary bit stream is the initial code, and the N-bit binary bit stream is the redundant check code.

[0181] The text is segmented to obtain a word sequence, and the word sequence is input into a first detection model to determine the grammatical structure to which multiple words in the sentence of the text to be processed belong. Based on the grammatical structure to which the multiple words belong, the position of the target word belonging to the predetermined grammatical structure is used as the first embedded bit, and zero-width spaces are inserted into the first embedded bit in sequence according to an M-bit binary stream, 1 for insertion and 0 for non-insertion.

[0182] Through the second detection model, the logical structures to which multiple sentences in the paragraph of the text to be processed belong are determined, and based on the logical structures to which the multiple sentences belong, the position of the target sentence belonging to the predetermined logical structure is used as the second embedded position; the corresponding zero-width space is inserted into the second embedded position according to the redundant check code.

[0183] During the embedding process, the watermark information inserted in all the first and second embedding bits in the paragraph is recorded in sequence to form a paragraph binary sequence. The paragraph binary bit stream is hashed to obtain the paragraph hash value. The paragraph hash is used as the leaf node, and the parent node hash is calculated layer by layer to finally obtain the root hash, which is then compressed into a low-order bit stream. The chapter structure to which multiple paragraphs in the chapter of the text to be processed belong is determined according to the third detection model. Based on the chapter structure to which multiple paragraphs belong, the position of the target paragraph belonging to the predetermined chapter structure is used as the third embedding bit. A zero-width space is inserted in the third embedding bit to complete the hierarchical watermark embedding of the text structure feature.

[0184] Furthermore, while inserting a zero-width space at the third embedding position, a separator (such as a carriage return key) can be inserted synchronously to separate each code to distinguish the type of inserted code, thereby facilitating subsequent restoration of the chapter watermark code.

[0185] For example, if the watermark code inserted at the end of a chapter is: 01011, the characters inserted are: \ru2008B\r\ru2008B\ru2008B. In this way, when restoring the chapter watermark later, "u2008B" is restored to the code "1". Because the "\r" is a character delimiter, the code "0" that was not inserted can be restored.

[0186] The text processing method provided in the embodiment of the present application also includes a watermark detection process. Figure 9 This section describes the watermark detection process.

[0187] Figure 9 The flowchart of the watermark information restoration method according to the embodiment of the present application is schematically shown. Figure 9 As shown, it includes operations S910 to S930.

[0188] In operation S910, watermark information restoration is performed on each of the multiple paragraphs included in the chapter of the target text to which the target watermark information is added, to obtain multiple groups of restored watermark information codes corresponding to the multiple paragraphs.

[0189] In operation S920, the multiple sets of restored watermark information codes are processed to obtain restored codes corresponding to the chapters. For example, the same method as that used to generate chapter codes can be used to perform hash operations on the multiple sets of restored watermark information codes to generate multiple sets of restored paragraph codes corresponding to the multiple paragraphs. Based on the restored paragraph codes, the restored codes corresponding to the chapters are obtained.

[0190] In operation S930 , a comparison and verification is performed between the chapter code extracted from the chapter including the plurality of paragraphs and the restored code.

[0191] In one example, the target text in operation S910 is the text to which the target watermark information is added. The watermark information embedded in each paragraph is restored. Specifically, a hidden Markov model is used to detect the text structure, identify grammar, logic, and text structure, and determine the embedding position (such as the adverbial position of the adverbial structure, the position after the conjunction of the progressive structure, etc.). The watermark information fragment at the embedding position is extracted. The watermark information fragments of all paragraphs are integrated to generate the restoration code of the chapter. The chapter code embedded in the document is extracted. The restoration code and the chapter code are compared to determine whether they are consistent to verify the integrity of the document. The hash value of the restoration code is calculated and compared with the pre-embedded chapter code extracted from the embedded position of the chapter for verification. If the two are consistent, the verification passes, indicating that the chapter has not been tampered with.

[0192] According to the embodiments of the present application, unlike conventional key verification methods, the method of the embodiments of the present application uses a hash tree to calculate the chapter code for storage when the watermark is embedded. This method can form a self-verification mechanism in the subsequent watermark verification process. The watermark information can be obtained through hash operations, and there is no need for additional storage and transmission of secret keys, which improves the verification efficiency and is particularly suitable for offline verification scenarios.

[0193] According to an embodiment of the present application, a method for restoring watermark information for each of multiple paragraphs included in a chapter of a target text to which target watermark information is added includes:

[0194] First, by performing text structure detection on the target text, at least one detection embedding position where watermark information can be added is determined in the target text;

[0195] Afterwards, extracting the initial codes and / or redundancy check codes added to each of the plurality of paragraphs;

[0196] Then, based on the initial codes and / or redundant check codes added to each of the multiple paragraphs and the position information of at least one detection embedding bit, the watermark information of each paragraph is restored to obtain multiple groups of restored watermark information codes corresponding to the multiple paragraphs.

[0197] In one example of text structure detection for the target text, the target text is subjected to text structure detection based on the aforementioned multi-level hidden Markov model to obtain the structural attribution probability of each word, sentence, and paragraph, and the detection embedding position is determined based on the probability. That is, the text is detected using a sentence-level hidden Markov model to identify the grammatical structure of each word; the text is detected using a paragraph-level hidden Markov model to identify the logical structure of each sentence; and the text is detected using a chapter-level hidden Markov model to identify the chapter structure of each paragraph. Determine at least one detection embedding position in the target text where watermark information can be added. The specific process for determining the detection embedding position is the same as the process for determining the watermark embedding position, and will not be repeated here.

[0198] Among them, the initial codes and / or redundant check codes added to each of the multiple paragraphs are extracted. Since the initial codes and / or redundant check codes are added to the target text in the form of predetermined characters, a character recognition tool can be used to extract all the added predetermined characters (such as zero-width spaces) in the text.

[0199] According to an embodiment of the present application, based on predetermined characters contained in each of the multiple paragraphs and position information of at least one detection embedding bit, restoring watermark information for each paragraph includes the following method:

[0200] First, based on the initial code added to the paragraph and the position information of at least one detection embedding bit, the first watermark information of the paragraph is restored to obtain an initial restored information code corresponding to the paragraph;

[0201] Afterwards, the second watermark information of the paragraph is restored based on the redundant check code added to the paragraph and the position information of at least one detection embedded bit to obtain an error correction restoration information code corresponding to the paragraph;

[0202] Then, the initial restored information code is corrected using the error correction restoration information code, and the restored watermark information code corresponding to the paragraph is obtained based on the initial restored information code after error correction.

[0203] During the watermark restoration process, error correction can be performed based on the extracted checksum (error correction mechanism of the RS encoding algorithm). If the watermark is partially damaged, it can be restored, ensuring the normal watermark verification and improving information reliability.

[0204] According to the embodiment of the present application, further, the method of restoring the first watermark information for each paragraph and the method of restoring the second watermark information for each paragraph are the same, and both can adopt the following method:

[0205] Extracting predetermined characters contained in each of the multiple paragraphs (initial encoding and / or redundant check codes are added to the target text in the form of predetermined characters); restoring watermark information for each paragraph based on the predetermined characters contained in each of the multiple paragraphs and position information of at least one detection embedding bit.

[0206] When embedding watermark information, the j-bit target code in the J-bit watermark code is converted into an invisible j-bit zero-width character, and the j target embedding bits are added to the multiple embedding bits. That is, the zero-width characters inserted into the paragraph only represent a portion of the code in the initial code and / or the redundant check code. For example, when embedding watermark information in binary code (such as 0110000…), the "1" in the code is converted into a zero-width character and embedded in a specific position, while the "0" in the code is not embedded. Therefore, after extracting the predetermined characters contained in each of the multiple paragraphs, the resulting code is only a partial code, all of which are the character "1", such as "111111…", and not the original complete watermark "0110000…". Therefore, it is necessary to combine the position information of at least one detection embedding bit to restore the watermark information for each paragraph.

[0207] For example, it is detected that a paragraph "Scientific and technological innovation has made great progress" has three embedded bits: the first bit is "scientific and technological innovation", the second bit is "achieved", and the third bit is "great progress"; when extracting the watermark, only "1" is extracted at the second bit "achieved"; then the complete watermark of this paragraph should be "010" after restoration, that is, the second bit corresponds to "1", and the first and third bits correspond to "0".

[0208] According to an embodiment of the present application, the predetermined characters contained in each of the multiple paragraphs include: a first predetermined character added to at least one detection embedded position, and a second predetermined character not in at least one detection embedded position. The first predetermined character and the second predetermined character are the same characters, for example, both are zero-width spaces, the difference is that the first predetermined character is added to the detection embedded position, and the second predetermined character is added to the non-detection embedded position. This means that the first predetermined character added to the detection embedded position belongs to the watermark information, and the second predetermined character added to the non-detection embedded position does not belong to the watermark information, and may have been tampered with and added by the user. In order to restore the correct watermark information, the second predetermined character added to the non-detection embedded position needs to be filtered out.

[0209] Specifically, restoring watermark information for each paragraph based on predetermined characters contained in each of the plurality of paragraphs and the position information of at least one detection embedding position includes restoring watermark information for each paragraph based on a first predetermined character and the position information of at least one detection embedding position. That is, only the first predetermined character added to the detection embedding position is extracted for processing, and the second predetermined character added to the non-detection embedding position is not processed.

[0210] In one example, the predetermined character may be, for example, a zero-width character. The initial encoding and / or redundant check code is added to the target text in the form of a zero-width character. The text is scanned to determine the positions of all zero-width characters. Since the text may be tampered with during dissemination or use, there are cases where zero-width characters appear in non-embedded positions. For example, in a paragraph, zero-width characters are extracted and located at embedded position 1, embedded position 3, and non-embedded position 2, respectively. At this time, the zero-width characters located at embedded position 1 and embedded position 3 are the first predetermined characters, and the zero-width character located at non-embedded position 2 is the second predetermined character. Therefore, it is only necessary to pay attention to the first predetermined character of the embedded position, and combine the position information of the embedded position to determine which characters belong to the initial encoding or redundant check code. Based on the zero-width character corresponding to the initial encoding and the position of the detected embedded position, the watermark information code is restored. Watermark information restoration is performed based on the first predetermined character and the embedded position information, avoiding the interference of non-embedded position characters, filtering illegal characters, and improving the accuracy and reliability of watermark restoration information.

[0211] Figure 10 The whole process of watermark embedding and watermark detection according to an embodiment of the present application is schematically shown.

[0212] The following, combined Figure 10 , which illustrates the method of watermark information embedding and restoration process.

[0213] like Figure 10As shown in the figure, during the watermark embedding phase, after the text is generated, text structure detection is performed on the text to obtain text structure information of the text to be processed. For example, positions with an observation probability greater than a threshold Q are detected using a multi-level hidden Markov model as candidate embedding positions. Positions with an observation probability less than a preset threshold Q are filtered out and skipped. The candidate embedding positions can also be further screened to determine the final target embedding position. After the embedding position is marked, the watermark information is embedded, that is, the watermark information is converted into a predetermined character and inserted into the embedding position.

[0214] In the watermark detection stage, the target text to be embedded with the watermark is firstly extracted from the predetermined characters (such as zero-width characters), and the multi-level hidden Markov model is used to detect the position where the observation probability is greater than the threshold Q as the detection embedding position, and the predetermined characters at this position are retained, and the zero-width characters added to the non-detection embedding position are filtered out (interference signals are filtered out) to obtain the correct watermark encoding information.

[0215] Below, taking the example of a user-designed watermark "©2024 Text AI Generated" embedded in a professional document generated by a large model, the watermark embedding and watermark detection methods of the embodiments of the present application are illustrated.

[0216] Exemplarily, the method of embedding watermark information is as follows:

[0217] Watermark preprocessing:

[0218] The original watermark "©2024 Text AI Generated" entered by the user through the client is converted into a binary bit stream (88 bits in total, using the first 16 bits as an example: 0010101101001001) after being encoded in UTF-8. RS(20,16) encoding (M=16 information bits, N=4 redundant bits) is used to generate a 20-bit bit stream.

[0219] Sentence-level embeddings:

[0220] For the sentence in the "Facts" section of the target paragraph: "Zhang San failed to fulfill his contractual payment obligations within a reasonable period and, despite repeated reminders from Li Si, still failed to pay," we perform grammatical analysis, segment the words, and then use sentence-level hidden Markov model decoding. We select positions with an observation probability greater than 0.6 as candidate embedding positions.

[0221] For example, "within a reasonable period" is a state-in-a-media structure with an observation probability of 0.8 (≥0.6), and is marked as candidate position 1. The determination of other legal embedding positions is the same as this and will not be repeated here.

[0222] Insert ZWS (\u200B) in candidate position 1 (between "within" and "within a reasonable period") to represent a 1, and then insert the remaining information positions in order. The final sentence becomes: "Zhang San failed to fulfill his payment obligations under the contract within a reasonable period of \u200B, and despite repeated reminders from Li Si, he still did not pay."

[0223] Paragraph-level embeddings:

[0224] For the entire target paragraph "Fact Finding" (for example, containing three sentences), Hidden Markov Model decoding is used to select positions with an observation probability greater than 0.6 as redundant embedding positions.

[0225] For example, the probability of the word "first" in the sentence "First, Zhang San failed to fulfill his payment obligation" is 0.7 (≥0.6), and it is marked as candidate position 1. The determination of other legal embedding positions is similar and will not be repeated here.

[0226] Insert ZWS at candidate position 1 to represent redundant bit 1, and the final paragraph becomes: "First\u200B, Zhang San failed to fulfill his payment obligation..."

[0227] Chapter-level embedding:

[0228] First, build a Merkle tree. Calculate the 128-bit hash value of each paragraph's bitstream (take the first 32 bits as an example), generate leaf nodes H(P1), H(P2), and H(P3), and calculate the parent node hash layer by layer to obtain a root hash, such as H_root = a1b2c3d4. This is compressed into an 8-bit bitstream (10110011).

[0229] The entire text conforms to the structure of "fact determination, determination basis, and determination results." Based on chapter-level hidden Markov model decoding, ZWS and \r delimiters are inserted at the end of the "Determination Results" section at legal locations (with an observation probability of 0.95 ≥ 0.6). For example, at the end of the section: "Determination results are as follows: payment of liquidated damages of 100,000 yuan" is inserted with \u200B\r...

[0230] Exemplarily, the watermark detection method is as follows:

[0231] Invisible character extraction: Use the character detection tool to scan the text and locate the ZWS (\u200B) position (1 at the sentence level, 1 at the paragraph level, 1 at the chapter level) and the \r delimiter (1).

[0232] Sentence-level detection Embedding position detection: The word sequence is input into the sentence-level hidden Markov model, and it is confirmed that "within a reasonable period" is an adverbial-predicate structure with an observation probability of 0.8 (≥0.6). The ZWS is retained.

[0233] Paragraph-level detection Embedding position detection: The paragraph logical structure is the total score (P1), the connecting word position probability of "First\u200B" is 0.7 (≥0.6), and the ZWS is retained.

[0234] Chapter-level detection embedded bit detection: the chapter order is "fact determination", "determination basis", "determination result", the observation probability is 0.95 (≥0.6), and the root hash bit stream is retained.

[0235] Layered bitstream extraction and error correction: Detection of embedded bits using sentence-level hidden Markov model and sentence-level ZWS to recover 16 bits (Example: 0010101101001001); 4 redundant bits are recovered based on the position information of the detection embedded bits determined by the paragraph-level hidden Markov model and the paragraph-level ZWS (Example 1010). 2-bit errors (such as bit stream errors caused by deletion) are corrected through RS decoding and restored to the original .

[0236] After recovery Reassemble the bitstreams into each paragraph (the splicing method is similar to the watermark embedding process), calculate the hash value to build the Merkle tree, and calculate and obtain the new root hash ,Compare the new root hash with the root hash bit stream embedded at the end of the chapter paragraph. If the comparison is consistent, the watermark verification passes, indicating that the watermark content has not been tampered with.

[0237] The text processing method provided in the embodiment of the present application performs text structure detection on the text to be processed through a pre-trained hidden Markov model, and can accurately identify the grammatical structure, logical structure and paragraph structure in the text. Determining the embedding position for adding watermark information based on the text structure information can ensure that the watermark information is embedded in a relatively hidden position in the text and has less impact on semantics. Not only does it improve the concealment of the watermark, it also reduces the interference of the watermark on the readability of the text, while ensuring that the watermark information can be stably embedded in the text. By performing multi-level embedding at the sentence level, paragraph level and chapter level, the copyright information of the text can be more comprehensively protected, and the watermark's anti-attack ability is enhanced. The multi-level embedding and verification mechanism improves the concealment and robustness of the watermark, making the watermark more difficult to detect and remove, and also facilitates the rapid and accurate verification of the integrity of the text and the authenticity of the copyright information during detection. It effectively solves the copyright protection problem of text generated by large language models: it ensures the concealment of watermarks, enhances robustness through hierarchical association and RS encoding, resists local deletion and modification, and avoids external storage dependence based on self-checking hash trees. Ultimately, it achieves efficient and reliable copyright protection with semantic losslessness and structural integrity, providing a highly adaptable technical solution for content traceability and attribution verification of high-value texts.

[0238] Based on the above text processing method, this application also provides a text processing device. Figure 11 The device is described in detail.

[0239] Figure 11 The structural block diagram of the text processing device according to an embodiment of the present application is schematically shown.

[0240] like Figure 11 As shown, the text processing device 1100 of this embodiment includes a detection module 1110 , a determination module 1120 and an adding module 1130 .

[0241] Detection module 1110 is configured to perform text structure detection on the text to be processed to obtain text structure information of the text to be processed, wherein the text structure information includes at least one of the following: the grammatical structure to which multiple words in a sentence of the text to be processed belong, the logical structure to which multiple sentences in a paragraph of the text to be processed belong, and the chapter structure to which multiple paragraphs in a chapter of the text to be processed belong. In one embodiment, detection module 1110 can be configured to perform operation S210 described above, which will not be further described here.

[0242] The determination module 1120 is configured to determine at least one embedding position for adding watermark information in the text to be processed based on the correlation between the text structure information and the contextual semantics of the text to be processed. In one embodiment, the determination module 1120 can be configured to perform the operation S220 described above, which will not be described in detail here.

[0243] The adding module 1130 is configured to add the target watermark information to at least one embedding position. In one embodiment, the adding module 1130 may be configured to execute the operation S230 described above, which will not be described in detail here.

[0244] According to an embodiment of the present application, the determination module includes a first determination submodule, a second determination submodule, and a third determination submodule.

[0245] The first determination submodule is configured to, based on the grammatical structures to which the multiple words belong, select as a first embedding position the position of a target word that belongs to a predetermined grammatical structure and whose contextual semantic relevance meets a predetermined condition.

[0246] The second determination submodule is configured to, based on the logical structures to which the multiple sentences belong, use as a second embedding position the position of the target sentence that belongs to the predetermined logical structure and whose contextual semantic relevance meets the predetermined conditions.

[0247] The third determination submodule is configured to, based on the chapter structures to which the multiple paragraphs belong, select the position of the target paragraph that belongs to the predetermined chapter structure and whose contextual semantic relevance meets predetermined conditions as the third embedding position.

[0248] According to an embodiment of the present application, the adding module includes a first generating submodule and a first adding submodule.

[0249] The first generating submodule is used to encode the target watermark information in a predetermined format to generate an M-bit initial code, where M is a positive integer.

[0250] The first adding submodule is configured to add at least part of the M-bit initial code to the first embedded bit.

[0251] According to an embodiment of the present application, the adding module further includes a second generating submodule and a second adding submodule.

[0252] The second generation submodule is used to encode the M-bit initial code generated based on the target watermark information based on the error correction mechanism to generate an (M+N)-bit error correction code, wherein the error correction code includes the M-bit initial code and an N-bit redundant check code, where N is a positive integer; the redundant check code is used to: in the case where the initial code is modified, the modified initial code is restored;

[0253] The second adding submodule is configured to add at least part of the N-bit redundant check code to the second embedded bit.

[0254] According to an embodiment of the present application, the adding module further includes a third generating submodule, a fourth generating submodule and a third adding submodule.

[0255] The third generation submodule is configured to generate a plurality of groups of reference information codes corresponding to the plurality of paragraphs according to the historical record information of adding target watermark information to the plurality of paragraphs included in each chapter of the text to be processed;

[0256] A fourth generating submodule, configured to generate chapter codes for the chapters based on the multiple sets of reference information codes;

[0257] The third adding submodule is configured to add at least a portion of the chapter code to the third embedding position.

[0258] According to an embodiment of the present application, the fourth generation submodule includes a first generation unit and a second generation unit.

[0259] The first generating unit is configured to perform hash operations on the multiple groups of reference information codes respectively to generate multiple groups of paragraph codes corresponding to the multiple paragraphs.

[0260] The second generating unit is configured to generate chapter codes for the chapters based on the multiple sets of paragraph codes.

[0261] According to an embodiment of the present application, the second generating unit includes a first determining subunit, a second determining subunit and a generating subunit.

[0262] A first determining subunit is configured to determine a structure of a hash tree used for performing a hash operation based on the multiple sets of reference information codes, wherein the hash tree includes multiple layers of nodes, each node representing a code value to be determined;

[0263] A second determining subunit is configured to determine the code value of each layer node layer by layer based on the multiple groups of reference information codes until the code value of the root node is determined;

[0264] The generation subunit is used to generate a chapter code based on the code value of the root node.

[0265] According to an embodiment of the present application, the second determination subunit is also used to calculate the coding value of the kth layer node using the coding value of the k-1th layer node, wherein multiple groups of reference information codes serve as the coding values ​​of multiple 1st layer nodes; k=2,…K.

[0266] According to an embodiment of the present application, the generation subunit is further configured to perform code bit compression on the encoding value of the root node to obtain a chapter code.

[0267] According to an embodiment of the present application, the device further includes: a restoration module, a restoration code generation module and a comparison and verification module.

[0268] a restoration module, configured to restore watermark information of each of the multiple paragraphs included in the chapter of the target text to which the target watermark information is added, and obtain multiple groups of restored watermark information codes corresponding to the multiple paragraphs;

[0269] A restoration code generation module is used to process multiple groups of restoration watermark information codes to obtain restoration codes corresponding to the chapters;

[0270] The comparison and verification module is used to compare and verify the chapter code extracted from a chapter containing multiple paragraphs and the restored code.

[0271] According to an embodiment of the present application, the restoration module includes a fourth determination submodule, an extraction submodule and a restoration submodule.

[0272] a fourth determining submodule, configured to determine at least one detection embedding position where watermark information can be added in the target text by performing text structure detection on the target text;

[0273] An extraction submodule, configured to extract the initial codes and / or redundant check codes added to each of the plurality of paragraphs;

[0274] The restoration submodule is used to restore the watermark information of each paragraph based on the initial code and / or redundant check code added to each of the multiple paragraphs and the position information of at least one detection embedding bit, so as to obtain multiple groups of restored watermark information codes corresponding to the multiple paragraphs.

[0275] According to an embodiment of the present application, the restoration submodule includes a first watermark information restoration unit, a second watermark information restoration unit and an error correction unit.

[0276] a first watermark information restoration unit, configured to restore the first watermark information of the paragraph according to the initial code added to the paragraph and the position information of at least one detection embedding bit, to obtain an initial restoration information code corresponding to the paragraph;

[0277] a second watermark information restoration unit, configured to restore the second watermark information of the paragraph according to the redundant check code added to the paragraph and the position information of at least one detection embedding bit, to obtain an error correction restoration information code corresponding to the paragraph;

[0278] The error correction unit is used to correct the initial restored information code using the error correction restoration information code, and obtain the restored watermark information code corresponding to the paragraph based on the initial restored information code after error correction.

[0279] According to an embodiment of the present application, the extraction submodule is further configured to extract predetermined characters contained in each of the plurality of paragraphs. The predetermined characters contained in each of the plurality of paragraphs include: a first predetermined character added to at least one detection embedding position, and a second predetermined character not in at least one detection embedding position.

[0280] According to an embodiment of the present application, the restoration submodule is further configured to restore watermark information for each paragraph based on predetermined characters contained in each of the multiple paragraphs and position information of at least one detection embedding position.

[0281] According to an embodiment of the present application, the restoration submodule is further configured to restore watermark information for each paragraph based on the first predetermined character and position information of at least one detection embedding position.

[0282] According to an embodiment of the present application, the adding module is also used to generate a J-bit watermark code of a predetermined format based on the target watermark information, where J is a positive integer; convert the j-bit target code in the J-bit watermark code into an invisible j-bit predetermined character, where the predetermined character includes a zero-width character; and sequentially add the j-bit predetermined character to j target embedded bits in a plurality of embedded bits, where j is a positive integer less than or equal to J.

[0283] According to an embodiment of the present application, the detection module includes a first detection submodule, a second detection submodule, and a third detection submodule.

[0284] A first detection submodule is configured to perform a first detection on a plurality of words based on a first detection model, obtain a plurality of first probabilities that each word belongs to a plurality of candidate grammatical structures, and determine the grammatical structure to which the word belongs based on the plurality of first probabilities;

[0285] a second detection submodule, configured to perform a second detection on the plurality of sentences based on a second detection model, obtain a plurality of second probabilities that each sentence belongs to a plurality of candidate logical structures, and determine the logical structure to which the sentence belongs based on the plurality of second probabilities;

[0286] The third detection submodule is configured to perform a third detection on the plurality of paragraphs based on a third detection model, obtain a plurality of third probabilities that each paragraph belongs to a plurality of candidate chapter structures, and determine the chapter structure to which the paragraph belongs based on the plurality of third probabilities.

[0287] According to an embodiment of the present application, the first detection submodule is further configured to determine, by using the first state transition information and the first observation probability information, a plurality of first probabilities that each word belongs to a plurality of candidate grammatical structures.

[0288] According to an embodiment of the present application, the device further includes: a first acquisition module, a first determination module and a first training module.

[0289] The first acquisition module is used to acquire a training text and perform word segmentation on a plurality of sample sentences in the training text to obtain a plurality of sample words;

[0290] A first determination module is used to determine the grammatical structure to which each of the plurality of sample words belongs, so as to generate a first label of the training text;

[0291] The first training module is used to train a basic model using multiple sample words and a first label to obtain a first detection model, wherein the model parameters of the first detection model include first state transition information and first observation probability information, wherein the first state transition information is used to represent the probability of transitioning from a first candidate grammatical structure to a second candidate grammatical structure, and the first observation probability information is used to represent the probability of each of multiple target sample words belonging to the second candidate grammatical structure, and the first candidate grammatical structure and the second candidate grammatical structure belong to at least one of the multiple grammatical structures to which the multiple sample words belong.

[0292] According to an embodiment of the present application, the device further includes: a second acquisition module, a second determination module and a second training module.

[0293] The second acquisition module is used to acquire a training text and perform sentence segmentation processing on multiple sample paragraphs in the training text to obtain multiple sample sentences;

[0294] A second determination module is used to determine the logical structure to which each of the plurality of sample sentences belongs, so as to generate a second label of the training text;

[0295] A second training module is used to train a basic model using multiple sample sentences and second labels to obtain a second detection model, wherein the model parameters of the second detection model include second state transition information and second observation probability information, wherein the second state transition information is used to characterize the probability of transitioning from the first candidate logical structure to the second candidate logical structure, and the second observation probability information is used to characterize the probability of each of multiple target sample sentences belonging to the second candidate logical structure, and the first candidate logical structure and the second candidate logical structure belong to at least one of the multiple logical structures to which the multiple sample sentences belong.

[0296] According to an embodiment of the present application, the device further includes: a third acquisition module, a third determination module and a third training module.

[0297] The third acquisition module is used to acquire the training text and segment the training text to obtain multiple sample paragraphs;

[0298] A third determination module is used to determine the chapter structure to which each of the multiple sample paragraphs belongs, so as to generate a third label of the training text;

[0299] The third training module is used to train the basic model using multiple sample paragraphs and third labels to obtain a third detection model, wherein the model parameters of the third detection model include third state transition information and third observation probability information, wherein the third state transition information is used to characterize the probability of transitioning from the first candidate chapter structure to the second candidate chapter structure, and the third observation probability information is used to characterize the probability of each of the multiple target sample paragraphs belonging to the second candidate chapter structure, and the first candidate chapter structure and the second candidate chapter structure belong to at least one of the multiple chapter structures to which the multiple sample paragraphs belong.

[0300] According to embodiments of the present application, any multiple modules among the detection module 1110, determination module 1120, and addition module 1130 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present application, at least one of the detection module 1110, determination module 1120, and addition module 1130 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of these. Alternatively, at least one of the detection module 1110, determination module 1120, and addition module 1130 may be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality.

[0301] Figure 12 A block diagram of an electronic device suitable for implementing a text processing method according to an embodiment of the present application is schematically shown.

[0302] like Figure 12 As shown, the electronic device 1200 according to an embodiment of the present application includes a processor 1201, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1202 or a program loaded from a storage unit 1208 into a random access memory (RAM) 1203. The processor 1201 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1201 may also include onboard memory for caching purposes. The processor 1201 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present application.

[0303] Various programs and data required for the operation of the electronic device 1200 are stored in the RAM 1203. The processor 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. The processor 1201 performs various operations of the method flow according to the embodiment of the present application by executing the programs in the ROM 1202 and / or the RAM 1203. It should be noted that the programs may also be stored in one or more memories other than the ROM 1202 and the RAM 1203. The processor 1201 may also perform various operations of the method flow according to the embodiment of the present application by executing the programs stored in the one or more memories.

[0304] According to an embodiment of the present application, electronic device 1200 may further include an input / output (I / O) interface 1205, which is also connected to bus 1204. Electronic device 1200 may also include one or more of the following components connected to I / O interface 1205: an input section 1206 including a keyboard, mouse, etc.; an output section 1207 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 1208 including a hard disk; and a communication section 1209 including a network interface card such as a LAN card or modem. Communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to I / O interface 1205 as needed. Removable media 1211, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 1210 as needed, so that computer programs read from the removable media can be installed into storage section 1208 as needed.

[0305] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of this application is implemented.

[0306] According to an embodiment of the present application, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, a computer-readable storage medium may include the ROM 1202 and / or RAM 1203 described above, and / or one or more memories other than ROM 1202 and RAM 1203.

[0307] The embodiments of the present application also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the text processing method provided in the embodiments of the present application.

[0308] When the computer program is executed by the processor, the above functions defined in the system / device of the embodiment of the present application are performed. According to the embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0309] In one embodiment, the computer program may be stored on a tangible storage medium, such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal over a network medium, downloaded and installed via a communication component, and / or installed from a removable medium. The program code contained in the computer program may be transmitted using any suitable network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0310] In such an embodiment, the computer program can be downloaded and installed from a network via the communication portion, and / or installed from a removable medium. When the computer program is executed by the processor, the above-mentioned functions defined in the system of the embodiment of the present application are performed. According to the embodiment of the present application, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.

[0311] According to an embodiment of the present application, the program code for executing the computer program provided by the embodiment of the present application can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0312] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0313] Those skilled in the art will appreciate that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.

[0314] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present application, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present application.

Claims

1. A text processing method, characterized in that: The method comprises: Performing text structure detection on the text to be processed to obtain text structure information of the text to be processed, wherein the text structure information includes at least one of the following: a grammatical structure to which multiple words in a sentence of the text to be processed belong, a logical structure to which multiple sentences in a paragraph of the text to be processed belong, and a chapter structure to which multiple paragraphs in a chapter of the text to be processed belong; Determining at least one embedding position for adding watermark information in the text to be processed based on the correlation between the text structure information and the context semantics in the text to be processed; Adding the target watermark information to be added to the at least one embedding position includes: generating multiple groups of reference information codes corresponding to multiple paragraphs based on historical record information of adding target watermark information to multiple paragraphs included in each chapter of the text to be processed; generating chapter codes for the chapters based on the multiple groups of reference information codes; adding at least part of the chapter codes to the third embedding position; wherein generating the chapter codes for the chapters based on the multiple groups of reference information codes includes: performing hash operations on the multiple groups of reference information codes respectively to generate multiple groups of paragraph codes corresponding to multiple paragraphs, and generating the chapter codes for the chapters based on the multiple groups of paragraph codes; wherein the third embedding position is the position of the target paragraph that belongs to a predetermined chapter structure and whose contextual semantic relevance meets predetermined conditions.

2. The method according to claim 1, characterized in that Determining, based on the correlation between the text structure information and the contextual semantics in the to-be-processed text, at least one embedding bit for adding watermark information in the to-be-processed text comprises at least one of the following: Based on the grammatical structures to which the multiple words belong, the position of the target word belonging to the predetermined grammatical structure and whose contextual semantic relevance meets the predetermined condition is used as the first embedding position; Based on the logical structures to which the multiple sentences belong, the position of the target sentence belonging to the predetermined logical structure and whose context semantic relevance meets the predetermined condition is used as the second embedding position.

3. The method according to claim 2, characterized in that Adding the target watermark information to be added to the at least one embedded position further comprises: Encoding the target watermark information in a predetermined format to generate an M-bit initial code, where M is a positive integer; At least part of the M-bit initial code is added to the first embedded bits.

4. The method according to claim 3, characterized in that Adding the target watermark information to be added to the at least one embedded position further comprises: performing an error correction mechanism on the M-bit initial code generated based on the target watermark information to generate an (M+N)-bit error correction code, wherein the error correction code includes the M-bit initial code and an N-bit redundant check code, where N is a positive integer; the redundant check code is used to restore the modified initial code when the initial code is modified; At least part of the N-bit redundant check code is added to the second embedded bit.

5. The method according to claim 1, wherein Generating a chapter code for the chapter based on the multiple sets of paragraph codes includes: Determining a structure of a hash tree used for performing a hash operation based on the multiple groups of reference information codes, wherein the hash tree includes multiple layers of nodes, each node representing a code value to be determined; Based on the multiple sets of paragraph codes, determining the code value of each layer of nodes layer by layer until the code value of the root node is determined; The chapter code is generated based on the code value of the root node.

6. The method according to claim 1, characterized in that The method further comprises: Restoring watermark information on each of the multiple paragraphs included in the chapter of the target text to which the target watermark information is added, and obtaining multiple groups of restored watermark information codes corresponding to the multiple paragraphs; Processing the multiple groups of restored watermark information codes to obtain restored codes corresponding to the chapters; A comparison and verification is performed between the chapter code extracted from the chapter containing the multiple paragraphs and the restored code.

7. The method according to claim 6, characterized in that Restoring watermark information for each of the multiple paragraphs included in the chapter of the target text to which the target watermark information is added includes: Determining at least one detection embedding position where watermark information can be added in the target text by performing text structure detection on the target text; extracting the initial codes and / or redundancy check codes added to each of the plurality of paragraphs; According to the initial codes and / or redundant check codes added to each of the multiple paragraphs, and the position information of the at least one detection embedding bit, watermark information is restored for each of the paragraphs to obtain multiple groups of restored watermark information codes corresponding to the multiple paragraphs.

8. The method according to claim 7, characterized in that Restoring the watermark information of each paragraph includes: Restoring the first watermark information of the paragraph according to the initial code added to the paragraph and the position information of the at least one detection embedding bit to obtain an initial restoration information code corresponding to the paragraph; Restoring the second watermark information of the paragraph according to the redundant check code added to the paragraph and the position information of the at least one detection embedded bit to obtain an error correction restoration information code corresponding to the paragraph; The error-correcting restoration information code is used to correct the initial restoration information code, and the restoration watermark information code corresponding to the paragraph is obtained based on the error-corrected initial restoration information code.

9. The method according to claim 7, characterized in that The initial code and / or the redundant check code are added to the target text in the form of predetermined characters; Extracting the initial codes and / or the redundant check codes added to each of the plurality of paragraphs includes: extracting the predetermined characters contained in each of the plurality of paragraphs; Restoring the watermark information for each of the paragraphs includes: restoring the watermark information for each of the paragraphs based on the predetermined characters contained in each of the plurality of paragraphs and the position information of the at least one detection embedding bit.

10. The method according to claim 9, characterized in that The predetermined characters contained in each of the plurality of paragraphs include: a first predetermined character added to the at least one detection embedding position, and a second predetermined character not included in the at least one detection embedding position; Restoring watermark information for each of the plurality of paragraphs based on the predetermined characters contained in each of the plurality of paragraphs and the position information of the at least one detection embedding bit includes: Watermark information of each paragraph is restored based on the first predetermined character and the position information of the at least one detection embedding position.

11. The method according to claim 1, characterized in that The adding the target watermark information to be added to the at least one embedding position further comprises: Generate a J-bit watermark code in a predetermined format based on the target watermark information, where J is a positive integer; Converting a j-bit target code in the J-bit watermark code into an invisible j-bit predetermined character, wherein the predetermined character includes a zero-width character; The j-bit predetermined characters are sequentially added to j target embedded positions in a plurality of embedded positions, wherein j is a positive integer less than or equal to J.

12. The method according to claim 1, characterized in that Performing text structure detection on the text to be processed to obtain text structure features of the text to be processed includes at least one of the following: Performing a first detection on the plurality of words based on a first detection model to obtain a plurality of first probabilities that each of the words belongs to a plurality of candidate grammatical structures, and determining the grammatical structure to which the word belongs based on the plurality of first probabilities; performing a second detection on the plurality of statements based on a second detection model to obtain a plurality of second probabilities that each of the statements belongs to a plurality of candidate logical structures, and determining the logical structure to which the statement belongs based on the plurality of second probabilities; A third detection is performed on the multiple paragraphs based on a third detection model to obtain multiple third probabilities that each of the paragraphs belongs to multiple candidate chapter structures, and the chapter structure to which the paragraph belongs is determined based on the multiple third probabilities.

13. The method according to claim 12, characterized in that The method further comprises: Acquire a training text, and perform word segmentation processing on a plurality of sample sentences in the training text to obtain a plurality of sample words; Determining the grammatical structure to which each of the plurality of sample words belongs, so as to generate a first label for the training text; A basic model is trained using the multiple sample words and the first label to obtain the first detection model, wherein the model parameters of the first detection model include first state transition information and first observation probability information, wherein the first state transition information is used to characterize the probability of transitioning from a first candidate grammatical structure to a second candidate grammatical structure, and the first observation probability information is used to characterize the respective probabilities of multiple target sample words belonging to the second candidate grammatical structure, and the first candidate grammatical structure and the second candidate grammatical structure belong to at least one of the multiple grammatical structures to which the multiple sample words belong.

14. The method according to claim 13, wherein: Performing a first detection on the multiple words based on the first detection model to obtain multiple first probabilities that each of the words belongs to multiple candidate grammatical structures includes: A plurality of first probabilities that each of the words belongs to a plurality of candidate grammatical structures are determined using the first state transition information and the first observation probability information.

15. The method according to claim 12, characterized in that The method further comprises: Acquire a training text, and perform sentence segmentation processing on a plurality of sample paragraphs in the training text to obtain a plurality of sample sentences; Determining the logical structure to which each of the plurality of sample sentences belongs, so as to generate a second label for the training text; A base model is trained using the multiple sample sentences and the second labels to obtain the second detection model, wherein model parameters of the second detection model include second state transition information and second observation probability information, wherein the second state transition information is used to characterize the probability of transitioning from the first candidate logical structure to the second candidate logical structure, and the second observation probability information is used to characterize the probability of each of the multiple target sample sentences belonging to the second candidate logical structure, and the first candidate logical structure and the second candidate logical structure belong to at least one of the multiple logical structures to which the multiple sample sentences belong.

16. The method according to claim 12, characterized in that The method further comprises: Acquire a training text, and segment the training text to obtain a plurality of sample paragraphs; Determining the chapter structure to which each of the plurality of sample paragraphs belongs, so as to generate a third label for the training text; The basic model is trained using the multiple sample paragraphs and the third label to obtain the third detection model, wherein the model parameters of the third detection model include third state transition information and third observation probability information, wherein the third state transition information is used to characterize the probability of transitioning from the first candidate paragraph structure to the second candidate paragraph structure, and the third observation probability information is used to characterize the probability of each of the multiple target sample paragraphs belonging to the second candidate paragraph structure, and the first candidate paragraph structure and the second candidate paragraph structure belong to at least one of the multiple paragraph structures to which the multiple sample paragraphs belong.

17. A text processing device, characterized in that: The device comprises: a detection module configured to perform text structure detection on a text to be processed to obtain text structure information of the text to be processed, wherein the text structure information includes at least one of the following: a grammatical structure to which a plurality of words in a sentence of the text to be processed belong, a logical structure to which a plurality of sentences in a paragraph of the text to be processed belong, and a chapter structure to which a plurality of paragraphs in a chapter of the text to be processed belong; a determination module, configured to determine, based on the correlation between the text structure information and the contextual semantics in the text to be processed, at least one embedding position in the text to be processed for adding watermark information; An adding module, configured to add target watermark information to be added to the at least one embedding position; The adding module includes a third generating submodule, a fourth generating submodule and a third adding submodule; The third generating submodule is used to generate a plurality of groups of reference information codes corresponding to the plurality of paragraphs according to the historical record information of adding target watermark information to the plurality of paragraphs included in each chapter of the text to be processed; The fourth generation submodule is configured to generate a chapter code for the chapter based on the multiple sets of reference information codes; the fourth generation submodule includes a first generation unit and a second generation unit, the first generation unit being configured to perform hash operations on the multiple sets of reference information codes respectively to generate multiple sets of paragraph codes corresponding to multiple paragraphs, and the second generation unit being configured to generate a chapter code for the chapter based on the multiple sets of paragraph codes; The third adding submodule is used to add at least part of the chapter code to a third embedding position; the third embedding position is the position of a target paragraph that belongs to a predetermined chapter structure and whose contextual semantic relevance meets predetermined conditions.

18. An electronic device, characterized in that: The electronic device comprises: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 16.

19. A non-volatile computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instructions are executed by a processor, the steps of the method according to any one of claims 1 to 16 are implemented.

20. A computer program product comprising a computer program, characterized in that The computer program implements the method according to any one of claims 1 to 16 when executed by a processor.

Citation Information

Patent Citations

  • Embedding method and extracting method for text watermark based on semantic role position mapping

    CN105205355A

  • Watermark adding method and device, storage medium and electronic equipment

    CN113538198A

  • Data generation method, text generation method and electronic equipment

    CN119323015A